Equipment state detection method and device, electronic equipment and storage medium
By inputting the bus device function triplet to the substrate controller, the number of errors can be corrected is collected, and the problem of speed reduction or bandwidth reduction cannot be alarmed during PCIe equipment operation is solved, real-time monitoring of device status and timely alarms are realized.
Patent Information
- Application Number
- CN202510495087.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when PCIe equipment has problems of speed reduction or bandwidth reduction during operation, it cannot be alerted in time.
By inputting the bus device function triplets of different root ports into the substrate controller, the substrate controller collects the number of correctable errors and determines the working status of the device based on the number of correctable errors, including normal, early warning and alarm.
Real-time monitoring of the operating status of PCIe devices and prompt alarms to avoid system downtime or data loss caused by speed reduction or bandwidth reduction problems.
Smart Images

Figure CN120336126A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a method, device, electronic device, and storage medium for detecting the status of a device. Background Art
[0002] The Peripheral Component Interconnect Express (PCIe) is a high-speed serial bus standard widely used in modern computer systems to connect hardware such as a central processing unit, a storage device, a graphics processing unit, and a network adapter. However, PCIe devices may experience problems such as speed reduction or bandwidth reduction during operation for various reasons.
[0003] The status of PCIe devices is monitored by the Basic Input / Output System (BIOS). When an abnormal status occurs in a PCIe device, the BIOS transmits the abnormal information to the Baseboard Management Controller (BMC), which is responsible for alarming. However, after the system starts up, the BIOS does not actively check the operating status of PCIe devices. Therefore, when a speed reduction or bandwidth reduction problem occurs in a PCIe device during the operation phase, alarming cannot be performed. Summary of the Invention
[0004] The present application provides a method, device, electronic device, and storage medium for detecting the status of a device, so as to at least solve the problem in the related art that alarming cannot be performed when a speed reduction or bandwidth reduction problem occurs in a device during the operation phase.
[0005] The present application provides a method for detecting the status of a device, including:
[0006] Inputting the bus device function triples of different root ports into a baseboard controller;
[0007] Collecting the number of correctable errors based on the baseboard controller through the bus device function triples; wherein, the correctable errors include the correctable errors of the root port and the correctable errors of the devices connected to the root port;
[0008] Determining the working status of the device according to the number of correctable errors; wherein, the working status includes normal, early warning, and alarm.
[0009] The present application further provides a device for detecting the status of a device, including:
[0010] An input unit, configured to input the bus device function triples of different root ports into a baseboard controller;
[0011] A collection unit for collecting the number of correctable errors based on a baseboard controller through a bus device function triple; wherein, the correctable errors include correctable errors of the root port and correctable errors of the devices connected to the root port.
[0012] A first determination unit for determining the working state of a device according to the number of correctable errors; wherein, the working state includes normal, warning, and alarm.
[0013] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the above-mentioned detection method of any device state when executing the computer program.
[0014] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of the above-mentioned detection method of any device state when executed by a processor.
[0015] This application also provides a computer program product including a computer program, and the computer program implements the steps of the above-mentioned detection method of any device state when executed by a processor.
[0016] Through this application, since the bus device function triples of different root ports are input into the baseboard controller, the baseboard controller can collect the correctable errors of the root port and the devices connected to the root port based on the bus device function triple, and determine the working state of the device according to the collected number of correctable errors, so as to evaluate the running status of the device. Therefore, the technical problem that the device cannot be alarmed in time when the speed or bandwidth decreases during operation in the prior art can be solved, and the technical effect of real-time monitoring of the device running state can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of a text generation method provided by an embodiment of this application;
[0019] Figure 2 It is an overall block diagram of collecting the number of correctable errors and states provided by an embodiment of this application;
[0020] Figure 3 It is a flowchart of collecting correctable errors provided by an embodiment of this application;
[0021] Figure 4 A schematic flowchart of a substrate controller collecting correctable errors provided by an embodiment of this application;
[0022] Figure 5 A schematic flowchart of determining the working state of a device provided by an embodiment of this application;
[0023] Figure 6 A schematic structural diagram of a device state detection device provided by an embodiment of this disclosure;
[0024] Figure 7 A schematic structural diagram of another device state detection device provided by an embodiment of this disclosure. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0026] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0027] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be made in conjunction with the accompanying drawings and specific implementation manners.
[0028] An embodiment of this application provides a device state detection method. In combination with the execution process of the device state detection method, the method will be described in detail.
[0029] Figure 1 A schematic flowchart of a text generation method provided by an embodiment of this application.
[0030] As Figure 1 shown, this method includes the following steps:
[0031] Step 101, input the bus device function triples of different root ports into the substrate controller.
[0032] The bus device function triple is a specific data structure used to identify and describe the functions of devices on a bus. It usually consists of three elements representing different information, including: the bus number, which is used to identify the specific bus to which the device is connected. In a complex computer system or other electronic device system, there may be multiple buses, and different buses are connected to different devices. The bus number can clarify which bus the device is on. The device number, on the same bus, there may be multiple devices connected, and the device number is used to uniquely identify each device on the bus. This can distinguish different devices, enabling the system to operate on specific devices, such as sending instructions, reading data, etc. The function number, each device may have multiple functions, and the function number is used to specify the specific function of the device. For example, a device may have both data transfer and control functions, and the function number can clarify which function of the device is to be used.
[0033] Through the combination of these three elements, the bus device function triple provides a clear and accurate identification method for the devices in the system, enabling relevant components such as the baseboard controller to easily identify, access, and control the devices on the bus, thereby realizing the detection of device status and other related operations. In the embodiment of this application, the bus device function triple is input into the baseboard controller, and the baseboard controller uses it to read the relevant registers of the devices connected to the root port and obtain information such as the number of correctable errors.
[0034] Step 102: Based on the baseboard controller, collect the number of correctable errors through the bus device function triple; among them, the correctable errors include the correctable errors of the root port and the correctable errors of the devices connected to the root port.
[0035] Correctable errors refer to a type of error situation that occurs during the operation of a computer system. Although these errors occur, the system has the ability to automatically repair them, without seriously affecting the normal operation of the system or causing serious consequences such as data loss or system crashes. In the embodiments of this application, correctable errors include two types. The first is the correctable error of the root port. The root port is the port in the computer system that is connected to the root complex, and it is the hub connecting the upper-level system and the lower-level devices in the device connection hierarchy of the system. During data transmission, due to reasons such as electrical interference, signal attenuation, or transient hardware failures, correctable errors may occur at the root port. For example, individual bits in data transmission may flip, but the system can detect and correct these errors through technical means such as error correction codes. The second is the correctable error of the devices connected to the root port. Various devices connected to the root port may also generate correctable errors during their own operation due to reasons such as internal circuit failures, software logic errors, or abnormal data interactions with the root port. For example, when a hard disk drive reads and writes data, data reading errors may occur due to slight defects on the surface of the disk platter, but the built-in error correction mechanism of the hard disk can correct these errors; another example is that when a network card receives network data packets, data errors may occur due to momentary fluctuations in the network signal.
[0036] Through the bus device function triple, the baseboard controller can traverse the root port and all the devices connected to it, collect the quantities of various correctable errors, and provide key data support for the subsequent state evaluation and maintenance of the system.
[0037] Step 103, determine the working state of the device according to the quantity of correctable errors; wherein, the working state includes normal, warning, and alarm.
[0038] In some embodiments, thresholds are preset as the basis for judging the working state of the device. When the quantity of correctable errors is at a relatively low level, that is, below the preset normal threshold range, it indicates that all components of the device are operating stably, data transmission is accurate, and it can perform various tasks normally according to the design requirements. When the quantity of correctable errors exceeds the normal threshold but has not reached the upper limit of the warning threshold, the system will determine the working state of the device as the warning state. This means that although the device is still running and the current correctable errors are not serious enough to affect its basic functions, it already shows that there may be some potential problems with the device, such as the aging of some components caused by long-term operation of the device, changes in conditions such as temperature, humidity, or electromagnetic interference in the operating environment that are not conducive to the stable operation of the device, or there are some small loopholes in the software system that have not yet caused serious errors.
[0039] When the number of correctable errors exceeds the upper limit of the warning threshold and reaches or exceeds the alarm threshold, the working state of the device is determined to be the alarm state. This indicates that the device is facing relatively serious problems, its performance has been significantly affected, and it may even malfunction at any time, resulting in serious consequences such as system downtime or data loss.
[0040] In some embodiments, when the system is in the warning and alarm states, preset measures can be provided to the user, such as displaying prominent prompt messages on the system management interface, sending warning emails or text messages to notify the system administrator, etc. Further, a strong audible and visual alarm can be issued.
[0041] In some embodiments, collecting the number of correctable errors by the baseboard controller through the bus device function triple includes:
[0042] In the case where the collection flag bit of the root port and the collection flag bit of the correctable error are unread, the baseboard controller reads the secondary bus and the subordinate bus of the root port; wherein, the collection flag bit is used to identify whether the root port and / or the correctable error has been collected;
[0043] The collection flag bit is a mechanism for identifying the status. In the embodiments of the present application, the collection flag bit is mainly used to mark whether the relevant information of the root port and the correctable error has been collected. In a computer system, the collection and processing of data are carried out in steps. In order to avoid duplicate collection or omission of information, a flag bit needs to be set. Each root port and correctable error has a corresponding collection flag bit, which is usually binary. The "unread" state can generally be represented by "0", meaning that the relevant information of the root port or the information of the correctable error has not been collected by the system; the "read" state is represented by "1", indicating that the information has been collected. It should be noted that this description method is only an exemplary illustration and is not a specific limitation on the marking of the collection flag bit.
[0044] The baseboard controller is one of the core components responsible for managing and controlling hardware resources in a computer system, and is used to coordinate the communication and data transmission between various hardware devices in the system. The baseboard controller can interact with different buses and devices, read their status information, and perform corresponding operations according to preset rules.
[0045] The secondary bus: is a bus connected under the root port, which further expands the connection ability of the root port. The secondary bus can connect multiple devices, and these devices communicate with the root port through the secondary bus, so as to realize interaction with the entire system.
[0046] Subordinate Bus: A bus that is subordinate to the secondary bus and can further subdivide the connected devices. Subordinate buses are usually used to connect devices with lower bandwidth requirements or relatively simple functions. Through this hierarchical bus structure, the device connections within the system can be effectively managed and organized.
[0047] When both the collection flag bit of the root port and the collection flag bit of the correctable error are unread, the baseboard controller performs a read operation. First, it reads from the BDF.dat file that stores the Bus Device Function (BDF) the number of root ports under each central processing unit and the BDF corresponding to each root port under each central processing unit and stores them; then it reads the secondary bus and subordinate bus corresponding to each root port and stores them.
[0048] Based on the baseboard controller reading the first correctable error count register of the root port, the correctable error count of the root port is obtained and stored in a preset table.
[0049] A register is a high-speed storage unit in a computer used to temporarily store data. The first correctable error count register is specifically used to store the correctable error count of the root port.
[0050] A preset table is a pre-defined data structure in the system used to store and manage relevant information of various devices. In the embodiments of this application, the preset table is used to store the correctable error count of the root port. The table usually has a specific format and fields, and each field corresponds to different information, such as the identifier of the root port, the correctable error count, etc. After the baseboard controller reads the correctable error count, it stores it in the corresponding position according to the format and rules of the preset table. In this way, the system can conveniently query, count, and analyze these data to timely understand the operating status of the root port and provide a basis for the maintenance and management of the system.
[0051] For example, the preset table may adopt the form of a database table, which contains two fields: "root port name" and "correctable error count". After the baseboard controller reads the correctable error count of a certain root port, it finds the corresponding record in the table according to the identifier of the root port and updates the correctable error count to the corresponding field.
[0052] In summary, based on the baseboard controller reading the first correctable error count register of the root port and storing the result in a preset table, it helps to ensure the stable operation and efficient management of the system.
[0053] Based on the baseboard controller, read the second correctable error count register of the device connected to the root port through the bus device function triple, secondary bus, and subordinate bus, and store the read result in a preset table.
[0054] Each device connected to the root port is equipped with a second correctable error count register, which is used to store the number of correctable errors that occur during the operation of the device. After the baseboard controller successfully reads the second correctable error count of the device, it will associate this data with the corresponding device identifier (determined by the bus device function triple) and store it in the corresponding position in the table according to the format and rules of the preset table. In this way, the system can conveniently manage, query, and analyze these data to timely grasp the operating status of the device and provide a basis for system maintenance and optimization.
[0055] Optionally, after reading the second correctable error count register of the device connected to the root port through the bus device function triple, secondary bus, and subordinate bus based on the baseboard controller and storing the read result in a preset table, the method further includes:
[0056] Store the bus device function triples of different root ports, the addresses of the first correctable error count registers, and the addresses of the second correctable error count registers in a preset data set, so that the baseboard controller can complete the collection of correctable errors by accessing the preset data set; where the preset data set includes the mapping relationship between the root port and the bus device function triple, the address of the first correctable error count register, and the address of the second correctable error count register;
[0057] Store the bus device function triples of different root ports, the addresses of the first correctable error count registers, and the addresses of the second correctable error count registers in a preset data set. Here, the bus device function triple, as mentioned before, is a unique identifier composed of a bus number, a device number, and a function number, which is used to accurately locate the device connected to the root port and its function. The first correctable error count register specifically stores the correctable error count of the root port itself, and the second correctable error count register is used to store the correctable error count of the device connected to the root port, and their addresses are the key identifiers for the baseboard controller to access these registers in the system memory.
[0058] The preset data set is a data set pre-set by the system for storing specific data. By storing the above key information in the preset data set, the system constructs a complete mapping relationship, that is, the preset data set. This preset data set details the correspondence between the root port and the bus device function triple, the address of the first correctable error count register, and the address of the second correctable error count register. When the baseboard controller performs subsequent correctable error collection, it only needs to access the preset data set to quickly and accurately obtain the required register addresses and the identification information of the devices, thereby greatly improving the efficiency and accuracy of data collection. For example, when the baseboard controller needs to collect the correctable error counts of a certain root port and its connected devices again, it can directly obtain the corresponding register addresses and device identifiers from the preset data set, without having to traverse each BDF under each RootPort of each central processor to determine whether the device exists.
[0059] Set the collection flag bit of the root port and the collection flag bit of the correctable error to read.
[0060] The collection flag bit of the root port is used to identify whether the relevant information of the root port has been collected, and the collection flag bit of the correctable error is used to identify whether the relevant information of the correctable error has been collected. Setting these two flag bits to read indicates that the system has completed the collection operation of the root port and its related correctable error information. It can avoid duplicate collection in subsequent data processing, and at the same time facilitate the system to manage and track the data collection status. For example, when the system is performing regular data checks or updates, it can quickly determine which root ports and correctable error information have been collected and which still need further processing by checking the status of these flag bits.
[0061] After completing the reading and storage of the correctable error count register of the device connected to the root port, storing the relevant information in the preset data set and setting the collection flag bit are important measures for the system to optimize the correctable error collection process and improve data management efficiency, which helps to ensure the stable operation of the system and the accuracy of data.
[0062] In some embodiments, based on the baseboard controller, reading the second correctable error count register of the device connected to the root port through the bus device function triple, the secondary bus, and the subordinate bus, and storing the reading result in the preset table further includes:
[0063] Based on the baseboard controller, reading the correctable error status register of the device connected to the root port through the bus device function triple, the secondary bus, and the subordinate bus to determine the status of the correctable error;
[0064] After determining the working status of the device according to the correctable error count, the method further includes:
[0065] When the operating state of the device is an alarm, determine the type of correctable error based on the state of the correctable error.
[0066] The correctable error status register is a storage unit inside the device for recording status information related to correctable errors. The baseboard controller accurately locates the target device using the bus device function triple, then communicates with it via the secondary bus and the slave bus, and sends a read instruction to the device to obtain the content in the correctable error status register.
[0067] The state of the correctable error includes various situations, such as the time when the error occurs, the frequency of error occurrence, the data area involved in the error, etc. These status information are crucial for deeply understanding the characteristics and potential impacts of the correctable error. For example, if the error occurs frequently, it may indicate potential hardware failures in the device or a poor operating environment; if the errors are concentrated in a specific data area, it may indicate problems with data storage or transmission in that area.
[0068] When it is determined that the operating state of the device is an alarm based on the number of correctable errors, it is necessary to further determine the type of correctable error based on the state of the correctable error.
[0069] The types of correctable errors are diverse. Common ones include parity errors, error checking and correcting single-bit errors, etc. Parity errors are usually caused by the flipping of bit positions during data transmission, resulting in a mismatch in the parity check result; error checking and correcting single-bit errors occur in a storage system that uses error checking and correcting technology, where a single bit error occurs but can be automatically corrected by the system.
[0070] Determining the error type by analyzing the state of the correctable error can provide targeted guidance for subsequent fault troubleshooting and repair. For example, if it is determined to be a parity error, it may be necessary to check whether there is interference in the data transmission line, whether the interface is loose, etc.; if it is an error checking and correcting single-bit error, it may be necessary to conduct a more in-depth detection of the storage device to determine whether there are problems such as hardware aging.
[0071] In some embodiments, before collecting the number of correctable errors by the baseboard controller using the bus device function triple, the method further includes:
[0072] When it is determined that the device fails to boot, clear the collection flag bit of the root port and the collection flag bit of the correctable error.
[0073] When it is determined that the device fails to boot, the collection flag bit of the root port and the collection flag bit of correctable errors are cleared. Since the Peripheral Component Interconnect Express (PCIe) devices may be plugged in or unplugged after each shutdown, the PCIe devices connected to the machine may be different after each boot. Therefore, after each boot, it is necessary to restart and read the Bus Device Function (BDF) of each Root Port of the CPU transmitted by the Basic Input Output System (BIOS), and traverse the connected PCIe devices according to the BDF.
[0074] When it is determined that the device boots successfully, determine whether the collection flag bit is set;
[0075] After it is determined that the device boots successfully, determine whether the collection flag bit is set. The setting of the collection flag bit is usually completed after receiving the collection instruction issued by the user. This step is to determine whether the user has a need to collect the number of correctable errors. If the collection flag bit is not set, it means that the user has not issued a collection instruction, and the system will not perform the operation of collecting the number of correctable errors.
[0076] When it is determined that the collection flag bit is set, determine the number of correctable errors collected based on the baseboard controller; wherein, the collection flag bit is set after receiving the collection instruction issued by the user;
[0077] When it is determined that the collection flag bit is set, determine the number of correctable errors collected based on the baseboard controller. This indicates that the user has issued a collection instruction, and the system will, according to the predetermined process, use the baseboard controller to collect the relevant number of correctable errors through the bus device function triplet.
[0078] Re-determine whether the device boots successfully. In the case of boot failure, set the collection status to failed, and clear the collection flag bit of the root port and the collection flag bit of correctable errors.
[0079] After deciding to collect the number of correctable errors, re-determine whether the device boots successfully. Ensure that the device is in a normal operating state during the collection process, because the normal operation of the device is the basis for accurately collecting the number of correctable errors. If it is found that the device fails to boot during this process, the system will set the collection status to failed, and at the same time clear the collection flag bit of the root port and the collection flag bit of correctable errors again. This can avoid continuing to perform invalid collection operations in the case of device abnormalities.
[0080] In summary, before collecting the number of correctable errors, the system ensures the accuracy and effectiveness of the collection operation by checking and processing the device power-on state and the collection flag bit. At the same time, it can also handle abnormal situations such as device power-on failure, ensuring the stability and reliability of the system.
[0081] In some embodiments, before reading the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triple, the secondary bus, and the subordinate bus, and storing the read result in a preset table, the method further includes:
[0082] Determining whether there is a connected device to the root port based on a tool instruction; wherein, the tool instruction is obtained through advance configuration.
[0083] The tool instruction is generated through advance configuration. During the development or deployment phase, technicians will configure the tool instruction according to factors such as the hardware architecture and device connection rules. This instruction can be a specific code logic, a preset command set, or a specific query method. The embodiments of the present application do not limit this. It can interact with the hardware and software of the system to obtain the connection status information of the root port.
[0084] In the case of determining that there is a connected device to the root port, based on the baseboard controller, read the second correctable error count register of the device connected to the root port through the bus device function triple, the secondary bus, and the subordinate bus, and store the read result in a preset table.
[0085] When it is determined that there is a connected device to the root port, subsequent read and storage operations will continue. That is, based on the baseboard controller, accurately find the target device through the bus device function triple, communicate with the device using the secondary bus and the subordinate bus, and read the value in its second correctable error count register. For the specific implementation method, please refer to the description of the above application embodiments in detail, and the embodiments of the present application will not elaborate here one by one.
[0086] If it is determined that there is no connected device to the root port, then there is no need to perform subsequent read and storage operations. The system will skip this step to avoid performing invalid operations, thereby improving the operating efficiency of the system.
[0087] In summary, before reading the second correctable error count register of the device connected to the root port, first determining whether there is a connected device to the root port based on the tool instruction can ensure the pertinence and effectiveness of subsequent operations, ensuring the stable and efficient operation of the system.
[0088] In some embodiments, in the case of failure in correctable error collection, set the collection status to failed and clear the collection flag bit.
[0089] During the entire correctable error collection process, collection failures may occur for various reasons. In the embodiments of the present application, the scenarios of correctable error failures include the following: device power-on failure; device power-on failure but the collection flag bit is not set; abnormal reading of the mobile phone flag bit, etc. When any of the above scenarios occurs, the collection flag bit will be cleared.
[0090] To clearly illustrate a method for detecting the status of a device provided in the embodiments of the present application, the following uses an embodiment for illustration.
[0091] The BMC provides three commands and a correctable error (CE) collection module.
[0092] Please refer to Figure 2 , Figure 2 , which is an overall block diagram for collecting the number and status of correctable errors provided in the embodiments of the present application. As Figure 1 shown, it includes: Command 1: Start collection command. When an operation and maintenance personnel accesses the BMC through this command, the BMC will set the CE collection flag bit to 1. The CE collection module is a thread that will continuously detect whether the flag bit is 1. If it is not 1, the thread will keep idling without performing any operations. If it is 1, it means the user hopes to collect the CE number and status. At this time, start collecting the CE number and status. After successful collection, clear the collection flag bit and set the collection status to "success". After collection failure, also clear the collection flag bit and set the collection status to "failure";
[0093] Command 2: Get collection status. When an operation and maintenance personnel accesses the BMC through this command, the BMC determines whether the collection flag bit is 0. If it is not 0, it means that collection is currently in progress and has not ended; if it is 0, it means that collection has ended. Then continue to judge the collection status. If the collection status is "failure", it means that collection has failed. If the collection status is "success", it means that collection has succeeded. If collection is successful, the collected CE number and status can be obtained;
[0094] Command 3: Download file command. If collection is successful when executing Command 2, the file of the collected CE number and status can be downloaded through this command. Then, whether there is a problem of speed reduction or bandwidth reduction can be predicted by parsing the CE number in the file, and the status of the CE can be viewed to determine the CE.
[0095] The CE collection module is further divided into two parts. The first part is that the BIOS transfers the device BDF (Bus Device Function) to the BMC during the boot process; the second part is that the BMC collects the CE number and status
[0096] Please refer toFigure 3 , Figure 3 is a flowchart for collecting errors that can be corrected provided by an embodiment of this application. Please refer to Figure 3 . It includes: During the boot process, the BIOS transfers the device BDF to the BMC; when booting, the BIOS traverses each RootPort of each CPU and transfers the bus device function triple of each RootPort to the BMC; after receiving the data passed by the BIOS, the BMC counts the number of RootPorts of each CPU, records the number and the bus device function triples of all RootPorts. In some embodiments, it can be recorded in a file, and the file name can be set to BDF.dat; it should be noted that this description method is only an exemplary illustration and not a specific limitation on the specific file name.
[0097] When shutting down or restarting, delete the BDF.dat file that records the bus device function triples of RootPorts. Because the BIOS transfers the BDF to the BMC every time it boots, if the file is not deleted when shutting down or restarting, the file will repeatedly save the BDF data.
[0098] Please refer to Figure 4 , Figure 4 is a schematic flowchart of a baseboard controller for collecting correctable errors provided by an embodiment of this application. As Figure 4 shown, it includes:
[0099] (1) Detect whether the boot is completed. If the boot is not completed, clear the BDF reading flag and the CE status reading flag, and clear the BDF information; After each shutdown, the PCIe devices may be plugged or unplugged, so the PCIe devices plugged into the machine may be different after each boot. Therefore, after each boot, it is necessary to restart and read the BDF of each RootPort of the CPU passed by the BIOS, and traverse the connected PCIe devices according to the BDF.
[0100] (2) Determine whether the collection flag bit is set. If it is set, continue to execute downward to start collection. If it is not set, the thread runs idle and does not execute collection downward.
[0101] (3) Detect again whether the boot is completed. If the boot is not completed, set the collection status to "failed", clear the collection flag bit, and end the collection; When the boot is not completed, the BIOS does not transfer the BDF information to the BMC, and the BMC cannot access the PCIe devices.
[0102] (4) Check if BDFReadFlag is 0. If it is not 0, it means the information has been read, and this step can be skipped directly to execute the next step. If it is 0, it means the information has not been read. First, read the number of RootPorts under each CPU and the BDF corresponding to each RootPort under each CPU from the BDF.dat file that stores BDF and store them. Then, read the SecondaryBus (secondary bus) and SubordinateBus (subordinate bus) corresponding to each RootPort and store them. After that, set BDFReadFlag to 1. (Each device connected under a RootPort is represented by BDF. The range of Bus is 0 to 255, the range of Device is 0 to 31, and the range of Function is 0 to 7. When traversing the devices under a RootPort, we can traverse the above ranges of Bus, Device, and Function to find if there is a device. However, the Bus range of the devices connected under each RootPort must be between SecondaryBus and SubordinateBus. Therefore, reading SecondaryBus and SubordinateBus first can narrow the traversal range and greatly shorten the traversal time. The reading method can be through the peciapp tool under BMC, such as peciapp -a CPUIndex rdpciconfig B D F Reg, or it can also be read using the API interface under BMC.)
[0103] (5) Create a file for storing the number and status of CEs. Before creating this file, delete it first to ensure that no information is stored in this file before each collection.
[0104] For each BDF, two pieces of information will be stored, such as:
[0105] Central Processing Unit: 0, RootPort: 0, B (Bus): 0x05, D (Device): 0x01, F (Function): 0x00, Number of CEs: 0x00000000.
[0106] (6) Check if CEStatusReadFlag is 0. If CEStatusReadFlag is not 0, it means the number and status of CEs have never been read, and execute step (7). Otherwise, it means it has been read once. At this time, the BDFs of the PCIe devices connected under all RootPorts under the CPU and the registers corresponding to the number and status of CEs that need to be read have been saved to the array BDFInfo. At this time, directly execute step (11).
[0107] (7) Traverse each RootPort under each CPU. First, read the CE count register storing the RootPort to obtain the CE count of this RootPort (the reading method is the same as that in (4)), and store the value in the file in the format in (5); if the value of CE is not 0, then clear the value of this register after reading to ensure that the CE read each time is newly generated during this period; finally, read the CE status register of each RootPort and store the value of the CE status register in the file in the format in (5).
[0108] (8) Determine whether a device is connected under this RootPort (it can be through the commandpeciapp-aCPUIndexrdpciconfigSecondoryBus 0x00 0x000x00. If the return is not error, it means that a PCIe device is connected. Of course, it can also be judged through the API interface).
[0109] (9) If a PCIe device is connected, then traverse each BDF of this RootPort and access this BDF (the range of Bus is from SecondaryBus to SubordinateBus calculated in (4), the range of Device (device) is 0 to 31, and the range of Function is 0 to 7). If this BDF exists, first calculate the offsets of the CE count register and the CE status register of this BDF; (there are many PCIe devices, so the offsets are not fixed. The CE count registers and CE status registers of these PCIe devices are stored in a linked list, so the linked list needs to be traversed to find the CE count register and the CE status register); then read the CE count register of this device and store it in the file. If the value of CE is not 0, then clear the value of this register after reading; finally, read the value of the CE status register of each RootPort and store it in the file.
[0110] (10) Store all the BDFs and the corresponding registers under each RootPort of each CPU into BDFInfo. After that, there is no need to traverse each BDF under each RootPort of each CPU to judge whether this BDF exists, because there are many BDFs and it takes a long time to traverse once.
[0111] (11) Traverse BDFInfo, take out the CPU RootPort, Bus, Device, Function and registers stored in it, obtain the count and status of the corresponding CE through the command, then store them in the file, and set CEStatusReadFlag to 1.
[0112] After the collection is completed, set the collection status to "success" and clear the collection flag to end the collection.
[0113] Any of the above steps may fail. If it fails, set the collection status to "failure" and clear the collection flag to end the collection.
[0114] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a process for determining the working state of a device provided by an embodiment of the present application. As Figure 5 shown, it includes:
[0115] The operation and maintenance personnel obtain files of the storage BDF, the corresponding CE quantity, and the CE status by providing three commands through the BMC every hour.
[0116] Parse the file to obtain the CE quantity corresponding to all BDFs.
[0117] Make a judgment based on the CE quantity. If it is 20 times or less per hour, it indicates a normal state. If it is between 20 times and 100 times per hour, it indicates a warning state, and a warning alarm for speed reduction and bandwidth reduction is generated. If it exceeds 100 times, it indicates a serious state, and a serious alarm for speed reduction and bandwidth reduction is generated.
[0118] When a warning alarm or a serious alarm is generated, parse the correctable error types corresponding to the BDFs in the file to check which BDFs have what kind of CEs, so that the operation and maintenance personnel can perform maintenance.
[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0120] An embodiment of the present application also provides a detection device for the device state. Figure 6 which is a schematic structural diagram of a detection device for the device state provided by an embodiment of the present disclosure. As Figure 6 shown, it includes:
[0121] An input unit 21, configured to input the bus device function triples of different root ports into the baseboard controller.
[0122] A collection unit 22, configured to collect the quantity of correctable errors based on the baseboard controller through the bus device function triples; wherein, the correctable errors include the correctable errors of the root ports and the correctable errors of the devices connected to the root ports.
[0123] A first determination unit 23, configured to determine the working state of the device according to the quantity of correctable errors; wherein, the working state includes normal, warning, and alarm.
[0124] Further, in a possible implementation manner of the embodiment of the present disclosure, the collection unit 22 is further configured to:
[0125] When the collection flag bit of the root port and the collection flag bit of the correctable error are unread, read the secondary bus and the slave bus of the root port based on the baseboard controller; wherein, the collection flag bit is used to identify whether the root port and / or the correctable error is collected;
[0126] Read the correctable error count register of the root port based on the baseboard controller to obtain the correctable error count of the root port, and store it in a preset table;
[0127] Based on the baseboard controller, read the second correctable error count register of the device connected to the root port through the bus device function triple, the secondary bus and the slave bus, and store the read result in a preset table.
[0128] Further, in a possible implementation manner of the embodiment of the present disclosure, as Figure 7 shown, the device further includes:
[0129] A storage unit 24, configured to, after the collection unit 22 reads the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triple, the secondary bus and the slave bus, and stores the read result in a preset table, store the bus device function triples of different root ports, the addresses of the first correctable error count registers, and the addresses of the second correctable error count registers in a preset data set, so that the baseboard controller completes the collection of correctable errors by accessing the preset data set; wherein, the preset data set includes the mapping relationship between the root port and the bus device function triple, the address of the first correctable error count register, and the address of the second correctable error count register;
[0130] A setting unit 25, configured to set the collection flag bit of the root port and the collection flag bit of the correctable error to read.
[0131] Further, in a possible implementation manner of the embodiment of the present disclosure, the collection unit 22 is further configured to:
[0132] Read the correctable error status register of the device connected to the root port based on the baseboard controller through the bus device function triple, the secondary bus and the slave bus to determine the status of the correctable error;
[0133] After determining the working status of the device according to the correctable error count, the device further includes:
[0134] When the working status of the device is an alarm, determine the type of the correctable error based on the status of the correctable error.
[0135] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 7 shown, the apparatus further includes:
[0136] A first processing unit 26, configured to clear the collection flag bit of the root port and the collection flag bit of correctable errors before the collection unit 22 collects the number of correctable errors based on the baseboard controller through the bus device function triple when it is determined that the device startup fails;
[0137] A second determination unit 27, configured to determine whether the collection flag bit is set when it is determined that the device startup is successful;
[0138] A third determination unit 28, configured to determine the number of correctable errors collected based on the baseboard controller when it is determined that the collection flag bit is set; wherein, the collection flag bit is set after receiving a collection instruction issued by the user;
[0139] A second processing unit 29 is further configured to re-determine whether the device startup is successful. When the startup fails, set the collection status to failed, and clear the collection flag bit of the root port and the collection flag bit of correctable errors.
[0140] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 7 shown, the apparatus further includes:
[0141] A fourth determination unit 210, configured to determine whether there is a device connected to the root port based on a tool instruction before the collection unit 22 reads the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triple, secondary bus, and subordinate bus, and stores the read result in a preset table; wherein, the tool instruction is obtained through pre-configuration;
[0142] A reading unit 211, configured to read the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triple, secondary bus, and subordinate bus, and store the read result in a preset table when it is determined that there is a device connected to the root port.
[0143] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 7 shown, the apparatus further includes:
[0144] A third processing unit 212 is further configured to set the collection status to failed and clear the collection flag bit when the collection of correctable errors fails.
[0145] For the description of the features in the corresponding embodiments of the device state detection device, reference can be made to the relevant description in the corresponding embodiments of the device state detection method, which will not be elaborated here one by one.
[0146] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the device state detection method.
[0147] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the device state detection method when running.
[0148] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical disks, etc., various media that can store computer programs.
[0149] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the device state detection method.
[0150] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the device state detection method.
[0151] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0152] The above has introduced in detail a method, device, electronic device, and storage medium for detecting the state of a device provided in this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for detecting the state of a device, characterized in that, Including: Input the bus device function triples of different root ports into the baseboard controller; Collect the number of correctable errors based on the baseboard controller through the bus device function triples; wherein, the correctable errors include the correctable errors of the root port and the correctable errors of the devices connected to the root port; Determine the working state of the device according to the number of correctable errors; wherein, the working state includes normal, warning and alarm.
2. The method for detecting the device state according to claim 1, wherein, The collecting the number of correctable errors based on the baseboard controller through the bus device function triples includes: When the collection flag bits of the root port and the correctable errors are unread, read the secondary bus and the subordinate bus of the root port based on the baseboard controller; wherein, the collection flag bit is used to identify whether the root port and / or the correctable errors are collected; Read the first correctable error count register of the root port based on the baseboard controller, obtain the number of correctable errors of the root port, and store it in a preset table; Based on the baseboard controller, read the second correctable error count register of the device connected to the root port through the bus device function triples, the secondary bus and the subordinate bus, and store the read result in the preset table.
3. The method for detecting the device state according to claim 2, characterized in that After reading the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triples, the secondary bus and the subordinate bus, and storing the read result in the preset table, the method further includes: Store the bus device function triples of different root ports, the address of the first correctable error count register, and the address of the second correctable error count register into a preset data set, so that the baseboard controller can complete the collection of correctable errors by accessing the preset data set; wherein, the preset data set includes the mapping relationship between the root port and the bus device function triples, the address of the first correctable error count register, and the address of the second correctable error count register; Set the collection flag bits of the root port and the correctable errors to read.
4. The method for detecting the device state according to claim 2, wherein The reading the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triples, the secondary bus and the subordinate bus, and storing the read result in the preset table further includes: Read the correctable error status register of the device connected to the root port based on the baseboard controller through the bus device function triples, the secondary bus and the subordinate bus, and determine the status of the correctable errors; After determining the working state of the device according to the number of correctable errors, the method further includes: When the working state of the device is alarm, determine the type of the correctable error based on the status of the correctable error.
5. The method for detecting the device state according to claim 2, wherein Before collecting the number of correctable errors based on the baseboard controller through the bus device function triples, the method further includes: When it is determined that the device fails to boot, clear the collection flag bit of the root port and the collection flag bit of the correctable error. When it is determined that the device boots successfully, determine whether the collection flag bit is set. When it is determined that the collection flag bit is set, determine the number of correctable errors collected based on the baseboard controller; wherein, the collection flag bit is set after receiving a collection instruction issued by the user. Re-determine whether the device boots successfully. When the boot fails, set the collection status to failed, and clear the collection flag bit of the root port and the collection flag bit of the correctable error.
6. The method for detecting the device state according to claim 2, wherein, Before reading the second correctable error count register of the device connected to the root port through the bus device function triple, the secondary bus, and the subordinate bus based on the baseboard controller and storing the read result in the preset table, the method further includes: Determine whether there is a device connected to the root port based on a tool instruction; wherein, the tool instruction is obtained through pre-configuration. When it is determined that there is a device connected to the root port, read the second correctable error count register of the device connected to the root port based on the baseboard controller through the bus device function triple, the secondary bus, and the subordinate bus, and store the read result in the preset table.
7. The method for detecting the device state according to any one of claims 1-6, characterized in that, The method further includes: When the collection of correctable errors fails, set the collection status to failed, and clear the collection flag bit.
8. A detection device for the state of a device, characterized in that, Includes: An input unit for inputting the bus device function triples of different root ports into the baseboard controller. A collection unit for collecting the number of correctable errors based on the baseboard controller through the bus device function triple; wherein, the correctable errors include the correctable errors of the root port and the correctable errors of the device connected to the root port. A first determination unit for determining the working state of the device according to the number of correctable errors; wherein, the working state includes normal, warning, and alarm.
9. An electronic device, characterized in that, Includes: A memory for storing a computer program. A processor for implementing the steps of the method for detecting the device state according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the method for detecting the device state according to any one of claims 1 to 7 when executed by a processor.