Equipment information management method and equipment information management device
By combining BIOS enumeration and BMC with the I2C/SMBus bus to obtain PCIe device information, the problem of incomplete information coverage is solved, achieving comprehensive information coverage and real-time monitoring, reducing operation and maintenance costs, and ensuring server stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAMEN YUANCHOU INTELLIGENT COMPUTING TECHNOLOGY CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
The incomplete information acquisition scope of existing PCIe devices leads to increased maintenance costs and risks such as server performance degradation and business interruption, which seriously restricts the stable operation and large-scale application of high-density PCIe device server clusters.
The manufacturing information of PCIe devices is obtained by enumerating the BIOS, and the functional information is obtained by the BMC through the I2C and SMBus buses. The two are then combined to form a target information set, achieving comprehensive information coverage.
It achieves comprehensive information coverage of PCIe devices, reduces operation and maintenance costs, improves the real-time nature and convenience of device status monitoring, and ensures stable server operation.
Smart Images

Figure CN121833587A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of server hardware management technology, and in particular to a method and apparatus for managing device information. Background Technology
[0002] Against the backdrop of the large-scale deployment of artificial intelligence (AI) servers, storage-intensive servers, and heterogeneous computing servers, the demand for PCIe device expansion on a single server continues to rise, typically requiring the simultaneous connection and management of 4 to 8 heterogeneous PCIe devices. A typical configuration includes 4 graphics processing units (GPUs), 2 high-performance solid-state drives (SSDs), 1 RAID controller, and 1 high-speed Ethernet / InfiniBand network interface card, covering the four core dimensions of computing, storage, control, and networking.
[0003] However, the information obtained by current PCIe devices is not comprehensive, which can lead to increased maintenance costs and risks such as server performance degradation and business interruption due to the failure to detect device anomalies in a timely manner. This seriously restricts the stable operation and large-scale application of high-density PCIe device server clusters. Summary of the Invention
[0004] This application provides a method and apparatus for managing device information, in order to at least solve the technical problem that the information obtained by current PCIe devices is not comprehensive in terms of coverage.
[0005] This application provides a method for managing device information, comprising: acquiring information about a high-speed serial computer extended bus standard device using a basic input / output system enumeration method to obtain first information, wherein the first information is manufacturing information of the high-speed serial computer extended bus standard device; acquiring information about the high-speed serial computer extended bus standard device using a baseboard management controller via an I2C bus and an SMBus bus to obtain second information, wherein the second information is functional information of the high-speed serial computer extended bus standard device; merging the first information and the second information to obtain a target information set; and sending the target information set to a target device used by a target object, so that the target object monitors the status of the device based on the target information set.
[0006] This application also provides a device information management apparatus, comprising: a first acquisition unit, configured to acquire information of a high-speed serial computer extended bus standard device using a basic input / output system enumeration method to obtain first information, wherein the first information is manufacturing information of the high-speed serial computer extended bus standard device; a second acquisition unit, configured to acquire information of the high-speed serial computer extended bus standard device using a baseboard management controller via an I2C bus and an SMBus bus to obtain second information, wherein the second information is functional information of the high-speed serial computer extended bus standard device; a merging module unit, configured to merge the first information and the second information to obtain a target information set; and a sending unit, configured to send the target information set to a target device used by a target object, so that the target object monitors the status of the device based on the target information set.
[0007] This application addresses the issue of comprehensive coverage of basic device information by enumerating PCIe devices through the BIOS and reading VPD information, including device manufacturing information such as device ID, manufacturer, and product name. VPD information is an inherent attribute of the device, extracted by the BIOS at the outset. The BMC uses the I2C and SMBus buses to read the device's FRU information, which covers functional information such as serial number, firmware version, and hardware status. By reading data through the BMC, the deficiencies of VPD information in functional description and status monitoring are compensated for. The VPD information obtained by the BIOS is merged with the FRU information obtained by the BMC to form a target information set. This target information set contains both manufacturing and functional information of the device, achieving comprehensive information coverage of the PCIe device. Therefore, it solves the technical problem of incomplete information coverage in current PCIe device acquisition technologies. Attached Figure Description
[0008] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a hardware structure block diagram of a mobile terminal for a device information management method according to an embodiment of this application;
[0010] Figure 2 This is a flowchart illustrating a method for managing device information according to an embodiment of this application;
[0011] Figure 3 This is a schematic diagram of the overall system architecture;
[0012] Figure 4 This is a schematic diagram of the overall process;
[0013] Figure 5 This is a diagram illustrating the information acquisition process;
[0014] Figure 6 This is a diagram illustrating BIOS access;
[0015] Figure 7 This is a diagram illustrating BMC access;
[0016] Figure 8 This is a structural block diagram of a device for managing device information according to an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0018] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0019] Existing PCIe device management solutions have three major technical shortcomings: First, the coverage is not comprehensive, with some solutions only supporting specific types of PCIe devices (such as managing only GPUs or storage devices), failing to achieve unified management of all types of devices; second, the information reliability is poor, as PCIe device information is obtained through only one method and lacks a multi-method comparison mechanism; and third, the device information collection is incomplete, as key information of PCIe devices is often stored in FRUs (Field Replaceable Units), and current technology lacks a stable, dynamic, and reliable means to obtain this information.
[0020] The aforementioned problems directly lead to prominent issues such as data inconsistency, long troubleshooting cycles, and frequent configuration conflicts during server operation and maintenance. This not only significantly increases the manpower and time costs of operation and maintenance but also poses risks such as server performance degradation and business interruption due to undetected equipment anomalies, severely restricting the stable operation and large-scale application of high-density PCIe device server clusters. Therefore, this invention aims to propose a comprehensive server PCIe device information management method that covers a wide range of devices, verifies information reliability through multiple methods, and addresses the core pain points of existing technologies.
[0021] Existing technical solutions provide a method and apparatus for dynamically identifying PCIe devices using a BMC. The method includes, after completing the power-on self-test (POST), the BIOS sending PCIe device information to the BMC, with the information sent according to a preset file format; the BMC parsing the PCIe device information to obtain the current PCIe device's location information and the corresponding I2C bus on the motherboard; and calling the corresponding I2C command to obtain the dynamic information of the current PCIe device. This invention, by pre-setting the content specifications for the BMC-BIOS protocol interaction, enables the BMC to obtain the I2C bus of each PCIe device on the motherboard, thereby achieving dynamic loading of PCIe devices to read dynamic and static information, meeting the needs of server administrators who want to efficiently manage PCIe devices through the BMC. It also solves the problem that the BMC cannot directly access the internal information of PCIe devices.
[0022] This solution essentially relies on the BMC (Browser Control Center) to obtain information about PCIe devices via the physical I2C bus. In this patent, the BIOS primarily collects and iterates through the Basic Location Information (BDF) of external PCIe devices connected to the server. This leads to the following problems:
[0023] 1. Poor flexibility. When external PCIe devices change, the BMC code needs to be adjusted to re-adapt to different configurations and ensure the accuracy of information acquisition;
[0024] 2. Limited coverage and device limitations. This technology is only applicable to PCIe devices with I2C interfaces. Some devices do not have an I2C interface or do not provide a dedicated FRU ROM, thus preventing the BMC from obtaining PCIe device information.
[0025] 3. Poor information accuracy. FRU information for PCIe devices is obtained solely through the BMC physical connection.
[0026] Secondly, existing technical solutions also provide a method and apparatus for detecting PCIe devices, a BIOS, and a storage medium. The method includes: acquiring a first information set, wherein the first information set includes device information of each PCIe device included in a first PCIe device set of the target machine; receiving a second information set sent by a Baseboard Management Controller (BMC), wherein the second information set includes device information of each PCIe device included in a second PCIe device set of the target machine acquired by the BMC; matching the first information set with the second information set, and determining the abnormal PCIe device if the first information set and the second information set do not match. This application solves the problem of errors in detecting PCIe device information in related technologies, thereby achieving a more accurate detection of PCIe device information.
[0027] This technology essentially involves the BIOS interacting with the system slot information stored in SMBIOS Type 9 and the PCIe information maintained by the BMC. This leads to the following problems:
[0028] 1. Incomplete information. According to the official SMBIOS documentation, SMBIOS Type 9 records system slot information, including BDF information, slot bandwidth, and whether it is in place. This information is insufficient for basic operation and maintenance needs. For example, key information such as the device PN and SN serial numbers, UUID, manufacturer information, and specific device model recorded in the FRU cannot be obtained.
[0029] 2. Limited coverage and device limitations. Smbios Type 9 can only record system slot information. For heterogeneous models (front panel + switch + rear panel), Type 9 cannot record device information on the rear panel.
[0030] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] The specific application environment architecture or specific hardware architecture on which the execution of the device information management method depends is described here.
[0032] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a device information management method according to an embodiment of this application. Figure 1 As shown, the server device may include one or more ( Figure 1Only one is shown in the image. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor (MCU), or programmable logic device (FPGA), etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0033] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the device information management method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0035] The embodiments of this application provide a method for managing device information. The method is described in detail below in conjunction with the execution flow of the device information management method.
[0036] The following explains the technical terms used in this application:
[0037] PCIE, short for Peripheral Component Interconnect Express, is a high-speed serial computer expansion bus standard.
[0038] BIOS: BIOS (Basic Input / Output System) is a type of firmware embedded in the ROM (Read-Only Memory) chip on the computer motherboard. It acts as a bridge between computer hardware and software, responsible for initializing hardware, testing hardware functions, loading the operating system and boot program, etc., when the computer is turned on. BIOS is one of the earliest programs loaded and executed during the computer startup process, providing the most basic and direct hardware control of the computer.
[0039] BMC, short for Baseboard Management Controller, is a core component in the server management system defined by the Intelligent Platform Management Interface (IPMI) protocol. It is a hardware manager integrated into servers, network devices, and other computer systems, primarily responsible for monitoring the hardware status of devices, performing remote management operations, and providing monitoring and control functions for devices.
[0040] RP: PCIe Root Port. In the PCIe architecture, the host (usually the CPU or memory controller) communicates with PCIe devices through the PCIe Root Complex. The PCIe Root Port is part of the Root Complex and can be seen as the starting point for the host system to extend outwards to PCIe devices.
[0041] FRU (Field Replaceable Unit) is a standardized hardware identity and status identification system designed for independently replaceable electronic devices. Its core value lies in providing a unified, traceable, and resolvable full lifecycle information carrier for PCIe devices, serving as a key data foundation for the unified management of server PCIe devices. It is typically stored in the device's EEPROM (Electrically Erasable Programmable Read-Only Memory) or SPI Flash, and can be read via the I2C bus or PCIe Configuration Space without relying on device drivers for loading.
[0042] I2C bus: The full name of I2C bus is Inter-Integrated Circuit Bus. It is a serial, half-duplex, multi-master multi-slave communication bus used for short-distance data transmission between chips.
[0043] SMBus bus: SMBus stands for System Management Bus, which is a serial bus protocol used for low-speed communication.
[0044] This embodiment provides a method for managing device information. Figure 2 This is a flowchart of a device information management method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:
[0045] Step S202: Information on the high-speed serial computer extended bus standard device is obtained by enumerating the basic input / output system to obtain first information, wherein the first information is the manufacturing information of the product of the high-speed serial computer extended bus standard device.
[0046] Specifically, when the server starts up, the Basic Input / Output System (BIOS) sends an enumeration command to the PCIe bus and identifies the connected PCIe devices one by one. By accessing the device's configuration space, the BIOS can read the device's manufacturing information, such as the Device ID and Vendor Code, as well as basic attributes such as the device name and version number.
[0047] This embodiment overcomes the limitations of a single information source. Through the built-in PCIe protocol stack, the BIOS can obtain the VPD information of all PCIe devices before any operating system is loaded, achieving preliminary comprehensive coverage of device information.
[0048] Step S204: The baseboard management controller obtains information about the high-speed serial computer expansion bus standard device through the I2C bus and SMBus bus to obtain second information, which is the functional information of the high-speed serial computer expansion bus standard device.
[0049] Specifically, after the server is powered on, the Baseboard Management Controller (BMC) initializes the I2C and SMBus bus controllers and obtains the FRU information of the PCIe device by sending a read command to a preset address. This information includes the device's functional information, such as its working status, health status, device serial number, and production batch number.
[0050] This embodiment solves the problem of limited information acquisition channels. By physically connecting the I2C and SMBus buses, the BMC can bypass the operating system and interact directly with the hardware to obtain more comprehensive functional information.
[0051] Step S206: Merge the first information and the second information to obtain the target information set;
[0052] Specifically, the first set of information (manufacturing information) obtained by the BIOS and the second set of information (functional information) obtained by the BMC are merged to form a target information set containing complete device information.
[0053] This embodiment solves the problem of information dispersion by aggregating information from different sources and providing a unified and comprehensive view of device information management.
[0054] Step S208: The target information set is sent to the target device used by the target object, so that the target object can monitor the status of the device based on the target information set.
[0055] Specifically, BMC sends the target information set to the target device, such as a server management system or a management platform used by operations and maintenance personnel, via IPMI, SMBIOS, or a RESTful API interface.
[0056] This embodiment improves the real-time performance and convenience of device monitoring; the real-time transmission and remote access of the target information set enable maintenance personnel to obtain the real-time status of the device without physical contact with the server, reducing maintenance costs and improving the real-time performance and ease of operation of server device monitoring.
[0057] Through the above steps, the BIOS enumerates PCIe devices and reads VPD information, including device manufacturing information such as device ID, manufacturer, and product name. VPD information is an inherent attribute of the device, extracted by the BIOS at the beginning, ensuring comprehensive coverage of basic device information. The BMC uses the I2C bus and SMBus bus to read the device's FRU information, which covers the device's functional information such as serial number, firmware version, and hardware status. By reading data through the BMC, the deficiencies of VPD information in functional description and status monitoring are compensated for. The VPD information obtained by the BIOS and the FRU information obtained by the BMC are merged to form a target information set. The target information set contains the device's manufacturing and functional information, achieving comprehensive information coverage of PCIe devices. Therefore, it can solve the technical problem of incomplete information coverage of current PCIe devices in related technologies.
[0058] Compared with similar technologies, the core of this application is the proposal to use BIOS to read server PCIe device configuration space information, especially the critical FRU VPD information, via MMIO. This overcomes the limitations of BMC hardware access and the incomplete information problem of BIOS smbios type9. Specifically:
[0059] 1. Addressing the data loss risk caused by the single-path approach in existing technologies (referring to the BMC control used in current mainstream technologies), this solution constructs a dual-main architecture of "BIOS-PCIe enumeration + BMC-I2C / SMBu reading," achieving full coverage of VPD+FRU information for all types of PCIe devices (GPU / SSD / RAID / network cards);
[0060] 2. To address the issues of information collaboration and real-time synchronization, this solution designs a multi-device information collaboration module, which realizes automatic verification, classified storage, and dynamic updating of VPD and FRU data, avoiding errors from manual comparison.
[0061] 3. To address multi-bus compatibility and hot-swap adaptation issues, this solution is compatible with I2C / SMBus / PCIe multi-bus protocols and supports real-time information updates during PCIe device hot-swapping to ensure uninterrupted service.
[0062] 4. To address the issue of low operation and maintenance efficiency in existing technologies, this solution provides a unified information access interface (IPMI+SMBIOS+RESTful API), supporting remote batch query and visual management, reducing tool switching and on-site operations.
[0063] This system is designed based on an x86 architecture server, and its overall architecture is as follows: Figure 3 As shown, the core consists of a four-layer structure: "hardware layer - bus layer - coordination layer - access layer", including:
[0064] Hardware layer: server motherboard, PCIe device set (GPU×4, SSD×2, RAID controller×1, network card×1), BMC chip, BIOS chip, information storage module (BMC-NVRAM + motherboard RAM).
[0065] Bus layer: PCIe bus (BIOS communicates with all PCIe devices), I2C bus (BMC communicates with network card FRU), SMBus bus (BMC communicates with RAID card FRU).
[0066] Collaboration Layer: Multi-device information collaboration module (integrated into BMC, with an ARM Cortex-M4 processor at its core, including enumeration scheduling, information verification, and hot-swappable adapter sub-modules);
[0067] Access layer: IPMI interface (remote operation and maintenance), SMBIOS interface (OS local call), Redfish (cloud platform integration).
[0068] The overall process of this system is as follows: Figure 4 As shown, after the server starts:
[0069] 1. BMC enumerates PCIe devices: BMC polls all PCIe external devices on the server through methods such as MCTP and records the BDF information of the PCIe devices. This information is unique and can be used as an identifier for device identification.
[0070] 2. Information Acquisition: Obtain detailed information about PCIe devices through multiple channels, specifically such as... Figure 5 As shown, information about the server's PCIe external devices is scanned and collected independently through both BIOS and BMC methods.
[0071] 3. BMC organizes basic PCIe device data: BMC organizes the collected basic PCIe device data for subsequent data omission verification.
[0072] 4. BMC processes information collected from multiple paths: BMC organizes the information collected from BIOS and BMC paths into a unified format and outputs it as a JSON file.
[0073] 5. Check for omissions: BMC compares the basic data with the multi-source data. If there are omissions, the information of the missing PCIe devices will be collected again; if there are no omissions, proceed to the next step.
[0074] 6. BMC presents data to users: BMC uses SMBIOS, REDIFSH, IPMI and other methods to collect PCIe device information and feeds it back to users.
[0075] Optionally, in this embodiment, the information of the high-speed serial computer extended bus standard device is obtained by using the basic input / output system enumeration method to obtain the first information, including: sending a read data command to the high-speed serial computer extended bus standard bus using the above-mentioned basic input / output system, reading the first field of the high-speed serial computer extended bus standard device according to the slot order, wherein the first field includes one or more of the device ID, device type, manufacturer information, product name, and serial number; and encapsulating the first field into JSON format information according to the device type-slot-field format to obtain the first information.
[0076] In this solution, the BIOS enumeration process ensures that all manufacturing information of PCIe devices is obtained at the initial stage of server startup, reducing the possibility of information omissions and enhancing the consistency and reliability of device information. By reading detailed device manufacturing information, maintenance personnel can more accurately identify and manage various types of PCIe devices in the server. Even with a large number and complex types of devices, device confusion can be effectively avoided, improving the accuracy and efficiency of device management. Encapsulating device manufacturing information in JSON format makes information management more convenient and improves readability and usability. For remote monitoring and maintenance systems, it enables faster parsing and display of device information, reducing the complexity of information processing.
[0077] After the server powers on and performs a self-test, the BIOS begins enumerating devices on the PCIe bus. The BIOS sends a configuration read command to the PCIe bus, reading the device manufacturing information in slot order. If a device is unresponsive, the BIOS automatically reduces the PCIe bus speed to retry, ensuring reliable information acquisition. After reading the manufacturing information of all devices, a first information set is formed. The BIOS sends a configuration read command to the PCIe bus again to locate the configuration space of each device; it then reads the device's manufacturing information fields according to the physical slot order; for cases of verification anomalies, the BIOS retryes reading to improve the success rate of information reading; it collects the first field information of all devices to prepare for subsequent information processing and transmission. After completing the reading of the device's first field information, the BIOS processes the information according to a preset format and converts the processed information to ISON format to ensure information structure and standardization. The first information in JSON format is sent to the BMC or information collaboration module via the LPC bus for subsequent information fusion and management; the information collaboration module receives and stores the first information to compare and supplement the second information obtained via the I2C / SMBus bus, ensuring the integrity and reliability of the final information set.
[0078] During server startup, the BIOS enumerates devices on the PCIe bus, sending configuration read commands to read the device's VPD information in slot order. This retrieves the device's manufacturing information, such as device ID, device type, manufacturer information, product name, and serial number. This process is automated and systematic, ensuring that the PCIe device's manufacturing information is accurately read early in the server's startup process. During enumeration, the BIOS sends read data commands to the PCIe bus, reading the first field of each device in slot order—the key information in the device's VPD, such as device ID, device type, manufacturer information, product name, and serial number. This information forms the device's basic identity and is crucial for device identification and management. The read first field information, i.e., the device's manufacturing information, is reorganized according to the "device type-slot-field" format and encapsulated into JSON format for easy subsequent processing and transmission. JSON format is lightweight, easy to read and write, and has a clear structure, making it suitable for transmitting and storing device information.
[0079] In high-density heterogeneous server cluster environments, operations and maintenance (O&M) personnel need to manage and monitor various PCIe devices on the servers in a unified manner. By implementing the above technical solution, O&M personnel can obtain the manufacturing information of all PCIe devices on the server in real time and accurately, including device ID, device type, manufacturer information, product name, and serial number. This information is stored and transmitted in JSON format, which facilitates parsing and display by the remote O&M platform, effectively reducing the information processing burden on O&M personnel and improving the efficiency of device management and troubleshooting. For example, on a certain server, when the BIOS is enumerating devices on the PCIe bus, it finds that the GPU device in slot 3 fails to be enumerated on the first attempt. At this time, the BIOS automatically reduces the PCIe bus speed to 16GB / s and retryes twice, eventually successfully reading the device's manufacturing information. This specific numerical application demonstrates the flexibility and reliability of the technical solution in dealing with device identification conflicts, ensuring the stability of information acquisition.
[0080] As an optional implementation, after sending a read data command to the high-speed serial computer expansion bus standard bus using the above-described basic input / output system and reading the first field of the high-speed serial computer expansion bus standard device according to the slot order, the method further includes: if the number of times the high-speed serial computer expansion bus standard device in the slot does not respond is greater than or equal to a first preset number, reducing the read rate of the high-speed serial computer expansion bus standard bus and resending the read data command; if the number of times the re-attempt to read data fails is greater than or equal to a second preset number, determining that the basic input / output system has failed to read data and generating a first fault log.
[0081] This solution effectively avoids read failures caused by poor signal quality or unstable device status during high-speed communication by reducing the PCIe bus speed and retrying the read when the device is unresponsive. This improves the stability and success rate of device information reading. By setting retry thresholds and a fault log generation mechanism, it ensures that maintenance personnel can promptly and accurately understand the fault situation when device information reading fails, enabling rapid response, reducing troubleshooting time, and avoiding negative impacts of device anomalies on server performance and business operations.
[0082] When the BIOS detects a device not responding, it records the number of times it has not responded. If the number of unresponsive attempts reaches or exceeds a first preset limit, the BIOS reduces the PCIe bus speed from 32GB / s to 16GB / s. At the lower bus speed, the BIOS resends the read data command to the device. After the device responds, the BIOS reads and stores the device information; if the device still does not respond, the above steps are repeated, up to a maximum of two retries. During the retries after reducing the bus speed, the BIOS continuously monitors the device's response. If the device still does not successfully return information within the second preset limit, the system marks it as a read failure. The BIOS records detailed information about the read failure, including the device slot location and the number of read failures. The read failure information is encapsulated in a first fault log for subsequent fault analysis and handling.
[0083] When the BIOS attempts to read PCIe device information in a slot, if the number of unresponsive attempts reaches or exceeds a first preset number, the BIOS will automatically reduce the PCIe bus read rate and then resend the read command. This strategy enhances the reliability of device information reading by reducing communication interference and improving the enumeration success rate when the device status is unstable or communication problems occur. After the BIOS attempts to reduce the PCIe bus rate and rereads the device information, if the read still fails after a second preset number of retries, the system will determine that the BIOS cannot read the device information and generate a first fault log. This mechanism ensures the traceability of device status anomalies, allowing maintenance personnel to quickly locate problems and take appropriate measures through the fault log.
[0084] In the operation and maintenance of large-scale server clusters, maintenance personnel face the challenge of managing high-density, heterogeneous PCIe devices. By implementing the above-mentioned technical solution, maintenance personnel can effectively handle device read conflicts or failures. Even in environments with fluctuating device status or abnormal PCIe bus communication, they can ensure the accurate and timely acquisition of device information, providing strong support for real-time monitoring of device status and troubleshooting. Suppose that in a server cluster, maintenance personnel encounter a situation where the GPU device in slot 3 remains unresponsive. After the initial failed attempt, the BIOS reduces the PCIe bus speed to 16GB / s and retryes reading data, performing two retries. However, the GPU device still does not respond. At this point, the BIOS determines that the data read has failed, generates the first fault log, and records information such as the device slot location, device type, and number of read failures, providing maintenance personnel with a basis for fault location. This specific numerical application demonstrates the practical effectiveness of this technical solution in handling abnormal device information reading situations, effectively improving maintenance efficiency and system stability.
[0085] Optionally, in this embodiment of the application, after encapsulating the first field into JSON format information according to the device type-slot-field format to obtain the first information, the method further includes: when the first information is obtained for the Nth time, recording the first duration of the first information obtained for the Nth time, where N≥1; when the first duration is greater than or equal to the first preset duration, sending the read data command to the high-speed serial computer extended bus standard bus again using the basic input / output system, so as to update the first information and obtain the updated first information, wherein the updated first information is used to merge with the second information to obtain the target information set.
[0086] In this solution, recording the duration of each information acquisition helps the system dynamically monitor the timeliness of information. For scenarios requiring real-time updates or monitoring of equipment status, it effectively avoids using outdated or altered information, ensuring the timeliness and accuracy of the information. By setting information update thresholds, the system can automatically detect and update equipment manufacturing information, ensuring that the combined result with functional information is the most up-to-date state, providing necessary data support for real-time monitoring of equipment status.
[0087] When the BIOS enumerates devices, it records the time of the first information acquisition. The system calculates the time difference between two information acquisitions. If the time difference exceeds a preset threshold (i.e., a first preset duration), the information is considered expired or needs to be updated. When the interval between two information acquisitions is detected to be greater than or equal to the first preset duration, the BIOS resends the read data command to the PCIe bus to update the device's manufacturing information. The BIOS collects the updated first information and prepares to merge it with the second information to form the target information set.
[0088] During server operation, the BIOS repeatedly enumerates PCIe devices to obtain their manufacturing information and encapsulates this information into a JSON-formatted first information set. To monitor the freshness of this information, the system needs to record the time point of each first information acquisition and calculate the duration between the Nth information acquisition and the next acquisition time point. If the time difference between two information acquisitions exceeds a preset duration, it indicates that the information may be outdated. In this case, the BIOS will send a read data command to the PCIe bus again to update the device's manufacturing information and obtain the updated first information set. The updated information will then be merged with the second information obtained through the BMC to form the target information set.
[0089] In dynamic server operation and maintenance environments, maintenance personnel need to monitor the status of PCIe devices in the server in real time. Considering that devices may be hot-swapped or their information status may change during operation, the above-mentioned technical solution can automatically detect the timeliness of the first piece of information and automatically update the device manufacturing information when the first duration exceeds a preset threshold, ensuring the timeliness and accuracy of the information. Assume the first preset duration is set to 30 minutes. During server operation, the BIOS enumerates PCIe devices and obtains manufacturing information every 10 minutes. When the system records the Nth time the first piece of information has been obtained, and no new information has been obtained within 30 minutes (i.e., the first duration has reached 30 minutes), the BIOS will send a read data command to the PCIe bus again to update the device's manufacturing information, ensuring the timeliness of the information. This specific application demonstrates the role of the technical solution in automatic information updating and real-time device status monitoring, effectively avoiding device management errors caused by outdated information and enhancing system stability and operational efficiency.
[0090] like Figure 6 As shown, the BIOS obtains detailed configuration information about a PCIe device by accessing its PCIe configuration space. Specifically, the VPD (Virtual Product Data) segment within the PCIe configuration space stores a wealth of valuable information about the device, which the BIOS can read by reading and writing the VPD address. More specifically:
[0091] 1) Enumeration: Send a "configuration read command" to the PCIe bus, allocate the configuration space base address in the order of slot 1→2→3→4 (e.g., slot 1: 0x001C0000, slot 2: 0x001D0000), and read the VPD core fields (Device ID, computing power / video memory).
[0092] 2) Conflict detection: If a device in a slot does not respond (e.g., GPU enumeration in slot 3 fails), the BIOS automatically reduces the PCIe bus speed (from 32GB / s to 16GB / s) and resends the enumeration command, retrying a maximum of 2 times;
[0093] 3) Repeat Steps 1 and 2 to read VPD data for all devices;
[0094] 4) Encapsulate the VPD data into JSON format (e.g., {"GPU":{"Slot1":{"Device ID":"0x2204","CUDA Cores":"6912"}}}) using the "Device Type-Slot-Field" format, and send it to the information collaboration module via the LPC bus, while updating the SMBIOSType41 table; Step 6: During server operation, receive the "VPD update request" from the BMC every 5 minutes, and reread the changed fields (e.g., SSD TBW, GPU firmware version) as needed to ensure data real-time performance.
[0095] As an optional implementation, the baseboard management controller obtains information about the aforementioned high-speed serial computer extended bus standard device via the I2C bus and SMBus bus to obtain the second information, including: sending a start address and a preset address to the I2C bus and the SMBus bus via the baseboard management controller, wherein the start address is the address where data reading begins and the preset address is the address where data reading stops; reading a second field from the start address to the preset address, wherein the second field includes one or more of the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version; and encapsulating the second field into JSON format information to obtain the second information.
[0096] In this solution, by sending commands to the I2C and SMBus buses via the BMC and specifying start and stop addresses, the scope of information reading can be precisely controlled, thereby improving the information acquisition rate and data accuracy. The reading of the second field covers detailed functional information of the device, helping maintenance personnel to fully understand the device status and improving the efficiency of fault diagnosis and preventative maintenance. By setting the information update cycle, the BMC can automatically detect changes in device status and update the second information in a timely manner, thus improving the real-time performance and efficiency of information management.
[0097] After the server powers on and performs a self-test, the BMC initializes the I2C and SMBus bus controllers. For each device, the BMC sends a read command to the start address of the device's FRU EEPROM via the corresponding bus. It continuously reads device information up to a preset stop address, ensuring that the read information includes all critical fields. The read device information undergoes CRC16 or PEC verification to ensure the integrity of data transmission. The BMC then sends another read command to the start address of the FRU EEPROM. It reads the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version to the preset address. All second-field information is read and encapsulated into JSON according to the format "Device Type-Slot-Field". During server operation, the BMC periodically checks the device status and updates the second-field information. The updated second-field information is sent to the information collaboration module or remote management platform via a protocol (such as IPMI). Maintenance personnel or the management system can monitor device status through real-time updates, achieving efficient management.
[0098] After the server powers on, the BMC initializes the I2C and SMBus bus controllers. It reads the FRU information of the devices by sending read commands to preset start and stop addresses. This information includes the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version. This process ensures the scope and completeness of the information obtained. When reading the FRU information via the I2C and SMBus buses, the BMC reads from the start address to the preset address, covering all second fields, including the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version. This information constitutes the device's "identity card," crucial for device management and maintenance. The read second field information is encapsulated in JSON format for easy data storage, transmission, and parsing. During server operation, the BMC automatically performs the reading and updating of the second field information at preset time intervals, such as every 30 seconds, maintaining the timeliness and accuracy of the information.
[0099] In the operation and maintenance of high-density server clusters, maintenance personnel need to monitor the functional information of all PCIe devices in the servers in real time. By implementing the above technical solution, BMC can automatically and in real time read and update the FRU information of the devices to the remote management platform, including the device serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version. This information is crucial for accurate monitoring of device status. For example, assuming the second preset time interval is set to 30 seconds, BMC updates the device's functional information every 30 seconds. Maintenance personnel or the remote management system can grasp the device's operating status and health information in real time, which has significant advantages for rapid response to device anomalies and preventive maintenance.
[0100] Optionally, in this embodiment of the application, after reading the second field from the above-mentioned starting address to the above-mentioned preset address, the method further includes: performing CRC verification on the second field to obtain a first verification result; performing PEC verification on the second field to obtain a second verification result; if the first verification result indicates that the verification failed and the second verification result indicates that the verification failed at least once, rereading the second field; if the number of times the verification failed is greater than or equal to a third preset number, determining that the baseboard management controller failed to read data, and generating a second fault log.
[0101] In this solution, CRC checksums detect errors during data transmission, ensuring that the device information read by the BMC has not been tampered with or damaged, thus enhancing data reliability and security. PEC checksums, as an additional verification method, combine with CRC checksums to form a dual verification mechanism, significantly improving the accuracy and reliability of data transmission. Instant rereading can quickly correct data errors detected by CRC or PEC checksums, preventing inaccurate device information management due to data errors. By setting a threshold for the number of failed checks, the system can automatically identify and record abnormal data reading situations, improving the accuracy of fault detection and the transparency of device management.
[0102] The BMC reads the second field information into the memory buffer. A CRC checksum algorithm is applied to calculate the checksum of the second field. The calculated checksum is compared with the pre-stored correct checksum. After reading the second field data packet transmitted on the SMBus bus, a PEC checksum is calculated. The calculated PEC checksum is compared with the pre-stored checksum. It is determined whether the first or second checksum result indicates a checksum failure. If at least one checksum result fails, a re-reading mechanism is triggered. The second field information is re-read into the BMC memory buffer. The number of data reads that fail the CRC or PEC checksum is counted. When the number of read failures reaches or exceeds a third preset number, the system determines a read anomaly. A second fault log is generated, recording the anomaly details.
[0103] After the BMC reads the FRU information of the PCIe device, it performs a CRC (Cyclic Redundancy Check) algorithm check on the read data to obtain the first check result, which is used to verify the integrity of the data transmission process. PEC (Packet Error Checking) is a verification mechanism on the SMBus bus used to further verify the integrity of the data packets and errors in the transmission process, obtaining a second check result as another layer of protection for data reliability verification. If either the CRC check or the PEC check shows that the data fails the check, that is, the first check result or the second check result indicates a check failure, the BMC will immediately reread the second field information to correct any possible data errors. If, after multiple attempts, the read data still fails the CRC or PEC check, that is, the number of failed checks reaches a third preset number, the system will determine that the BMC has failed to read the data, generate a second fault log, and record specific fault information, including the device location, the type of failed check, etc., to facilitate subsequent fault investigation and repair.
[0104] In large-scale server operations and maintenance, factors such as signal interference and hardware failures can affect the accuracy of data retrieval. By implementing the above technical solution, BMC automatically detects data transmission errors after reading device information through a dual CRC and PEC verification mechanism, and corrects data errors immediately, improving the stability and reliability of device information management. For example, if a third preset number of attempts is set to three, and the data fails the CRC or PEC verification after three consecutive attempts by BMC to read device information, the system will generate a second fault log. Maintenance personnel can quickly locate the problematic device based on the log and take corresponding troubleshooting measures to ensure the normal operation of the server equipment.
[0105] As an optional implementation, after encapsulating the second field into JSON format information to obtain the second information, the method further includes: when the second information is obtained for the Mth time, recording the second duration of the second information obtained for the Mth time, where M≥1; when the second duration is greater than or equal to the second preset duration, reading the second field from the starting address to the preset address again to update the second information and obtain updated second information, wherein the updated second information is used to merge with the first information to obtain the target information set.
[0106] In this solution, by recording the duration of each information acquisition, the system can automatically determine when information needs to be updated, avoiding the problems of outdated information or excessively frequent updates. This ensures both the timeliness of information and optimizes resource utilization. By setting a time limit for information updates, the system ensures that the target information set formed by merging the second and first pieces of information always remains up-to-date, which is crucial for real-time monitoring of equipment status and rapid fault response.
[0107] Obtain the current time as the end time of M reads. Calculate the duration since the last successful read of the second information. Record the duration of this read. Check if the current duration has reached or exceeded the second preset duration. If the condition is met, BMC re-initiates the process of reading the second field information. The updated second information is encapsulated and prepared to be merged with the first information. The updated second information is merged with the first information to form the target information set.
[0108] During the process of acquiring the FRU information (second information) of the device, the BMC tracks the timestamp of each information acquisition and calculates the duration since the last successful information acquisition (second duration) to ensure the continuity and regularity of information updates. When the BMC detects that the duration since the last successful reading of the second information exceeds the second preset duration, the system will automatically trigger the information update process, read the device's FRU information again via the I2C and SMBus buses, and update the second information to ensure the real-time performance and accuracy of the information, so as to merge it with the first information acquired by the BIOS to form a complete target information set.
[0109] In data center server operation and maintenance management, maintenance personnel need to regularly monitor the status information of all PCIe devices in the server, including the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version. Since the server environment can be affected by various factors, such as system load and hot-plug events, the timeliness of device status information directly impacts the accuracy of maintenance decisions. To ensure the real-time nature and accuracy of the information, the system implements the above-mentioned technical solution to automatically detect the duration of the second information field. When the duration exceeds a preset 10 minutes, the system automatically triggers a re-reading and update of the second information field, ensuring that the target information set formed by merging the first information with the latest status information. This provides maintenance personnel with real-time and accurate device status monitoring data.
[0110] like Figure 7As shown, for the BMC, after the server powers on, the BMC initializes the I2C / SMBus controller, identifies the device protocol type through "bus probing" (sending an I2C probe command to 0xA0 and an SMBus probe command to 0x50), and reads the information stored in the PCIe device's FRU EEPROM (the device's "identity card" pre-programmed by the manufacturer). Details are as follows:
[0111] 1) Send I2C start bit → read command from address 0xA0 + start offset address 0x00; receive serial number, manufacturer and other fields (64 bytes in total), and perform CRC16 check (compare the check value of FRU Offset 0x3F-0x40);
[0112] 2) Switch to SMBus mode, send start bit → read command from address 0x50 → start offset address 0x00; receive fields such as warranty period and hardware version (32 bytes in total), and perform PEC (Packet Error Checking) verification;
[0113] 3) Verification passed: The FRU data is encapsulated into JSON and sent to the information collaboration module; Verification failed: Retry 3 times. If it still fails, record the fault log ("SMBus failed to read RAIDFRU, reason: PEC verification error") and trigger the backup logic (extract the hardware version field from VPD).
[0114] 4) Real-time refresh: Steps 1-4 are executed automatically every 30 seconds to ensure that the FRU information is synchronized with the device status (such as the "firmware version" field in the FRU is updated after the RAID card firmware is upgraded).
[0115] Optionally, in this embodiment of the application, after obtaining the information of the high-speed serial computer extended bus standard device through the I2C bus and SMBus bus using the baseboard management controller to obtain the second information, the method further includes: comparing whether the first number of devices and the second number of devices are the same, wherein the first number of devices is the number of devices identified by the basic input / output system and the second number of devices is the number of devices identified by the baseboard management controller; if the first number of devices and the second number of devices are different, rereading the first information and the second information.
[0116] This solution obtains and compares device quantity information through both BIOS and BMC paths, enabling timely detection of discrepancies and preventing misjudgments or management errors due to inconsistencies. When inconsistencies arise, re-reading the information helps eliminate errors in the device identification process, improving the accuracy and reliability of server device information management.
[0117] The BIOS completes PCIe device enumeration and records the first number of devices. The BMC completes device information reading on the I2C / SMBus bus and records the second number of devices. The system compares the first and second device counts for equality. It then checks if the first and second device counts match. If the device counts do not match, the BIOS and BMC re-execute the device information reading process. The re-read information is compared again until they match.
[0118] During server initialization and operation, the BIOS and BMC independently acquire information about the PCIe devices in the server. The BIOS obtains the number of devices through enumeration (first device count), while the BMC reads device information via the I2C and SMBus buses (second device count). After acquiring the information, the system compares the device counts obtained through the two methods to confirm the completeness and accuracy of the information. If the number of PCIe devices identified by the BIOS differs from the number identified by the BMC, the system will trigger a re-reading process, where the BIOS and BMC will again acquire the device information separately to ensure data consistency and integrity.
[0119] In server hot-swappable maintenance, administrators frequently need to hot-swap PCIe devices while the server is running, such as replacing a faulty GPU or upgrading an SSD. Since hot-swapping can affect device identification and information acquisition, inconsistencies may arise between the device information obtained by the BIOS and BMC. By implementing the above technical solution, after the initial reading of device information, the system automatically compares the number of devices identified by the BIOS and BMC. If a discrepancy is found (e.g., the BIOS identifies 16 devices while the BMC only identifies 15), the system will trigger a re-reading mechanism. The BIOS and BMC will then re-obtain the device information separately until the two sets of device counts match.
[0120] As an optional implementation, after merging the first information and the second information to obtain the target information set, the method further includes: when there are multiple high-speed serial computer expansion bus standard devices, classifying all the target information sets of the multiple high-speed serial computer expansion bus standard devices to obtain a classified target information set, wherein the classification method includes one or more of the following: classification by GPU, classification by SSD card, and classification by network card; and storing the classified target information sets into different storage devices.
[0121] This solution categorizes and stores information about different types of devices, enabling maintenance personnel to more efficiently locate and manage specific types of devices, such as GPUs, SSDs, or network devices, thus improving the efficiency of information retrieval and device management. Storing the categorized information across different devices not only achieves classified information management but also enhances information security, preventing information loss due to the failure of a single storage point. Furthermore, maintenance personnel can access information about specific types of devices more quickly, improving overall maintenance efficiency.
[0122] Device information is obtained through BIOS and BMC to form a target information set. The target information set is then categorized according to device type (GPU, SSD, network card, etc.). The categorized information set is then prepared for storage. The storage devices for each type of device information are determined. The categorized target information set is stored on the corresponding storage devices according to device type. Information on the storage devices is stored and managed according to a preset format for easy subsequent information retrieval and device monitoring.
[0123] When a server contains multiple PCIe devices (such as GPUs, SSDs, and network cards), the system categorizes the target information set formed by the first and second pieces of information obtained and merged through the BIOS and BMC. This categorization can be based on device type, such as GPU, SSD, or network card. The categorized target information set is then stored in the corresponding storage device according to the device type. For example, GPU information is stored in a GPU information management storage device, SSD information in an SSD information management storage device, and network card information in a network device information management storage device.
[0124] In the operation and maintenance management of large-scale heterogeneous server clusters, data centers typically deploy a large number of different types of PCIe devices, such as graphics processing units (GPUs), storage devices (SSD cards), control devices (RAID controllers), and network devices (high-speed Ethernet / InfiniBand network cards). Operations personnel need to efficiently manage and monitor the status information of these devices. By implementing the above technical solution, after merging the first and second sets of information to form the target information set, the system automatically categorizes the information set and stores it according to device type. For example, assuming the server contains 4 GPUs, 2 SSD cards, and 1 network device, the system stores the GPU information set in dedicated GPU storage, the SSD card information set in dedicated SSD card storage, and the network device information set in dedicated network device storage. This strategy not only enables operations personnel to quickly locate and access detailed information for specific types of devices but also enhances information storage security and management targeting, improving operational efficiency and the quality of server cluster management.
[0125] The beneficial effects of this plan are:
[0126] 1. Improved reliability of multi-device management: Batch enumeration + dual-bus reading + conflict retry design increases the success rate of PCIe device information reading from 85% of the existing technology to 99.5%. A test on an AI server showed that 100 servers with 16 GPUs each ran continuously for 30 days. The existing technology had 45 information reading failures, while this solution only had 5 failures (all of which were recovered through retries). Moreover, the success rate of hot-swappable device information updates was 100%.
[0127] 2. Improved Operation and Maintenance Efficiency: The unified information interface (IPMI + RESTful API) supports remote batch management without the need for tool switching or manual comparison. Operation and maintenance personnel can view the information of all PCIe devices (VPD + FRU) on 10,000 servers through a single web management platform. The query time for information on a single server is reduced from 15 minutes to 2 minutes, and the operation and maintenance time for 10,000 servers is reduced from 2,500 hours to 625 hours.
[0128] 3. Information integrity and consistency guarantee across all devices: (GPU / SSD / RAID / Network Card) + automatic verification mechanism, increasing information coverage from 40% of existing technologies to 100%, and 100% accuracy of inconsistency alarms. In a test, the FRU version information of the RAID card of 5 servers was deliberately modified. This solution triggered alarms within 30 seconds, allowing maintenance personnel to quickly locate and correct the problem, avoiding configuration errors caused by incorrect information.
[0129] 4. Strong scenario compatibility: Supports PCIe 3.0 / 4.0 / 5.0 protocols and I2C / SMBus buses, and is compatible with more than 98% of mainstream PCIe devices such as NVIDIA / AMD GPUs, Samsung / Intel SSDs, and LSI / Broadcom RAID cards. No hardware or software modifications are required for different manufacturers. For example, after replacing the NVIDIA A100 with an AMD MI250 GPU (PCIe 4.0), this solution can read VPD data and synchronize it to the collaborative module normally without reconfiguration.
[0130] The main advantages of this solution are:
[0131] 1. Dual-subject multi-bus information reading architecture: The BIOS enumerates all types of PCIe devices in batches and reads VPD data through the PCIe bus, while the BMC reads the FRU information of some devices through the I2C / SMBus bus, achieving "full device - full dimension" information coverage;
[0132] 2. Multi-device information collaborative verification mechanism: Automatic comparison of VPD and FRU data is achieved through cross-correlation fields (hardware version, compatibility list). When inconsistencies are found, multi-level alarms are triggered to ensure information accuracy.
[0133] 3. Unified information management interface and extended design: compatible with IPMI / SMBus / Redfish API, supports PCIeSwitch cascading expansion (managing up to 32 PCIe devices) and AI health prediction, and is suitable for ultra-large-scale AI / storage server scenarios.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0135] Embodiments of this application also provide a device for managing device information. Figure 8 This is a structural block diagram of a device for managing device information according to an embodiment of this application, such as... Figure 8 As shown, the device includes:
[0136] The first acquisition unit 802 is used to acquire information about the high-speed serial computer extended bus standard device by using the basic input / output system enumeration method, and obtain first information, wherein the first information is the manufacturing information of the product of the high-speed serial computer extended bus standard device.
[0137] The second acquisition unit 804 is used to acquire information of the high-speed serial computer expansion bus standard device via the I2C bus and SMBus bus using the baseboard management controller, and obtain second information, wherein the second information is the functional information of the high-speed serial computer expansion bus standard device.
[0138] The merging module unit 806 is used to merge the first information and the second information to obtain the target information set;
[0139] The sending unit 808 is used to send the target information set to the target device used by the target object, so that the target object can monitor the status of the device based on the target information set.
[0140] Through the above steps, the BIOS enumerates PCIe devices and reads VPD information, including device manufacturing information such as device ID, manufacturer, and product name. VPD information is an inherent attribute of the device, extracted by the BIOS at the beginning, ensuring comprehensive coverage of basic device information. The BMC uses the I2C bus and SMBus bus to read the device's FRU information, which covers the device's functional information such as serial number, firmware version, and hardware status. By reading data through the BMC, the deficiencies of VPD information in functional description and status monitoring are compensated for. The VPD information obtained by the BIOS and the FRU information obtained by the BMC are merged to form a target information set. The target information set contains the device's manufacturing and functional information, achieving comprehensive information coverage of PCIe devices. Therefore, it can solve the technical problem of incomplete information coverage of current PCIe devices in related technologies.
[0141] Optionally, in this embodiment, the first acquisition unit includes a first reading module and a first encapsulation module. The first reading module is used to send a read data command to the high-speed serial computer extended bus standard bus using the above-described basic input / output system, and read the first field of the above-described high-speed serial computer extended bus standard device according to the slot order. The first field includes one or more of the following: device ID, device type, manufacturer information, product name, and serial number. The first encapsulation module is used to encapsulate the first field into JSON format information according to the device type-slot-field format to obtain the first information.
[0142] In this solution, the BIOS enumeration process ensures that all manufacturing information of PCIe devices is obtained at the initial stage of server startup, reducing the possibility of information omissions and enhancing the consistency and reliability of device information. By reading detailed device manufacturing information, maintenance personnel can more accurately identify and manage various types of PCIe devices in the server. Even with a large number and complex types of devices, device confusion can be effectively avoided, improving the accuracy and efficiency of device management. Encapsulating device manufacturing information in JSON format makes information management more convenient and improves readability and usability. For remote monitoring and maintenance systems, it enables faster parsing and display of device information, reducing the complexity of information processing.
[0143] As an optional implementation, the above-described apparatus further includes a reading unit and a first generation unit. The reading unit is configured to, after sending a read data command to the high-speed serial computer expansion bus standard bus using the basic input / output system and reading the first field of the high-speed serial computer expansion bus standard device according to the slot order, reduce the reading rate of the high-speed serial computer expansion bus standard bus and resend the read data command if the number of times the high-speed serial computer expansion bus standard device in the slot fails to respond is greater than or equal to a first preset number. The first generation unit is configured to, if the number of times the re-attempt to read data fails is greater than or equal to a second preset number, determine that the basic input / output system has failed to read data and generate a first fault log.
[0144] This solution effectively avoids read failures caused by poor signal quality or unstable device status during high-speed communication by reducing the PCIe bus speed and retrying the read when the device is unresponsive. This improves the stability and success rate of device information reading. By setting retry thresholds and a fault log generation mechanism, it ensures that maintenance personnel can promptly and accurately understand the fault situation when device information reading fails, enabling rapid response, reducing troubleshooting time, and avoiding negative impacts of device anomalies on server performance and business operations.
[0145] Optionally, in this embodiment of the application, the above-mentioned device further includes a recording unit and a first updating unit. The recording unit is used to record the first duration of the first information obtained for the Nth time after encapsulating the first field into JSON format information according to the device type-slot-field format to obtain the first information, where N≥1. The first updating unit is used to send the read data command to the high-speed serial computer extended bus standard bus again using the basic input / output system when the first duration is greater than or equal to the first preset duration, so as to update the first information and obtain the updated first information, wherein the updated first information is used to merge with the second information to obtain the target information set.
[0146] In this solution, recording the duration of each information acquisition helps the system dynamically monitor the timeliness of information. For scenarios requiring real-time updates or monitoring of equipment status, it effectively avoids using outdated or altered information, ensuring the timeliness and accuracy of the information. By setting information update thresholds, the system can automatically detect and update equipment manufacturing information, ensuring that the combined result with functional information is the most up-to-date state, providing necessary data support for real-time monitoring of equipment status.
[0147] As an optional implementation, the second acquisition unit includes an initiation module, a second reading module, and a second encapsulation module. The initiation module is used to send a start address and a preset address to the I2C bus and the SMBus bus using the aforementioned baseboard management controller, wherein the start address is the address where data reading begins, and the preset address is the address where data reading stops. The second reading module is used to read a second field from the start address to the preset address, wherein the second field includes one or more of the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version. The second encapsulation module is used to encapsulate the second field into JSON format information to obtain the second information.
[0148] In this solution, by sending commands to the I2C and SMBus buses via the BMC and specifying start and stop addresses, the scope of information reading can be precisely controlled, thereby improving the information acquisition rate and data accuracy. The reading of the second field covers detailed functional information of the device, helping maintenance personnel to fully understand the device status and improving the efficiency of fault diagnosis and preventative maintenance. By setting the information update cycle, the BMC can automatically detect changes in device status and update the second information in a timely manner, thus improving the real-time performance and efficiency of information management.
[0149] Optionally, in this embodiment, the device further includes a first verification unit, a second verification unit, a first rereading unit, and a second generation unit. The first verification unit is used to perform CRC verification on the second field after reading the second field from the starting address to the preset address to obtain a first verification result. The second verification unit is used to perform PEC verification on the second field to obtain a second verification result. The rereading unit is used to reread the second field if the first verification result indicates that the verification failed and the second verification result indicates that the verification failed at least once. The second generation unit is used to determine that the substrate management controller has failed to read data and generate a second fault log if the number of failed verifications is greater than or equal to a third preset number.
[0150] In this solution, CRC checksums detect errors during data transmission, ensuring that the device information read by the BMC has not been tampered with or damaged, thus enhancing data reliability and security. PEC checksums, as an additional verification method, combine with CRC checksums to form a dual verification mechanism, significantly improving the accuracy and reliability of data transmission. Instant rereading can quickly correct data errors detected by CRC or PEC checksums, preventing inaccurate device information management due to data errors. By setting a threshold for the number of failed checks, the system can automatically identify and record abnormal data reading situations, improving the accuracy of fault detection and the transparency of device management.
[0151] As an optional implementation, the above-described apparatus further includes a recording unit and a second updating unit. The recording unit is used to record the second duration of the second information obtained for the Mth time after encapsulating the second field into JSON format information and obtaining the second information, where M≥1. The second updating unit is used to read the second field from the starting address to the preset address again when the second duration is greater than or equal to the second preset duration, so as to update the second information and obtain the updated second information, wherein the updated second information is used to merge with the first information to obtain the target information set.
[0152] In this solution, by recording the duration of each information acquisition, the system can automatically determine when information needs to be updated, avoiding the problems of outdated information or excessively frequent updates. This ensures both the timeliness of information and optimizes resource utilization. By setting a time limit for information updates, the system ensures that the target information set formed by merging the second and first pieces of information always remains up-to-date, which is crucial for real-time monitoring of equipment status and rapid fault response.
[0153] Optionally, in this embodiment, the above-mentioned device further includes a comparison unit and a second rereading unit. The comparison unit is used to compare whether the first number of devices and the second number of devices are the same after obtaining the information of the high-speed serial computer extended bus standard device through the I2C bus and SMBus bus using the baseboard management controller and obtaining the second information. The first number of devices is the number of devices identified by the basic input / output system, and the second number of devices is the number of devices identified by the baseboard management controller. The second rereading unit is used to reread the first information and the second information if the first number of devices and the second number of devices are different.
[0154] This solution acquires and compares device quantity information through both BIOS and BMC paths, enabling timely detection of discrepancies and preventing misjudgments or management errors due to inconsistencies. When inconsistencies arise, re-reading the information helps eliminate errors in device identification, improving the accuracy and reliability of server device information management.
[0155] As an optional implementation, the above-described apparatus further includes a classification unit and a storage unit. The classification unit is used to classify all the target information sets of the multiple high-speed serial computer expansion bus standard devices after merging the first information and the second information to obtain a target information set, in the case of multiple high-speed serial computer expansion bus standard devices, to obtain a classified target information set. The classification method includes one or more of the following: classification by GPU, classification by SSD card, and classification by network card. The storage unit is used to store the classified target information sets into different storage devices.
[0156] This solution categorizes and stores information about different types of devices, enabling maintenance personnel to more efficiently locate and manage specific types of devices, such as GPUs, SSDs, or network devices, thus improving the efficiency of information retrieval and device management. Storing the categorized information across different devices not only achieves classified information management but also enhances information security, preventing information loss due to the failure of a single storage point. Furthermore, maintenance personnel can access information about specific types of devices more quickly, improving overall maintenance efficiency.
[0157] For a description of the features in the embodiment corresponding to the device for managing device information, please refer to the relevant description in the embodiment corresponding to the method for managing device information, which will not be repeated here.
[0158] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the device information management method.
[0159] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the device information management method.
[0160] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0161] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the device information management method.
[0162] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the device information management method.
[0163] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0164] The above provides a detailed description of a device information management method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for managing equipment information, characterized in that, include: The information of the high-speed serial computer extended bus standard device is obtained by using the basic input / output system enumeration method to obtain the first information, wherein the first information is the manufacturing information of the product of the high-speed serial computer extended bus standard device. The baseboard management controller obtains information about the high-speed serial computer expansion bus standard device through the I2C bus and SMBus bus to obtain second information, which is the functional information of the high-speed serial computer expansion bus standard device. The first information and the second information are combined to obtain the target information set; The target information set is sent to the target device used by the target object, so that the target object can monitor the status of the device based on the target information set.
2. The method according to claim 1, characterized in that, Information about standard devices on the high-speed serial computer extended bus is obtained by enumerating basic input / output systems, yielding the first piece of information, including: The basic input / output system is used to send a read data command to the high-speed serial computer expansion bus standard bus, and the first field of the high-speed serial computer expansion bus standard device is read according to the slot order. The first field includes one or more of the following: device ID, device type, manufacturer information, product name, and serial number. The first field is encapsulated into JSON format information according to the device type-slot-field format to obtain the first information.
3. The method according to claim 2, characterized in that, After sending a read data command to the high-speed serial computer expansion bus standard bus using the basic input / output system, and reading the first field of the high-speed serial computer expansion bus standard device according to the slot order, the method further includes: If the number of times the high-speed serial computer expansion bus standard device in the slot fails to respond is greater than or equal to a first preset number, the read rate of the high-speed serial computer expansion bus standard bus is reduced, and the read data command is resent. If the number of unsuccessful attempts to read data is greater than or equal to the second preset number, it is determined that the basic input / output system has failed to read data, and a first fault log is generated.
4. The method according to claim 2, characterized in that, After encapsulating the first field into JSON format information according to the device type-slot-field format to obtain the first information, the method further includes: In the case of obtaining the first information for the Nth time, record the first duration of the first information obtained for the Nth time, where N≥1; If the first duration is greater than or equal to the first preset duration, the basic input / output system sends the read data command to the high-speed serial computer extended bus standard bus again to update the first information and obtain the updated first information, wherein the updated first information is used to merge with the second information to obtain the target information set.
5. The method according to claim 1, characterized in that, The baseboard management controller obtains information about the high-speed serial computer expansion bus standard device via the I2C bus and SMBus bus to obtain second information, including: The baseboard management controller sends a start address and a preset address to the I2C bus and the SMBus bus, wherein the start address is the address at which data reading begins and the preset address is the address at which data reading stops. Read the second field from the starting address to the preset address, wherein the second field includes one or more of the device's serial number, manufacturer, model, production date, warranty period, hardware version, and firmware version; The second field is encapsulated into JSON format information to obtain the second information.
6. The method according to claim 5, characterized in that, After reading the second field from the starting address to the preset address, the method further includes: Perform a CRC check on the second field to obtain the first check result; Perform PEC validation on the second field to obtain the second validation result; If at least one of the following conditions is met: the first verification result indicates that the verification failed and the second verification result indicates that the verification failed, the second field will be read again. If the number of failed verifications is greater than or equal to a third preset number, the baseboard management controller is determined to have failed to read data, and a second fault log is generated.
7. The method according to claim 5, characterized in that, After encapsulating the second field into JSON format information to obtain the second information, the method further includes: In the case of obtaining the second information for the Mth time, record the second duration of the second information obtained for the Mth time, where M≥1; If the second duration is greater than or equal to the second preset duration, the second field from the starting address to the preset address is read again to update the second information and obtain updated second information, wherein the updated second information is used to merge with the first information to obtain the target information set.
8. The method according to any one of claims 1 to 7, characterized in that, After obtaining the second information by using a baseboard management controller to acquire information about the high-speed serial computer extended bus standard device via the I2C bus and SMBus bus, the method further includes: Compare whether the first number of devices and the second number of devices are the same, wherein the first number of devices is the number of devices identified by the basic input / output system, and the second number of devices is the number of devices identified by the baseboard management controller; If the number of the first device and the number of the second device are different, the first information and the second information are read again.
9. The method according to any one of claims 1 to 7, characterized in that, After merging the first information and the second information to obtain the target information set, the method further includes: When there are multiple high-speed serial computer expansion bus standard devices, the target information sets of all the multiple high-speed serial computer expansion bus standard devices are classified to obtain classified target information sets. The classification method includes one or more of the following: classification by GPU, classification by SSD card, and classification by network card. The classified target information sets are stored in different storage devices.
10. A device for managing equipment information, characterized in that, include: The first acquisition unit is used to acquire information about the high-speed serial computer extended bus standard device by using the basic input / output system enumeration method, and obtain first information, wherein the first information is the manufacturing information of the product of the high-speed serial computer extended bus standard device. The second acquisition unit is used to acquire information of the high-speed serial computer expansion bus standard device through the I2C bus and SMBus bus using the baseboard management controller, and obtain second information, the second information being the functional information of the high-speed serial computer expansion bus standard device. The merging module unit is used to merge the first information and the second information to obtain the target information set; The sending unit is used to send the target information set to the target device used by the target object, so that the target object can monitor the status of the device based on the target information set.
Citation Information
Cited By
Method and device for managing JBOG by BMC
CN122111801A