A method for real-time and accurate monitoring of the location status and quantity of key server components
Patent Information
- Application Number
- CN202611181214.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-01
AI Technical Summary
然而,随着服务器架构日益复杂、部件密度不断提高以及业务对可用性要求的持续提升,现有监控方案暴露出以下显著不足:
1、 本发明中,通过构建预置的硬件接口与物理位置映射表,以键值对形式精确记录GPIO引脚号、I2C总线地址与物理位置字符串的对应关系,将传统方案仅能报告的有或无二元状态突破至插槽级精确定位,能够准确输出CPU插槽编号、DIMM插槽编号、风扇槽位编号或电源槽位编号等层级化物理位置信息;同时,通过为每个被监控部件建立独立的状态机实例,定义在位且正常、在位但异常、不在位、故障等多维状态集合,结合CPLD并行采集PRSNT#引脚电平信号与BMC主动扫描I2C总线设备地址的双模式检测机制,有效区分物理在位、电气连接和功能运行三个维度的状态差异,彻底解决了传统方案无法精确报告具体哪个插槽空闲或故障、无法区分物理缺失与电气故障的问题,使运维人员无需人工对照丝印进行排查,直接通过图形化界面定位故障部件,显著提升了运维效率和故障诊断的准确性;
Smart Images

Figure CN122673052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server hardware management, and in particular to a method for real-time and accurate monitoring of the location status and quantity of key server components. Background Technology
[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, the scale of servers in data centers is growing exponentially. The stability and maintainability of server systems have become core concerns for data center operation and maintenance management. In the complex architecture of servers, critical components such as CPUs, memory modules, hard drives, fans, power supply modules, and PCIe expansion cards are fundamentally determined by their correct location and accurate quantity configuration, directly impacting the server's computing performance and data storage reliability. Therefore, real-time and precise monitoring of the location and quantity of critical server components is a prerequisite for achieving automated server operation and maintenance, rapid fault location, and refined asset management.
[0003] In existing technologies, server management controllers typically report the presence and status of components through standard protocols such as intelligent platform management interfaces, providing basic hardware monitoring capabilities for operations and maintenance personnel. However, with increasingly complex server architectures, ever-increasing component density, and continuously rising availability requirements from businesses, existing monitoring solutions have revealed the following significant shortcomings: 1. The information granularity is coarse, usually only able to report whether it is present or absent, and unable to accurately report which specific slot is idle or faulty; 2. Poor real-time performance; status updates rely on periodic polling, which introduces delays and prevents immediate response to hot-plug events. 3. Reliability is questionable; some implementations rely on operating system drivers or higher-level protocols, and may fail to function when the system malfunctions (such as freezing or OS becoming unresponsive). 4. It lacks intelligent association and cannot automatically associate the status of components with specific physical location information. Manual inspection is required by comparing with the silkscreen, resulting in low maintenance efficiency. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method and system for real-time and accurate monitoring of the location status and quantity of key server components, thus solving the above problems.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for real-time and accurate monitoring of the location status and quantity of key server components, comprising the following steps: S1. Directly acquire physical level signals or bus communication signals of key components through hardware interfaces; S2. Call the hardware interface and physical location mapping table preset in the BMC memory. The mapping table records the correspondence between GPIO pin number, I2C bus address and physical location string in key-value pair form. Use the signal identifier collected in step S1 as the input key value to query the mapping table and output the corresponding physical location string. The physical location string includes CPU socket number, DIMM socket number, fan slot number or power supply slot number. S3. Establish an independent state machine instance for each monitored component, and define the state set as: in place and normal, in place but abnormal, out of place, and faulty. S4. Configure the GPIO of CPLD and BMC to work in interrupt-triggered mode. When a rising or falling edge change is detected in the PRSNT# pin level, a hardware interrupt signal is immediately generated. In the interrupt service routine, the interrupt event identifier is written to the lockless circular queue. The interrupt event identifier includes the GPIO pin number that changed, the direction of the level change, and the timestamp. S5. Write the component status data obtained in step S3 into a lightweight database deployed in the BMC flash memory. The status data includes component type, physical location, current status, component model, serial number, and status change timestamp. Output structured data to the external management system through a Redfish-compliant RESTful API. The structured data is encapsulated in JSON format and includes a Component field indicating the component type, a Location field indicating the physical location string, a Present field indicating the on-site status, and a PartNumber field indicating the component number.
[0006] Preferably, in step S1, the hardware interface is divided into the general purpose input / output (GPIO) pins of the programmable logic device (CPLD), the I2C bus interface of the baseboard management controller (BMC), the GPIO pins of the CPLD or BMC, and the SMBus bus of the BMC. Specifically, the GPIO pins of the CPLD are used to acquire the level signal of the CPU socket PRSNT# pin. When the CPU is inserted, this pin is pulled low to ground, and when removed, it is pulled high by a pull-up resistor. The I2C bus interface of the BMC is used to actively scan the serial presence detection (SPD) device addresses on each memory channel. If an EEPROM device response is detected, the memory module is determined to be present. The GPIO pins of the CPLD or BMC are used to acquire the level signal of the fan connector PRSNT# pin to determine the fan's presence status, and simultaneously read the fan's TACH speed signal. The SMBus bus of the BMC is used to establish PMBus communication with the power module and read the PRESENT and POWER_GOOD status bits of the power module.
[0007] Preferably, the CPLD's parallel acquisition of the level signal of the CPU socket PRSNT# pin includes: The CPLD detects pin level changes with a microsecond-level response time and operates independently of the BMC and host CPU, continuously monitoring component insertion and removal events even when the server crashes or the operating system becomes unresponsive. The CPLD integrates a hardware debouncing filter module to debouncing the PRSNT# pin signal, eliminating level oscillations caused by mechanical contact jitter.
[0008] Preferably, the BMC's I2C bus interface actively scans the SPD device addresses on each memory channel, including: Iterate through the pre-configured list of I2C bus numbers and the list of slave device addresses; For each bus address combination, execute an I2C probe command. If an acknowledgment signal is received, record that the DIMM slot corresponding to that address is in place, and further read the memory parameter information in the SPD EEPROM.
[0009] Preferably, the hardware interface-physical location mapping table in step S2 is loaded into the BMC memory during system initialization, and its data structure includes: GPIO mapping entries: The keys are the GPIO controller number and pin number, and the values are physical location strings; I2C mapping entry: The key is the I2C bus number and device address, and the value is a physical location string; The mapping table supports hot updates, and the mapping relationship can be dynamically modified through the management interface when the server hardware configuration changes.
[0010] Preferably, the state machine judgment in step S3 further includes state transition logic for the power module: when PMBus communication is successful and the PRESENT state bit is 1, the state is in place; when the PRESENT state bit is 1 but the POWER_GOOD state bit is 0, the state is transitioned to fault; when PMBus communication times out or the PRESENT state bit is 0, the state is transitioned to out of place.
[0011] Preferably, the lock-free circular queue in step S4 is implemented using a single producer and single consumer model, wherein: the interrupt service routine acts as the producer and writes the event identifier to the tail of the queue; the background thread acts as the consumer and reads the event identifier from the head of the queue.
[0012] A real-time and accurate monitoring system for the location status and quantity of key server components includes: The hardware signal sensing layer includes a CPLD chip or FPGA chip, a BMC chip and their peripheral circuits. The CPLD or FPGA chip is connected to the PRSNT# signal of the CPU socket, the PRSNT# signal of the DIMM socket, the PRSNT# signal of the fan connector and the TACH signal through GPIO pins. The BMC chip is connected to the SPD interface of the DIMM socket through the I2C bus and to the PMBus interface of the power module through the SMBus bus. The signal acquisition and logic parsing layer includes a parallel signal acquisition module, a hardware digital filtering module, a debouncing module, and an interrupt generation module deployed in a CPLD or FPGA, as well as an I2C bus scan driver module, an address mapping query module, a multi-dimensional state machine engine module, a lockless queue management module, a hierarchical interrupt handling module, and a status confirmation module deployed in a BMC. The unified state management layer includes a monitoring daemon running in the BMC operating system, an embedded lightweight database, an event handling engine, a history analysis module, a Redfish API server, a web server, and a graphical display engine. The CPLD chip is powered independently of the host CPU and BMC reset signals. When the server crashes or the operating system crashes, the CPLD can still continue to perform signal acquisition and interrupt generation tasks, and notify the BMC to handle backlog events after the BMC recovers.
[0013] This invention provides a method and system for real-time and accurate monitoring of the location status and quantity of key server components. Compared with existing technologies, it has the following advantages: 1. In this invention, by constructing a pre-built hardware interface and physical location mapping table, the correspondence between GPIO pin numbers, I2C bus addresses, and physical location strings is accurately recorded in key-value pairs. This breaks through the traditional solution's limitation of only reporting the binary state of presence or absence to precise slot-level positioning. It can accurately output hierarchical physical location information such as CPU slot number, DIMM slot number, fan slot number, or power supply slot number. At the same time, by establishing an independent state machine instance for each monitored component, defining a multi-dimensional set of states such as present and normal, present but abnormal, absent, and fault, and combining the dual-mode detection mechanism of CPLD parallel acquisition of PRSNT# pin level signal and BMC active scanning of I2C bus device addresses, it effectively distinguishes the state differences in three dimensions: physical presence, electrical connection, and functional operation. This completely solves the problem that traditional solutions cannot accurately report which specific slot is idle or faulty, and cannot distinguish between physical absence and electrical fault. This allows maintenance personnel to locate faulty components directly through a graphical interface without manually checking the silkscreen, significantly improving maintenance efficiency and fault diagnosis accuracy. 2. In this invention, a CPLD is used to detect pin level changes in parallel with a microsecond-level response time. The GPIO of the CPLD and BMC is configured to work in interrupt-triggered mode. When a rising or falling edge change of the PRSNT# pin level is detected, a hardware interrupt signal is immediately generated. Combined with a lockless circular queue, a single producer-single consumer event processing mechanism is implemented. The latency of capturing and reporting state changes is shortened from the second to minute level of the traditional periodic polling method to the millisecond level, realizing an instant response to hot-plug events. In addition, the CPLD runs independently of the BMC and the host CPU, and its power supply is independent of the reset signal of the host CPU and the BMC. The internal hardware debouncing filter module performs debouncing processing on the signal. Even in extreme cases such as server crash, operating system unresponsiveness, or abnormal BMC reset, it can still continuously monitor component insertion and removal events and accurately record timestamps. After the BMC recovers, the backlog of events is processed through interrupt notification. This completely breaks through the reliability bottleneck of traditional solutions that rely on operating system drivers or high-level protocols and cannot work when the system is abnormal. Attached Figure Description
[0014] Figure 1 This is a flowchart of a method and system for real-time and accurate monitoring of the location status and quantity of key server components proposed in this invention. Figure 2 This invention presents a system framework diagram of a method and system for real-time and accurate monitoring of the location status and quantity of key server components. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Please see Figures 1-2 The present invention provides two technical solutions, specifically including the following embodiments: Example 1: A method for real-time and accurate monitoring of the location status and quantity of key server components includes the following steps: S1. The physical level signals or bus communication signals of key components are directly acquired through the hardware interface. The hardware interface is divided into the general purpose input / output (GPIO) pins of the programmable logic device (CPLD), the I2C bus interface of the baseboard management controller (BMC), the GPIO pins of the CPLD or BMC, and the SMBus bus of the BMC. Among them, the GPIO pins of the CPLD are used to acquire the level signal of the CPU socket PRSNT# pin. When the CPU is inserted, this pin is pulled low to ground level, and when it is removed, it is pulled up to high level by the pull-up resistor. The I2C bus interface of the baseboard management controller (BMC) is used to actively scan the serial presence detection (SPD) device address on each memory channel. If the EEPROM device response is detected, the memory module is determined to be in place. The GPIO pins of the CPLD or BMC are used to acquire the level signal of the fan connector PRSNT# pin to determine the fan's presence status, and at the same time read the fan's TACH speed signal. The SMBus bus of the BMC is used to establish PMBus communication with the power module and read the PRESENT status bit and POWER_GOOD status bit of the power module. S2. The hardware interface and physical location mapping table pre-installed in the BMC memory is invoked. This mapping table records the correspondence between GPIO pin numbers, I2C bus addresses, and physical location strings in key-value pairs. The signal identifier acquired in step S1 is used as the input key to query the mapping table, and the corresponding physical location string is output. The physical location string includes the CPU socket number, DIMM socket number, fan slot number, or power supply slot number. The CPLD acquires the level signal of the CPU socket PRSNT# pin in parallel, including: The CPLD detects pin level changes with a microsecond-level response time and operates independently of the BMC and host CPU, continuously monitoring component insertion and removal events even when the server crashes or the operating system becomes unresponsive. The CPLD integrates a hardware debouncing filter module to debouncing the PRSNT# pin signal, eliminating level oscillations caused by mechanical contact jitter. The BMC's I2C bus interface actively scans the SPD device addresses on each memory channel, including: Iterate through the pre-configured list of I2C bus numbers and the list of slave device addresses; For each bus address combination, an I2C probe command is executed. If an acknowledgment signal is received, the presence of the corresponding DIMM slot is recorded, and memory parameter information in the SPD EEPROM is further read. The hardware interface-physical location mapping table is loaded into the BMC memory during system initialization, and its data structure includes: GPIO mapping entries: The keys are the GPIO controller number and pin number, and the values are physical location strings; I2C mapping entry: The key is the I2C bus number and device address, and the value is a physical location string; The mapping table supports hot updates, allowing the mapping relationship to be dynamically modified through the management interface when the server hardware configuration changes. S3. Establish an independent state machine instance for each monitored component, and define the state set including: in and normal, in but abnormal, out, and fault. When PMBus communication is successful and the PRESENT status bit is 1, the state is in; when the PRESENT status bit is 1 but the POWER_GOOD status bit is 0, the state changes to fault; when PMBus communication times out or the PRESENT status bit is 0, the state changes to out. S4. Configure the GPIO of the CPLD and BMC to work in interrupt-triggered mode. When a rising or falling edge change is detected in the PRSNT# pin level, a hardware interrupt signal is immediately generated. In the interrupt service routine, the interrupt event identifier is written to the lockless circular queue. The lockless circular queue is implemented in a single producer-single consumer mode, where: the interrupt service routine acts as the producer and writes the event identifier to the tail of the queue; the background thread acts as the consumer and reads the event identifier from the head of the queue. The interrupt event identifier includes the GPIO pin number that changed, the direction of the level change, and the timestamp. S5. Write the component status data obtained in step S3 into a lightweight database deployed in the BMC flash memory. The status data includes component type, physical location, current status, component model, serial number, and status change timestamp. Output structured data to the external management system through a Redfish-compliant RESTful API. The structured data is encapsulated in JSON format and includes a Component field indicating the component type, a Location field indicating the physical location string, a Present field indicating the on-site status, and a PartNumber field indicating the component number.
[0017] A real-time and accurate monitoring system for the location status and quantity of key server components includes: The hardware signal sensing layer includes a CPLD chip or FPGA chip, a BMC chip and their peripheral circuits. The CPLD or FPGA chip is connected to the PRSNT# signal of the CPU socket, the PRSNT# signal of the DIMM socket, the PRSNT# signal of the fan connector and the TACH signal through GPIO pins. The BMC chip is connected to the SPD interface of the DIMM socket through the I2C bus and to the PMBus interface of the power module through the SMBus bus. The signal acquisition and logic parsing layer includes a parallel signal acquisition module, a hardware digital filtering module, a debouncing module, and an interrupt generation module deployed in a CPLD or FPGA, as well as an I2C bus scan driver module, an address mapping query module, a multi-dimensional state machine engine module, a lockless queue management module, a hierarchical interrupt handling module, and a status confirmation module deployed in a BMC. The unified state management layer includes a monitoring daemon running in the BMC operating system, an embedded lightweight database, an event handling engine, a history analysis module, a Redfish API server, a web server, and a graphical display engine. The CPLD chip is powered independently of the host CPU and BMC reset signals. When the entire server crashes or the operating system crashes, the CPLD can still continue to perform signal acquisition and interrupt generation tasks, and notify the BMC to handle backlogged events through interrupts after the BMC recovers.
[0018] Example 2: Based on Example 1, let's take a data center deploying a dual-socket server (Node: Rack05-U10) running high-performance computing tasks as an example. This server is configured with 2 CPUs, 16 DIMM memory modules, 4 hot-swappable fans, and 2 redundant power supplies. During operation, the system frequently experiences memory access errors. Traditional solutions can only report memory faults, failing to pinpoint the exact cause and resulting in multiple ineffective maintenance steps. S1. The physical level signals or bus communication signals of key components are directly acquired through the hardware interface. The CPLD acquires all key component signals in parallel through the GPIO pins: CPU1_PRSNT# (GPIO_A12) and CPU2_PRSNT# (GPIO_A13) are both low, confirming that the two CPUs are physically present; the PRSNT# signals of fans 1-4 are all low, and the TACH speed signals are 4200RPM, 4150RPM, 0RPM, and 4180RPM respectively, indicating that fan 3 is present but stopped; the BMC scans through the I2C bus and detects the SPD device response at addresses such as Bus2_Addr0x50 and Bus3_Addr0x51, but intermittent communication timeouts occur at Bus4_Addr0x52 (corresponding to DIMM_B2); S2. The hardware interface and physical location mapping table pre-installed in the BMC memory is invoked. This mapping table records the correspondence between GPIO pin numbers, I2C bus addresses, and physical location strings in key-value pairs. The signal identifier acquired in step S1 is used as the input key to query the mapping table, and the corresponding physical location string is output. The physical location string includes the CPU socket number, DIMM socket number, fan slot number, or power supply slot number. The CPLD acquires the level signal of the CPU socket PRSNT# pin in parallel, including: The CPLD detects pin level changes with a microsecond-level response time and operates independently of the BMC and host CPU, continuously monitoring component insertion and removal events even when the server crashes or the operating system becomes unresponsive. The CPLD integrates a hardware debouncing filter module to debouncing the PRSNT# pin signal, eliminating level oscillations caused by mechanical contact jitter. The BMC's I2C bus interface actively scans the SPD device addresses on each memory channel, including: Iterate through the pre-configured list of I2C bus numbers and the list of slave device addresses; For each bus address combination, an I2C probe command is executed. If an acknowledgment signal is received, the presence of the corresponding DIMM slot is recorded, and memory parameter information in the SPD EEPROM is further read. The hardware interface-physical location mapping table is loaded into the BMC memory during system initialization, and its data structure includes: GPIO mapping entries: The keys are the GPIO controller number and pin number, and the values are physical location strings; I2C mapping entry: The key is the I2C bus number and device address, and the value is a physical location string; The mapping table supports hot updates, allowing the mapping relationship to be dynamically modified through the management interface when the server hardware configuration changes.
[0019] The exact location of the fault is: Rack05-U10, the second memory slot in the B channel of the second CPU, and system fan 3; S3. Establish an independent state machine instance for each monitored component, defining the state set as: in and normal, in but abnormal, out, and fault. When PMBus communication is successful and the PRESENT status bit is 1, the state is in; when the PRESENT status bit is 1 but the POWER_GOOD status bit is 0, the state transitions to fault; when PMBus communication times out or the PRESENT status bit is 0, the state transitions to out. DIMM_B2 state machine: I2C communication intermittent timeout (unstable electrical connection); PRESNT# pin continuously low level (physical in); state transition to: in but abnormal - poor contact; FAN3 state machine: PRESNT# low level (physical in); PWM control signal is issued normally, TACH=0 (functional fault); state transition to: fault - stopped; power module state machine: PSU1 and PSU2 both PRESENT are 1, POWER_GOOD are both 1; state remains: in and normal. S4. Configure the GPIO of the CPLD and BMC to work in interrupt-triggered mode. When a rising or falling edge change is detected in the PRSNT# pin level, a hardware interrupt signal is immediately generated. In the interrupt service routine, the interrupt event identifier is written to the lockless circular queue. The lockless circular queue is implemented in a single producer-single consumer mode, where: the interrupt service routine acts as the producer and writes the event identifier to the tail of the queue; the background thread acts as the consumer and reads the event identifier from the head of the queue. The interrupt event identifier includes the GPIO pin number that changed, the direction of the level change, and the timestamp. S5. Write the component status data obtained in step S3 into a lightweight database deployed in the BMC flash memory. The status data includes component type, physical location, current status, component model, serial number, and status change timestamp. Output structured data to the external management system through a Redfish-compliant RESTful API. The structured data is encapsulated in JSON format and includes a Component field indicating the component type, a Location field indicating the physical location string, a Present field indicating the on-site status, and a PartNumber field indicating the component number.
[0020] The web management interface refreshes the server backplane topology in real time: the DIMM_B2 location flashes orange, indicating a poor contact warning; the FAN_Bay_3 location is highlighted in red, indicating a fault requiring immediate attention; other normal components are displayed in green; when maintenance personnel click on the DIMM_B2 location, a detailed information panel pops up, showing: Current status: Unstable electrical connection; Historical records: 12 momentary disconnection events occurred in this slot in the past 72 hours; Trend analysis: The frequency of poor contact is on the rise; Recommended action: Check the oxidation of the memory module's gold fingers, and replace the slot or memory module if necessary.
[0021] Maintenance Execution and Verification: Based on the precise location information, the maintenance personnel directly operated the DIMM_B2 slot of the Rack05-U10 server: the memory module was removed, and slight oxidation was found on the gold fingers. The gold fingers and slot were cleaned with anhydrous ethanol, and the memory module was reinserted to ensure it was fully seated. The system detected in real time that I2C communication had returned to normal, and the status was automatically updated to be in place and normal. At the same time, the faulty fan 3 was replaced, and the system detected that the TACH speed had returned to 4200RPM, and the status was updated to be in place and normal.
[0022] The following table compares the results with those of the traditional approach:
[0023] Conclusion: Through a specific application scenario of precise memory fault location and predictive maintenance in a dual-socket server, the significant technical advantages and practical value of this invention in actual data center operation and maintenance are fully verified. In this example, the system successfully located the fuzzy memory fault information that traditional solutions could only report to the DIMM_B2 slot (CPU2, Channel B, Slot2), and simultaneously identified the fan 3's shutdown fault, achieving independent differentiation and precise location of concurrent faults in multiple components. Through CPLD microsecond-level signal detection combined with an interrupt-driven mechanism, state changes can be captured, parsed, and reported within 50ms, which is 20-100 times faster than the second- to minute-level latency of traditional polling solutions. In terms of actual operation and maintenance results, fault diagnosis time is shortened from 2-3 hours of traditional manual diagnosis to within 1 second of automatic location, avoiding ineffective operations such as mistakenly disassembling normal components, reducing the number of maintenance operations from 3 to 1 precise operation, and significantly reducing downtime from 4 hours to 15 minutes, thus significantly improving operation and maintenance efficiency.
[0024] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for real-time and accurate monitoring of the location status and quantity of key server components, characterized in that: Includes the following steps: S1. Directly acquire physical level signals or bus communication signals of key components through hardware interfaces; S2. Call the hardware interface and physical location mapping table preset in the BMC memory. The mapping table records the correspondence between GPIO pin number, I2C bus address and physical location string in key-value pair form. Use the signal identifier collected in step S1 as the input key value to query the mapping table and output the corresponding physical location string. The physical location string includes CPU socket number, DIMM socket number, fan slot number or power supply slot number. S3. Establish an independent state machine instance for each monitored component, and define the state set as: in place and normal, in place but abnormal, out of place, and faulty. S4. Configure the GPIO of CPLD and BMC to work in interrupt-triggered mode. When a rising or falling edge change is detected in the PRSNT# pin level, a hardware interrupt signal is immediately generated. In the interrupt service routine, the interrupt event identifier is written to the lockless circular queue. The interrupt event identifier includes the GPIO pin number that changed, the direction of the level change, and the timestamp. S5. Write the component status data obtained in step S3 into a lightweight database deployed in the BMC flash memory. The status data includes component type, physical location, current status, component model, serial number, and status change timestamp. Output structured data to the external management system through a Redfish-compliant RESTful API. The structured data is encapsulated in JSON format and includes a Component field indicating the component type, a Location field indicating the physical location string, a Present field indicating the on-site status, and a PartNumber field indicating the component number.
2. The method for real-time and accurate monitoring of the location status and quantity of key server components according to claim 1, characterized in that: In step S1, the hardware interface is divided into the general purpose input / output (GPIO) pins of the programmable logic device (CPLD), the I2C bus interface of the baseboard management controller (BMC), the GPIO pins of the CPLD or BMC, and the SMBus bus of the BMC. Specifically, the GPIO pins of the CPLD are used to acquire the level signal of the CPU socket PRSNT# pin. When the CPU is inserted, this pin is pulled low to ground, and when removed, it is pulled high by a pull-up resistor. The I2C bus interface of the BMC is used to actively scan the serial presence detection (SPD) device addresses on each memory channel. If an EEPROM device response is detected, the memory module is determined to be present. The GPIO pins of the CPLD or BMC are used to acquire the level signal of the fan connector PRSNT# pin to determine the fan's presence status, and simultaneously read the fan's TACH speed signal. The SMBus bus of the BMC is used to establish PMBus communication with the power module and read the PRESENT and POWER_GOOD status bits of the power module.
3. The method for real-time and accurate monitoring of the location status and quantity of key server components according to claim 2, characterized in that: The CPLD parallel acquisition of the level signal of the CPU socket PRSNT# pin includes: The CPLD detects pin level changes with a microsecond-level response time and operates independently of the BMC and host CPU, continuously monitoring component insertion and removal events even when the server crashes or the operating system becomes unresponsive. The CPLD integrates a hardware debouncing filter module to debouncing the PRSNT# pin signal, eliminating level oscillations caused by mechanical contact jitter.
4. The method for real-time and accurate monitoring of the location status and quantity of key server components according to claim 2, characterized in that: The BMC's I2C bus interface actively scans the SPD device addresses on each memory channel, including: Iterate through the pre-configured list of I2C bus numbers and the list of slave device addresses; For each bus address combination, execute an I2C probe command. If an acknowledgment signal is received, record that the DIMM slot corresponding to that address is in place, and further read the memory parameter information in the SPD EEPROM.
5. The method for real-time and accurate monitoring of the location status and quantity of key server components according to claim 1, characterized in that: The hardware interface and physical location mapping table in step S2 is loaded into the BMC memory during system initialization, and its data structure includes: GPIO mapping entries: The keys are the GPIO controller number and pin number, and the values are physical location strings; I2C mapping entry: The key is the I2C bus number and device address, and the value is a physical location string; The mapping table supports hot updates, and the mapping relationship can be dynamically modified through the management interface when the server hardware configuration changes.
6. The method for real-time and accurate monitoring of the location status and quantity of key server components according to claim 1, characterized in that: The state machine judgment in step S3 also includes the state transition logic of the power module: when PMBus communication is successful and the PRESENT state bit is 1, the state is in place; when the PRESENT state bit is 1 but the POWER_GOOD state bit is 0, the state is transitioned to fault; when PMBus communication times out or the PRESENT state bit is 0, the state is transitioned to out of place.
7. The method for real-time and accurate monitoring of the location status and quantity of key server components according to claim 1, characterized in that: The lock-free circular queue in step S4 is implemented using a single producer and single consumer pattern, wherein: the interrupt service routine acts as the producer and writes the event identifier to the tail of the queue; the background thread acts as the consumer and reads the event identifier from the head of the queue.
8. A real-time accurate monitoring system for the location status and quantity of key server components, based on any one of the real-time accurate monitoring methods for the location status and quantity of key server components according to claims 1-7, characterized in that: include: The hardware signal sensing layer includes a CPLD chip or FPGA chip, a BMC chip and their peripheral circuits. The CPLD or FPGA chip is connected to the PRSNT# signal of the CPU socket, the PRSNT# signal of the DIMM socket, the PRSNT# signal of the fan connector and the TACH signal through GPIO pins. The BMC chip is connected to the SPD interface of the DIMM socket through the I2C bus and to the PMBus interface of the power module through the SMBus bus. The signal acquisition and logic parsing layer includes a parallel signal acquisition module, a hardware digital filtering module, a debouncing module, and an interrupt generation module deployed in a CPLD or FPGA, as well as an I2C bus scan driver module, an address mapping query module, a multi-dimensional state machine engine module, a lockless queue management module, a hierarchical interrupt handling module, and a status confirmation module deployed in a BMC. The unified state management layer includes a monitoring daemon running in the BMC operating system, an embedded lightweight database, an event handling engine, a history analysis module, a Redfish API server, a web server, and a graphical display engine. The CPLD chip is powered independently of the host CPU and BMC reset signals. When the server crashes or the operating system crashes, the CPLD can still continue to perform signal acquisition and interrupt generation tasks, and notify the BMC to handle backlog events after the BMC recovers.