Memory fault diagnosis method and system, and baseboard management controller
Patent Information
- Application Number
- CN202611325090.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]为了解决现有技术中存在的上述问题和缺陷的至少一个方面,本发明的实施例提供了一种内存故障诊断方法及其系统、基板管理控制器,通过基板管理控制器经由基本输入输出系统周期性采集DDR5内存的错误巡检及纠正寄存器数据,解析为包含Bank组地址、Bank地址、行地址、最大行错误计数和累计错误计数的五元组数据,再通过按地址粒度从细到粗逐层匹配的内存故障诊断算法,实现了对DDR5可纠正错误CE故障的主动监测与精准诊断,克服了DDR5片内ECC导致CPU无法感知错误从而传统中断上报机制失效的技术难题
[0022]本发明的实施例提供的内存故障诊断方法及其系统、基板管理控制器具有以下优点中的至少一个或至少一个优点的一部分:
Smart Images

Figure CN122816979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of memory testing, and in particular to a memory fault diagnosis method and system, and a baseboard management controller. Background Technology
[0002] As a core component of a computing system, the reliability of memory directly affects data integrity and system stability. Memory errors typically begin with a single-bit correctable error (CE). If not detected and handled in time, it may evolve into an uncorrectable error (UCE), leading to system crashes or data corruption.
[0003] In DDR4 and earlier memory architectures, the main method used was to combine the ECC error correction mechanism with the Machine Check Architecture (MCA) of the Central Processing Unit (CPU). When an error occurred, a System Management Interrupt (SMI) was triggered, which was then reported by the Basic Input Output System (BIOS) to the Baseboard Management Controller (BMC). This information was then combined with information from Model Specific Registers (MSRs) to locate and diagnose the fault.
[0004] However, DDR5 memory introduced on-chip ECC technology, which moved the error correction function to within the memory chip itself. Single-bit errors are automatically corrected by the memory chip before data is transmitted to the CPU, so the data input to the CPU does not show a fault, making the CPU unable to detect the error and unable to trigger an SMI interrupt. The traditional monitoring mechanism that relies on CPU reporting thus becomes ineffective, the system cannot respond in a timely manner, and the fault may continue to accumulate and eventually evolve into UCE, seriously threatening system stability.
[0005] Although DDR5 defines registers related to Error Check and Scrub (ECS) (mode registers MR15 to MR20), the information they can reflect is relatively limited due to resource constraints. Specifically: MR16-MR18 only reflect the address information of the row with the largest error in a single round of inspection; MR19's maximum row error count (REC) threshold is 4, and counting only occurs when the row error exceeds 4, reflecting a count range rather than an exact value; MR20's cumulative error count (EC) is also a range value. Therefore, based solely on a single ECS data set, it is only possible to determine if a memory problem exists, but not to determine the specific type of fault.
[0006] In view of this, a novel memory fault diagnosis method and system, as well as a baseboard management controller, are proposed to solve the above problems in whole or in part. Summary of the Invention
[0007] To address at least one of the aforementioned problems and deficiencies in the existing technology, embodiments of the present invention provide a memory fault diagnosis method and system, and a baseboard management controller. The baseboard management controller periodically collects error inspection and correction register data from DDR5 memory via the basic input / output system, parses it into a five-tuple of data including Bank group address, Bank address, row address, maximum row error count, and cumulative error count, and then uses a memory fault diagnosis algorithm that matches errors layer by layer from fine to coarse address granularity to achieve proactive monitoring and accurate diagnosis of DDR5 correctable error (CE) faults. This overcomes the technical challenge of the CPU being unable to detect errors due to DDR5 on-chip ECC, thus rendering the traditional interrupt reporting mechanism ineffective. Based on this, by combining software and hardware repair capabilities, row faults are isolated in a graded manner, or Bank faults are isolated through Bank isolation, ensuring stable system operation without replacing the memory. The technical solution is as follows:
[0008] According to one aspect of the present invention, a memory fault diagnosis method is provided, the memory fault diagnosis method comprising:
[0009] The baseboard management controller periodically reads the error check and correction register of the memory chip through the basic input / output system and obtains the error check and correction data;
[0010] The baseboard management controller parses the error inspection and correction data, extracts the Bank group address, Bank address, row address, maximum row error count, and cumulative error count, and obtains the quintuple data;
[0011] The baseboard management controller arranges the quintuple data corresponding to each cycle in chronological order to obtain the current quintuple data and historical quintuple data.
[0012] Based on the hierarchical address information in the current quintuple data, the current quintuple data is matched with the historical quintuple data in order of address granularity from fine to coarse. The memory fault type is determined based on the first successfully matched level.
[0013] According to another aspect of the present invention, a memory fault diagnosis system is provided, including a baseboard management controller and a basic input / output system, the memory fault diagnosis system further comprising:
[0014] The baseboard management controller is configured to periodically read the error check and correction register of the memory chip through the basic input / output system and obtain error check and correction data.
[0015] Then, the error inspection and correction data is parsed to extract the Bank group address, Bank address, row address, maximum row error count, and cumulative error count, resulting in a quintuple of data.
[0016] Next, the quintuple data corresponding to each period are arranged in chronological order to obtain the current quintuple data and historical quintuple data;
[0017] Based on the hierarchical address information in the current quintuple data, the current quintuple data is matched with the historical quintuple data in order of address granularity from fine to coarse. The fault type is determined based on the level of the first successful match.
[0018] According to another aspect of the present invention, a substrate management controller is provided, the substrate management controller comprising:
[0019] The data acquisition module is configured to periodically read the error check and correction register of the memory chip through the basic input / output system and acquire error check and correction data.
[0020] The data parsing module is configured to parse error inspection and correction data, extract the Bank group address, Bank address, row address, maximum row error count, and cumulative error count to obtain quintuple data;
[0021] The fault diagnosis module is configured to arrange the five-tuple data corresponding to each cycle in chronological order to obtain the current five-tuple data and historical five-tuple data. Based on the hierarchical address information in the current five-tuple data, the module performs hierarchical matching between the current five-tuple data and the historical five-tuple data in order of address granularity from fine to coarse, and determines the memory fault type based on the first successfully matched hierarchical level.
[0022] The memory fault diagnosis method and system, and the baseboard management controller provided in the embodiments of the present invention have at least one or a portion of the following advantages:
[0023] (1) The error inspection and correction register data of the memory chip is periodically collected by the baseboard management controller through the basic input / output system and parsed into a five-tuple data containing the Bank group address, Bank address, row address, maximum row error count and cumulative error count. Then, the fault type is determined by the memory fault diagnosis algorithm. It can realize the active monitoring and identification of DDR5 CE faults without relying on CPU interrupt reporting, thus improving the accuracy of DDR5 CE fault detection.
[0024] (2) By using the hierarchical address information in the current quintuple data as a benchmark, and following the order of address granularity from fine to coarse (row address - Bank address - Bank group address - device), the current quintuple data is matched hierarchically with the historical quintuple data. The first successfully matched level is determined as the fault type. This can diagnose different types of DDR5 CE faults such as row faults, Bank faults, Bank group faults and device faults, thus improving the granularity of fault diagnosis.
[0025] (3) The timed task is started through the baseboard management controller and a call request is sent to the basic input / output system through the intelligent platform management interface. The basic input / output system responds to the call request by reading the error inspection and correction register of the corresponding memory chip through the memory controller and returning the data. The collection period is set to 30 minutes to 24 hours, which realizes low-overhead and periodic collection of fault data and ensures the real-time fault detection without affecting the system's operating performance.
[0026] (4) By parsing the Bank group address, Bank address and row address from mode register 16 to mode register 18 according to the DDR5 specification, parsing the maximum row error count from mode register 19, and parsing the cumulative error count from mode register 20, the standardized parsing of DDR5 ECS register data is realized, providing an accurate data basis for subsequent fault diagnosis.
[0027] (5) After the fault type is diagnosed by the baseboard management controller, the corresponding fault isolation method is selected to perform isolation operation on the faulty memory area in combination with the pre-acquired soft repair capability and / or hard repair capability and the remaining resources. This achieves accurate isolation of the faulty area without replacing the memory, ensuring the continuous and stable operation of the system.
[0028] (6) When the fault type is a row fault, temporary repair is performed by prioritizing soft isolation. If soft isolation has been performed, it is further determined whether the hard isolation conditions are met (the maximum row error count is the maximum value under the same Bank group address and Bank address and is greater than or equal to the preset threshold, and the historical data of the same address shows an increasing trend of the maximum row error count in the time series). If the conditions are met, permanent repair is performed through hard isolation, thus realizing the hierarchical repair strategy for row faults and optimizing the utilization of hard repair resources while ensuring the repair effect.
[0029] (7) When the fault type is Bank fault, the Bank isolation parameters are sent to the basic input / output system through the baseboard management controller. During the power-on startup phase, the basic input / output system calculates all system addresses corresponding to the Bank that needs to be isolated based on the Bank isolation parameters and marks them as reserved, thus achieving Bank-level fault isolation at the cost of sacrificing a small amount of memory capacity.
[0030] (8) The baseboard management controller determines the fault quantity range based on the setting of the maximum row error count in the mode register 19, and sets the corresponding alarm level according to the fault quantity range, thereby realizing the graded early warning of fault severity, which makes it easier for maintenance personnel to take differentiated response measures according to the fault level.
[0031] (9) The risk level of memory chips is classified by the baseboard management controller according to the setting of mode register 20, realizing the graded evaluation of memory chips based on ECS data, which makes it easier for maintenance personnel to intuitively grasp the health status of memory chips and take differentiated maintenance strategies. Attached Figure Description
[0032] These and / or other aspects and advantages of the present invention will become apparent and readily understood from the following description of preferred embodiments taken in conjunction with the accompanying drawings, in which:
[0033] Figure 1 This is a flowchart illustrating the overall execution framework of a memory fault diagnosis method according to an embodiment of the present invention.
[0034] Figure 2 for Figure 1 The flowchart shown below illustrates the specific execution process of the memory fault diagnosis algorithm within the overall execution flow.
[0035] Figure 3 for Figure 1 The diagram shows the specific execution flow for fault isolation within the overall execution process.
[0036] Figure 4 In order to be in Figure 3The graph shows the trend of the minimum value (REC_min) of REC in the mode register MR19 as it is set during the fault isolation process. Detailed Implementation
[0037] The technical solution of the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings. In this specification, the same or similar reference numerals indicate the same or similar components. The following description of the embodiments of the present invention with reference to the accompanying drawings is intended to explain the overall inventive concept of the present invention and should not be construed as a limitation thereof.
[0038] In existing technologies, although DDR5 has ECS-related registers (MR15-MR20), due to resource constraints, the information they can reflect is relatively limited. The main problems are as follows:
[0039] MR16-MR19 reflect the address and count of the largest error line in a round of inspection, but they cannot reflect all the fault information and historical information of the inspection.
[0040] MR19 is the maximum error row failure count. Its threshold (RETC) is 4, meaning MR19 only counts when the number of row failures exceeds 4. Furthermore, MR19 reflects a count range, not an exact value. Its minimum value formula is... The formula for the maximum value is: The definition of the MR19 register and the corresponding intervals after each bit is set are shown in Table 1 and Table 2, respectively:
[0041] Table 1 MR19 Register Definitions
[0042]
[0043] Table 2. Range of fault numbers corresponding to bit positions
[0044]
[0045] MR20 reflects the number of erroneous rows in row mode and the number of erroneous codewords in codeword mode. Similar to MR19, the threshold for MR20 is related to the ETC field of MR15 and the density of the memory chip. The formula for calculating the minimum value of each interval is as follows: The formula for calculating the maximum value is: The definitions of the MR20 register and the corresponding intervals after each bit position (taking ETC as 256 and Density as 16G as examples) are shown in Tables 3 and 4 respectively:
[0046] Table 3 MR20 Register Definitions
[0047]
[0048] Table 4. Range of fault numbers corresponding to bit positions
[0049]
[0050] Therefore, the accuracy of existing CE fault detection and analysis technologies is not high. Based on a single ECS message, it can only determine whether a problem exists in the memory, but cannot determine the specific type of fault. To address the above-mentioned problems in existing technologies, this invention provides a memory fault diagnosis method based on the Error Check and Scrub (ECS) register.
[0051] See Figure 1 This demonstrates the overall execution flow of the memory fault diagnosis method. It showcases the complete closed-loop process from the initialization and preparation of the Baseboard Management Controller (BMC) to the final completion of fault diagnosis and fault isolation.
[0052] like Figure 1 As shown, this memory fault diagnosis method mainly includes the following steps:
[0053] Step S110: The baseboard management controller periodically reads the error inspection and correction register of the memory chip through the basic input / output system and obtains the error inspection and correction data;
[0054] Step S120: The baseboard management controller parses the error inspection and correction data, extracts the Bank group address, Bank address, row address, maximum row error count and cumulative error count, and obtains the quintuple data;
[0055] Step S130: The substrate management controller arranges the quintuple data corresponding to each cycle in chronological order to obtain the current quintuple data and historical quintuple data;
[0056] Step S140: Based on the hierarchical address information in the current quintuple data, perform hierarchical matching between the current quintuple data and the historical quintuple data in order of address granularity from fine to coarse, and determine the memory fault type based on the first successfully matched hierarchical level.
[0057] The Baseboard Management Controller (BMC), as the core out-of-band management component of the server, works in conjunction with the Basic Input Output System (BIOS) to periodically collect, parse, diagnose, and isolate memory error check and scrub (ECS) data (ECS data). The BMC sends call requests to the BIOS through the Intelligent Platform Management Interface (IPMI) or other interfaces for communication and data exchange. The following details the execution steps of the memory fault diagnosis method of this invention:
[0058] Before initiating the memory fault diagnosis algorithm, the BMC first assesses the repair capabilities of the memory chips. Specifically, the BMC reads the Serial Presence Detect (SPD) and Memory Mode Register (MR) data to obtain information on the capabilities of Soft Post Package Repair (sPPR) and Hard Post Package Repair (hPPR), as well as the remaining resource status.
[0059] sPPR is a temporary repair method; the repair becomes invalid after a host reboot. hPPR, on the other hand, is a permanent repair method. Both sPPR and hPPR repair row-level failures. For sPPR, typically one isolation is supported per Bank Group (BG), and some vendors support one isolation per Bank; this capability information can be obtained from SPD (Service Packet Management). Furthermore, sPPR usually shares resources with hPPR; therefore, if hPPR has exhausted resources, even if the BG has not yet undergone sPPR, it is usually impossible to perform sPPR again.
[0060] In addition to using certain bits in the MR register to query the sPPR and / or hPPR capabilities, the sPPR and / or hPPR capabilities can also be queried using the BIOS / UEFI firmware, or using firmware or tools provided by the memory manufacturer.
[0061] It should be noted that this step is optional. If the user's requirement is to isolate the faulty memory after a memory failure is detected, then sPPR and / or hPPR capabilities can be obtained before detecting whether the memory is faulty.
[0062] In one specific embodiment, the specific implementation process of step S110 is as follows:
[0063] The BMC works with the BIOS to periodically collect ECS register data. The specific steps are as follows: The BMC starts a scheduled task to periodically collect ECS data. The BMC sends a request to the BIOS via IPMI or other methods. In response to the request, the BIOS reads the corresponding ECS register from the memory controller and returns the error check and correction data to the BMC. The ECS data collection period can be flexibly set as needed, typically a minimum of 30 minutes and a maximum of 24 hours.
[0064] In one specific embodiment, the specific implementation process of step S120 is as follows:
[0065] After receiving the call request, the BIOS preferably reads the mode registers 16 to 20 (i.e., MR16-MR20) of the corresponding memory chip through the memory controller and sends them to the BMC. The BMC parses the contents of the above registers and extracts five fields: Bank Group Address (hereinafter referred to as BG), Bank Address (hereinafter referred to as BA), Row Address (hereinafter referred to as Row), Maximum Row Error Count (hereinafter referred to as REC), and Accumulated Error Count (hereinafter referred to as EC), to obtain the quintuple data.
[0066] Specifically, according to the DDR5 specification (JESD79-5), BG, BA and Row are first parsed from MR16-MR18, where Row is contained in the three registers MR16-MR18, and MR18 also contains BG and BA information; then REC is parsed from MR19; and finally EC is parsed from MR20.
[0067] It should be noted that, although according to the DDR5 specification, the actual values of REC and EC need to be calculated according to formulas, such as REC calculation being related to the RETC threshold and EC calculation being related to the ETC threshold and memory chip density, and the specific parameters in different memory calculation formulas are also different, since both REC and EC are fault counting intervals, in the solution of this invention, it is not necessary to calculate the actual error count, but only to use the values of the corresponding bits of MR19 and MR20.
[0068] In one specific embodiment, the specific implementation process of steps S130-S140 is as follows:
[0069] BMC arranges the quintuple data corresponding to each cycle according to timestamps or chronological order to obtain the current quintuple data and historical quintuple data. Based on the hierarchical address information in the current quintuple data, it performs hierarchical matching between the current quintuple data and the historical quintuple data in order of address granularity from fine to coarse, and determines the memory fault type based on the first successfully matched hierarchical level.
[0070] See Figure 2 It shows that in Figure 1 The specific execution flow of the memory fault diagnosis algorithm mentioned above in the overall execution flow is shown.
[0071] like Figure 2 As shown, this memory fault diagnosis algorithm narrows down the fault range layer by layer from top to bottom by comparing the current (current) ECS data with historical ECS data. The core logic of this algorithm is to use a "layer-by-layer narrowing" matching strategy, that is, first look at Row, then Bank, then Bank Group, and finally attribute it to Device. Therefore, the specific fault types include row fault type, Bank fault type, Bank group fault type, and Device fault type. The specific steps are as follows:
[0072] The first step is data filtering: First, determine whether the REC value in the collected ECS data is greater than 0 (i.e., whether the fault data exceeds the threshold). If REC ≤ 0 in the current ECS data, it means there is no fault, and the fault diagnosis ends directly; if REC > 0, then filter out all other quintuple data with REC > 0 from the historical data of this memory chip, excluding the current data.
[0073] The second step is to determine row faults: If there is one or more records in the selected historical 5-tuple data that are completely consistent with the BG, BA, and Row of the current ECS, then the memory chip is determined to have a row fault. The fault address is recorded as the corresponding BG, BA, and Row, and the fault diagnosis process ends.
[0074] The third step is to determine a Bank fault: If no data in the historical data matches the BG, BA, and Row of the current ECS, continue searching for data that matches the BG and BA of the current ECS. If such data exists, the memory chip is determined to have experienced a Bank fault, the fault address is recorded as the corresponding BG and BA, and the fault diagnosis process ends.
[0075] It should be noted that when historical ECS data contains data with the same BG, BA, and Row, as well as data with the same BG and BA across different Rows, according to the layer-by-layer matching logic of this algorithm, since higher-level Bank faults will suppress lower-level Row faults, it will be preferentially identified as a Bank fault.
[0076] The fourth step is to determine a Bank group failure: If the historical ECS data does not contain data with the same BG and BA as the current ECS data, continue searching for data with the same BG as the current ECS data. If such data exists, the memory chip is determined to have experienced a Bank group failure, and the fault diagnosis process ends. Similarly, higher-level Bank group failures will suppress the determination of lower-level Bank failures and row failures, meaning they are prioritized as Bank group failures.
[0077] The fifth step is to determine the device failure: If there is no matching historical data at any of the above levels, it is determined that the memory chip has a device failure, that is, the entire memory chip has a defect, and the fault diagnosis process ends.
[0078] It should be noted that this algorithm requires traversing the data of every single memory chip in the memory to diagnose the overall memory fault. If the server has been restarted, historical data needs to be cleared and new data needs to be collected for diagnosis; if the memory configuration has not changed, historical diagnostic records can be used as a reference for a new round of diagnosis.
[0079] In a specific embodiment, an example of memory fault type determination is shown in Table 5:
[0080] Table 5 Examples of each fault type
[0081]
[0082] In the row failure example, if the BG, BA, and Row in the historical ECS data are the same as those in the current ECS data, it is determined to be a row failure, and soft isolation can be performed using sPPR; if the REC value continues to repeat and increases (for example, from 1 to 2 and then to 4) and exceeds the threshold, hard isolation using hPPR should be considered.
[0083] In the example of a Bank failure, a Bank failure is identified when the historical ECS data does not contain the same BG, BA, and Row as the current ECS data, but contains the same BG and BA. Even if some historical data contains the same BG, BA, and Row as the current data, it is still identified as a Bank failure because a higher-level Bank failure will suppress a lower-level Row failure.
[0084] In the Bank group failure example, if the historical ECS data does not contain the same BG and BA as the current ECS data, but the same BG exists, it is determined to be a Bank group failure. Similarly, a higher-level Bank group failure will suppress lower-level Bank failures and row failures, and will still be determined to be a Bank group failure.
[0085] Furthermore, after determining the memory fault type, the fault level can be set and an alarm can be triggered based on the REC bit setting of the mode register MR19, i.e., based on the range of fault numbers. Examples of specific settings are shown in Table 6:
[0086] Table 6. Fault Level and Alarm Level Settings Based on REC Setting Status of MR19
[0087]
[0088] Furthermore, after determining the type of memory fault, the memory fault diagnosis method also includes a memory chip classification step:
[0089] BMC classifies memory chips into risk levels based on the setting of mode register 20;
[0090] When no bits are set in mode register 20, the memory chip is determined to be at the error-free level.
[0091] When the low-order bit of mode register 20 is set and the number of faulty rows is less than the number of rows in a bank, the memory chip is determined to be of low risk level.
[0092] When the middle bit of mode register 20 is set and the number of erroneous rows is greater than the number of rows in a bank, the memory chip is determined to be of a high-risk level.
[0093] When the high-order bit of mode register 20 is set and the number of erroneous rows exceeds the number of rows in a Bank group, the memory chip is determined to be in a dangerous state.
[0094] For example, one executable strategy is to hierarchically classify memory chips based on the number of banks and bank groups. Taking a memory chip with 16 row address bits and 4 banks per bank group as an example, the number of rows per bank is 2. 16 =65536, the number of rows in each Bank group is 4×2 16 =262144. Therefore, the grade classification is:
[0095] When no bits are set in the mode register MR20, i.e. MR20 equals 0, the classification is A: Error Free.
[0096] If any bit in EC0-EC3 of the mode register MR20 is set, it indicates that the number of erroneous rows is less than the number of rows in a bank, and the classification is B: Low Risk.
[0097] If any bit in EC5-EC6 of the mode register MR20 is set, it indicates that the number of erroneous rows has exceeded the number of rows in a bank, and the classification is C: High Risk.
[0098] If any bit in EC6-EC7 of the mode register MR20 is set, it indicates that the number of erroneous rows has exceeded the number of rows in a Bank group, and the classification is D: Danger Condition.
[0099] It should be noted that the above error count does not refer to the number of error rows in a single bank or bank group, but rather the number of error rows in the entire memory chip.
[0100] In a specific embodiment, after step S140, the memory fault diagnosis method further includes: when the BMC detects that the fault diagnosis result is a fault in a certain area of memory, it selects the appropriate fault isolation method according to the fault type, soft repair capability and / or hard repair capability and the remaining resources.
[0101] See Figure 3 It shows that in Figure 1 The diagram illustrates the specific execution flow for fault isolation following the overall execution flow. It demonstrates the different isolation strategies and judgment conditions adopted for different types of faults (specifically, line faults and bank faults).
[0102] If the memory fault diagnosis result is a row fault:
[0103] First, attempt to perform sPPR soft isolation. Determine if the Bank group corresponding to the row address of the faulty row has already undergone sPPR soft isolation. If not, perform fault isolation or fault repair via sPPR; if memory supports performing sPPR once per Bank, determine whether sPPR can be performed based on the Bank.
[0104] Next, if sPPR soft isolation has already been performed, the condition for hPPR hard isolation is then determined. Specifically, for data diagnosed as row faults, the data is grouped by BG and BA. The ECS data with the largest REC within that BG and BA is selected as ECS_max. Then, the data with the same BG, BA, and Row as ECS_max are sorted in chronological order. If the REC shows an increasing trend over time, hPPR fault isolation for that BG, BA, and Row can be considered during BIOS startup.
[0105] See Figure 4 The figure shows the trend curve of the minimum value (REC_min) of REC in the mode register MR19 as it changes with the setting bit.
[0106] Given the limited resources of hPPR, hPPR fault isolation can be considered only when REC_min must be greater than a certain value.
[0107] like Figure 4 As shown, the trend of the minimum value of REC in MR19 as it changes with the setting position is illustrated: the horizontal axis x represents the setting index of REC[x] in MR19 (from 0 to 5); the vertical axis REC[x]_min represents the minimum fault count value; the blue curve indicates that according to the formula REC[x]_min=RETC×2 x (The default RETC threshold is 4) The minimum value calculated, for example, when x=2, the minimum value is 16; the red dashed line represents a decision threshold line set in the algorithm (y=16, i.e. x=2). This optional condition is set before performing hPPR isolation to avoid excessive consumption of limited hPPR resources.
[0108] In summary, the hPPR hard isolation conditions include: ① The REC of the current ECS is at its maximum value under the corresponding BG and BA conditions, and the REC value is greater than or equal to the preset threshold (e.g., Figure 2 The threshold shown is 16); ② Historical ECS data for the same BG, BA, and Row show an increasing trend in REC over time. If both of the above conditions are met, hPPR is performed for permanent repair.
[0109] If the memory fault diagnosis result is a Bank fault:
[0110] like Figure 3 As shown, Bank isolation is performed directly. The Bank isolation process is as follows:
[0111] 1) The BMC sends the Bank isolation parameters to the BIOS, including memory slot information, device information, Bank group (BG), and Bank address (BA).
[0112] 2) During the boot process, the BIOS calculates all system addresses corresponding to the banks that need to be isolated according to the parameters issued by the BMC and the DDR5 memory address mapping rules.
[0113] 3) The BIOS marks the calculated system addresses as reserved. This means that these isolated memory addresses will no longer be allocated to the operating system, thus achieving fault isolation at the cost of sacrificing a small portion of memory capacity.
[0114] For other situations that are neither row faults nor bank faults (such as bank group faults or device faults), the isolation process is terminated directly.
[0115] To further realize the memory fault diagnosis method of the present invention, the present invention also provides specific embodiments of a memory fault diagnosis system and a baseboard management controller therein.
[0116] The baseboard management controller is configured to: first, periodically read the error inspection and correction registers of the memory chips through the basic input / output system and obtain error inspection and correction data; then, parse the error inspection and correction data to extract the Bank group address, Bank address, row address, maximum row error count, and cumulative error count to obtain quintuple data; next, arrange the quintuple data corresponding to each cycle in chronological order to obtain the current quintuple data and historical quintuple data; finally, based on the hierarchical address information in the current quintuple data, perform hierarchical matching between the current quintuple data and the historical quintuple data in order of address granularity from fine to coarse, and determine the memory fault type based on the first successfully matched hierarchical level.
[0117] Specifically, the baseboard management controller includes a data acquisition module, a data parsing module, and a fault diagnosis module.
[0118] The data acquisition module is configured to periodically read the error check and correction registers of the memory chips through the basic input / output system (PIS) and acquire error check and correction data. Specifically, the data acquisition module starts a timed task and sends a call request to the PIS through the intelligent platform management interface; in response to the call request, the PIS reads the error check and correction registers of the corresponding memory chips through the memory controller and returns the read error check and correction data to the baseboard management controller.
[0119] The data parsing module is configured to parse error inspection and correction data, extracting the Bank group address, Bank address, row address, maximum row error count, and cumulative error count to obtain a quintuple of data. Specifically, the data parsing module parses the Bank group address, Bank address, and row address from mode registers 16 to 18 according to the DDR5 specification, parses the maximum row error count from mode register 19, and parses the cumulative error count from mode register 20.
[0120] The fault diagnosis module is configured to arrange the quintuple data corresponding to each cycle in chronological order to obtain the current quintuple data and historical quintuple data. Based on the hierarchical address information in the current quintuple data, it sequentially matches the current quintuple data with the historical quintuple data in order of address granularity from fine to coarse. The fault type of the memory is determined based on the first successfully matched level. The specific judgment logic of the fault diagnosis module has been described in detail in step S140 and will not be repeated here.
[0121] Each module in the aforementioned baseboard management controller can be implemented individually or jointly by a field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), processor-executed firmware or software, etc., and those skilled in the art can flexibly choose according to actual design requirements.
[0122] The memory fault diagnosis method and system, and the baseboard management controller provided in the embodiments of the present invention have at least one or a portion of the following advantages:
[0123] (1) The error inspection and correction register data of the memory chip is periodically collected by the baseboard management controller through the basic input / output system and parsed into a five-tuple data containing the Bank group address, Bank address, row address, maximum row error count and cumulative error count. Then, the fault type is determined by the memory fault diagnosis algorithm. It can realize the active monitoring and identification of DDR5 CE faults without relying on CPU interrupt reporting, thus improving the accuracy of DDR5 CE fault detection.
[0124] (2) By using the hierarchical address information in the current quintuple data as a benchmark, and following the order of address granularity from fine to coarse (row address - Bank address - Bank group address - device), the current quintuple data is matched hierarchically with the historical quintuple data. The first successfully matched level is determined as the fault type. This can diagnose different types of DDR5 CE faults such as row faults, Bank faults, Bank group faults and device faults, thus improving the granularity of fault diagnosis.
[0125] (3) The timed task is started through the baseboard management controller and a call request is sent to the basic input / output system through the intelligent platform management interface. The basic input / output system responds to the call request by reading the error inspection and correction register of the corresponding memory chip through the memory controller and returning the data. The collection period is set to 30 minutes to 24 hours, which realizes low-overhead and periodic collection of fault data and ensures the real-time fault detection without affecting the system's operating performance.
[0126] (4) By parsing the Bank group address, Bank address and row address from mode register 16 to mode register 18 according to the DDR5 specification, parsing the maximum row error count from mode register 19, and parsing the cumulative error count from mode register 20, the standardized parsing of DDR5 ECS register data is realized, providing an accurate data basis for subsequent fault diagnosis.
[0127] (5) After the fault type is diagnosed by the baseboard management controller, the corresponding fault isolation method is selected to perform isolation operation on the faulty memory area in combination with the pre-acquired soft repair capability and / or hard repair capability and the remaining resources. This achieves accurate isolation of the faulty area without replacing the memory, ensuring the continuous and stable operation of the system.
[0128] (6) When the fault type is a row fault, temporary repair is performed by prioritizing soft isolation. If soft isolation has been performed, it is further determined whether the hard isolation conditions are met (the maximum row error count is the maximum value under the same Bank group address and Bank address and is greater than or equal to the preset threshold, and the historical data of the same address shows an increasing trend of the maximum row error count in the time series). If the conditions are met, permanent repair is performed through hard isolation, thus realizing the hierarchical repair strategy for row faults and optimizing the utilization of hard repair resources while ensuring the repair effect.
[0129] (7) When the fault type is Bank fault, the Bank isolation parameters are sent to the basic input / output system through the baseboard management controller. During the power-on startup phase, the basic input / output system calculates all system addresses corresponding to the Bank that needs to be isolated based on the Bank isolation parameters and marks them as reserved, thus achieving Bank-level fault isolation at the cost of sacrificing a small amount of memory capacity.
[0130] (8) The baseboard management controller determines the fault quantity range based on the setting of the maximum row error count in the mode register 19, and sets the corresponding alarm level according to the fault quantity range, thereby realizing the graded early warning of fault severity, which makes it easier for maintenance personnel to take differentiated response measures according to the fault level.
[0131] (9) The risk level of memory chips is classified by the baseboard management controller according to the setting of mode register 20, realizing the graded evaluation of memory chips based on ECS data, which makes it easier for maintenance personnel to intuitively grasp the health status of memory chips and take differentiated maintenance strategies.
[0132] While some embodiments of the present general inventive concept have been shown and described, those skilled in the art will understand that changes may be made to these embodiments without departing from the principles and spirit of the present general inventive concept, the scope of which is defined by the claims and their equivalents.
Claims
1. A method for diagnosing memory faults, characterized in that, The memory fault diagnosis method includes: The baseboard management controller periodically reads the error check and correction register of the memory chip through the basic input / output system and obtains the error check and correction data; The baseboard management controller parses the error inspection and correction data, extracts the Bank group address, Bank address, row address, maximum row error count, and cumulative error count, and obtains quintuple data; The substrate management controller arranges the quintuple data corresponding to each cycle in chronological order to obtain the current quintuple data and historical quintuple data. Based on the hierarchical address information in the current quintuple data, the current quintuple data is matched hierarchically with the historical quintuple data in order of address granularity from fine to coarse, and the memory fault type is determined according to the first successfully matched level.
2. The memory fault diagnosis method according to claim 1, characterized in that, The fault types include row fault types, bank fault types, bank group fault types, and device fault types. The process involves using the hierarchical address information in the current 5-tuple data as a benchmark, and sequentially matching the current 5-tuple data with the historical 5-tuple data in order of address granularity from fine to coarse. The memory fault type is determined based on the first successfully matched level, including: Obtain the value of the maximum row error count in the current quintuple data. If the value of the maximum row error count is less than or equal to 0, determine that there is no memory fault and end the diagnosis. If the maximum row error count is greater than 0, then select all target historical quintuple data with a maximum row error count greater than 0 from the historical quintuple data. If at least one historical quintuple in the target historical quintuple data has the same Bank group address, Bank address, and row address as the current quintuple data, then the memory fault type is determined to be the row fault type. If the target historical quintuple data does not contain any historical quintuple data whose Bank group address, Bank address, and row address are all the same as those in the current quintuple data, but contains at least one historical quintuple data whose Bank group address and Bank address are both the same as those in the current quintuple data, then the memory fault type is determined to be the Bank fault type. If the target historical quintuple data does not contain any historical quintuple data whose Bank group address and Bank address are the same as those in the current quintuple data, but there is historical quintuple data whose Bank group address is the same as those in the current quintuple data, then the memory fault type is determined to be the Bank group fault type. If there is no historical quintuple data in the target historical quintuple data whose Bank group address is the same as the Bank group address in the current quintuple data, then the memory fault type is determined to be the device fault type.
3. The memory fault diagnosis method according to claim 2, characterized in that, The baseboard management controller periodically reads the error check and correction registers of the memory chips through the basic input / output system and obtains error check and correction data, including: The baseboard management controller starts a timed task and sends a call request to the basic input / output system through the intelligent platform management interface; In response to the call request, the basic input / output system reads the error inspection and correction register of the corresponding memory chip through the memory controller, and returns the read error inspection and correction data to the baseboard management controller. The duration of the cycle is from 30 minutes to 24 hours.
4. The memory fault diagnosis method according to claim 2, characterized in that, The error inspection and correction registers include mode registers 15 to 20; The baseboard management controller parses the error inspection and correction data, extracts the Bank group address, Bank address, row address, maximum row error count, and cumulative error count, and obtains a quintuple of data, including: The baseboard management controller parses the Bank group address, Bank address, and row address from mode registers 16 to 18 according to the DDR5 specification, parses the maximum row error count from mode register 19, and parses the cumulative error count from mode register 20 to obtain the quintuple data.
5. The memory fault diagnosis method according to claim 2, characterized in that, Before the baseboard management controller periodically reads the error inspection and correction register of the memory chip through the basic input / output system and obtains the error inspection and correction data, the memory fault diagnosis method further includes: The substrate management controller acquires the soft repair capability and / or hard repair capability of the memory chip and the remaining resource status; The baseboard management controller selects the appropriate fault isolation method based on the fault type, the soft repair capability and / or hard repair capability, and the remaining resources. The baseboard management controller obtains the soft capabilities and / or hard capabilities, as well as the remaining resource status, by reading serial presence detection and / or memory registers.
6. The memory fault diagnosis method according to claim 2, characterized in that, After determining the memory fault type based on the hierarchical address information in the current quintuple data and sequentially matching the current quintuple data with the historical quintuple data in order of address granularity from fine to coarse, the memory fault diagnosis method further includes: The baseboard management controller selects the appropriate fault isolation method based on the fault type, the soft repair capability and / or hard repair capability, and the remaining resources.
7. The memory fault diagnosis method according to claim 6, characterized in that, When the fault type is a row fault type Determine whether the Bank group address corresponding to the row address of the row fault type has been soft isolated. If soft isolation is not implemented, then fault isolation will be achieved through soft isolation. If soft isolation has been implemented, then continue to determine whether the hard isolation conditions are met. If they are met, then perform fault isolation through hard isolation. The hard isolation conditions include: The maximum row error count in the current 5-tuple data is the maximum value under the same Bank group address and Bank address, and the maximum row error count is greater than or equal to a preset threshold. and Historical quintuple data for the same Bank group address, Bank address, and row address show an increasing trend in the maximum row error count over time.
8. The memory fault diagnosis method according to claim 6, characterized in that, When the fault type is a Bank fault type The baseboard management controller sends the Bank isolation parameters to the basic input / output system, wherein the Bank isolation parameters include memory slot information, device information, Bank group address, and Bank address; During the power-on startup phase, the basic input / output system calculates all system addresses corresponding to the banks that need to be isolated based on the bank isolation parameters. The basic input / output system marks the calculated system address as reserved to isolate the Bank.
9. The memory fault diagnosis method according to claim 4, characterized in that, The memory fault diagnosis method also includes a fault level determination step: The baseboard management controller determines the fault quantity range based on the setting of the maximum row error count in the mode register 19, and sets the corresponding alarm level according to the fault quantity range.
10. The memory fault diagnosis method according to claim 9, characterized in that, The memory fault diagnosis method also includes memory chip grading: The baseboard management controller classifies memory chips according to the setting status of mode register 20; When no bit is set in the mode register 20, the memory chip is determined to be at the error-free level. When the low-order bit of the mode register 20 is set and the number of erroneous rows is less than the number of rows in a Bank, the memory chip is determined to be of low risk level. When the middle bit of the mode register 20 is set and the number of erroneous rows is greater than the number of rows in a Bank, the memory chip is determined to be of a high-risk level. When the high-order bit of the mode register 20 is set and the number of erroneous rows exceeds the number of rows in a Bank group, the memory chip is determined to be in a dangerous state.
11. A memory fault diagnosis system, the memory fault diagnosis system comprising a baseboard management controller and a basic input / output system, the memory fault diagnosis system being used to execute the memory fault diagnosis method according to any one of claims 1-10, characterized in that, The substrate management controller is configured to: First, the error detection and correction register of the memory chip is periodically read through the basic input / output system to obtain error detection and correction data; Then, the error inspection and correction data is parsed to extract the Bank group address, Bank address, row address, maximum row error count, and cumulative error count, resulting in quintuple data; Next, the quintuple data corresponding to each period are arranged in chronological order to obtain the current quintuple data and historical quintuple data; Based on the hierarchical address information in the current quintuple data, the current quintuple data is matched hierarchically with the historical quintuple data in order of address granularity from fine to coarse, and the memory fault type is determined according to the first successfully matched level.
12. A baseboard management controller, wherein the baseboard management controller is disposed in the memory fault diagnosis system according to claim 11, characterized in that, The substrate management controller includes: The data acquisition module is configured to periodically read the error check and correction register of the memory chip through the basic input / output system and acquire error check and correction data. The data parsing module is configured to parse the error inspection and correction data, extract the Bank group address, Bank address, row address, maximum row error count and cumulative error count, and obtain quintuple data; The fault diagnosis module is configured to arrange the five-tuple data corresponding to each cycle in chronological order to obtain the current five-tuple data and historical five-tuple data. Based on the hierarchical address information in the current five-tuple data, the module performs hierarchical matching between the current five-tuple data and the historical five-tuple data in order of address granularity from fine to coarse, and determines the memory fault type based on the first successfully matched level.