Memory fault injection test method and electronic device

CN122795698APending Publication Date: 2026-09-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611241066.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]本发明提供了一种内存故障注入测试方法及电子设备,以至少解决相关技术中内存故障注入测试存在的注入失真、覆盖盲目、验证不可靠的技术问题

Benefits of technology

[0006]通过本发明,首先获取待测电子设备的硬件平台信息,根据硬件平台信息加载匹配的故障语义模型,该模型将具有物理成因的内存失效模式描述为包含触发条件、位翻转模式和物理位置约束的参数化模板,使注入的故障与实际硬件失效模式相一致;然后基于加载的故障语义模型和硬件平台信息建立由多个故障维度定义的多维故障空间并初始化覆盖状态矩阵,该覆盖状态矩阵记录了多维故障空间中各个子空间的测试覆盖情况,将抽象的故障空间转化为可量化的结构;再依据覆盖状态矩阵确定与物理位置约束相匹配且覆盖程度符合预设条件的子空间作为下一注入目标并执行注入,能够自动识别测试覆盖不足的区域进行补充注入。这样能够在不修改操作系统或应用程序的前提下,高效、可复现地验证内存子系统的可靠性、可用性与容错能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122795698A_ABST
    Figure CN122795698A_ABST
Patent Text Reader

Abstract

The application discloses a memory fault injection test method and electronic equipment, and relates to the technical field of testing. The method comprises the following steps: obtaining hardware platform information of an electronic equipment to be tested; loading a matched fault semantic model from a preset fault semantic model library according to the hardware platform information, wherein the model describes a memory failure mode with a physical cause as a parameterized template containing a trigger condition, a bit flip mode and a physical location constraint; establishing a multi-dimensional fault space defined by multiple fault dimensions based on the loaded model and the hardware platform information, and initializing a coverage state matrix for recording test coverage states of each subspace; and determining a subspace as an injection target according to the coverage state matrix, wherein the subspace is matched with the physical location constraint of the model and the coverage degree meets preset conditions, and performing injection. In this way, the problems of injection distortion, blind coverage and unreliable verification in traditional memory fault testing are solved, and the authenticity, controllability and sufficiency of the test are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of testing technology, and in particular to a memory fault injection testing method and electronic device. Background Technology

[0002] The reliability of the memory subsystem of an electronic device is directly related to the stable operation of the entire system. To verify the fault tolerance and recovery mechanism of electronic devices when memory hardware failures occur, fault injection testing methods are commonly used. This involves artificially injecting simulated memory errors into the system and observing whether the system's response behavior meets expectations. Traditional memory fault injection methods are mainly divided into two categories: one is implemented through software simulation, which intercepts memory access requests and modifies data bits at the operating system level. This method relies on a software middleware layer, cannot trigger the underlying hardware error handling process, and is affected by the operating system's scheduling, making the injection timing uncontrollable and resulting in significant uncertainty in the test results. The other method is implemented through physical means such as shorting hardware pins. This method can only simulate bit errors, has a single type of fault, and lacks systematic coverage metrics in the testing process, making it difficult to verify the reliability and availability of electronic devices under complex memory fault scenarios. Summary of the Invention

[0003] This invention provides a memory fault injection testing method and electronic device to at least solve the technical problems of injection distortion, blind coverage, and unreliable verification in related technologies.

[0004] This invention provides a memory fault injection testing method, comprising: Obtain the hardware platform information of the electronic device under test; Based on the hardware platform information, a matching fault semantic model is loaded from a pre-set fault semantic model library; the fault semantic model is used to describe memory failure modes with physical causes as parameterized templates containing triggering conditions, bit flip patterns and physical location constraints. Based on the loaded fault semantic model and the hardware platform information, a multi-dimensional fault space defined by multiple fault dimensions is established, and a coverage state matrix is ​​initialized; the coverage state matrix is ​​used to record the test coverage state of each subspace in the multi-dimensional fault space. Based on the coverage state matrix, a subspace that matches the physical location constraints of the loaded fault semantic model and whose coverage meets the preset conditions is determined from the multidimensional fault space as the next injection target, and a memory fault injection operation is performed using the next injection target.

[0005] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described memory fault injection test methods when executing the computer program.

[0006] This invention first acquires the hardware platform information of the electronic device under test. Based on this information, a matching fault semantic model is loaded. This model describes memory failure modes with physical causes as parameterized templates containing triggering conditions, bit flip patterns, and physical location constraints, ensuring the injected faults match the actual hardware failure modes. Then, based on the loaded fault semantic model and hardware platform information, a multi-dimensional fault space defined by multiple fault dimensions is established, and a coverage state matrix is ​​initialized. This matrix records the test coverage of each subspace within the multi-dimensional fault space, transforming the abstract fault space into a quantifiable structure. Next, based on the coverage state matrix, subspaces matching physical location constraints and meeting preset coverage conditions are selected as the next injection targets and injection is performed. This automatically identifies areas with insufficient test coverage and performs supplementary injection. Thus, without modifying the operating system or application, the reliability, availability, and fault tolerance of the memory subsystem can be efficiently and reproducibly verified.

[0007] In addition, the present invention also provides a corresponding electronic device for the memory fault injection test method, which has the same or corresponding technical features as the memory fault injection test method mentioned above, and has the same effect. Attached Figure Description

[0008] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A flowchart of a memory fault injection test method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the memory fault injection testing device provided in an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0011] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0012] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] The specific application environment architecture or specific hardware architecture on which the execution of the memory fault injection test method depends is described here.

[0014] This invention provides a memory fault injection testing method. This method can be applied to various electronic devices under test (DUTs) that possess error handling mechanisms and memory subsystems. The DUTs may include, but are not limited to, servers, switching devices, storage devices, personal computers, and workstations. Servers may be based on Intel architecture, AMD architecture, or be connected to Compute Express Link (CXL) expansion devices. All of the above devices may include a processor, a memory controller, and a memory module, and the processor natively supports a hardware-level fault injection interface or possesses corresponding fault injection capabilities.

[0015] Reliability, availability, and serviceability (RAS) refers to the set of capabilities of an electronic device system to maintain correct operation, recover quickly, and be easy to maintain when hardware failures occur. The method of this invention is used to verify whether the RAS capability of the electronic device under test (DUT) meets design expectations under these hardware failure conditions. This method, by calling the hardware error injection interface natively supported by the DUT, accurately triggers various memory failures without modifying the operating system kernel or application, thereby verifying the fault tolerance, error recovery capability, and serviceability of the DUT under real hardware failure conditions.

[0016] The following section describes the method in detail, taking into account the execution flow of the memory fault injection test method. Figure 1 A flowchart of the memory fault injection test method provided in the embodiments of the present invention is shown below. Figure 1 As shown, the method includes: S101. Obtain the hardware platform information of the electronic device under test.

[0017] It should be noted that the hardware platform information may include the processor type, memory specifications, hardware error handling capabilities, and memory resource distribution of the electronic device under test. This invention uses the electronic device under test as the object of testing.

[0018] When performing step S101, the present invention obtains the hardware platform information of the electronic device under test, which can ensure that the fault injection test method matches the hardware characteristics of the electronic device under test and avoid test failure or result distortion caused by hardware differences.

[0019] S102. Based on the hardware platform information, load the matching fault semantic model from the pre-set fault semantic model library; the fault semantic model is used to describe memory failure modes with physical causes as parameterized templates containing triggering conditions, bit flip patterns and physical location constraints.

[0020] It should be noted that a fault semantic model refers to a computable model built for a given physical failure mechanism, transforming the abstract physical failure process into a programmable parameterized description. Memory failure modes with physical causes can include row hammering, electromigration, uncorrectable errors (UE), and adjacent multi-bit errors. The parameterized template is stored in structured data format. Triggering conditions refer to the access frequency threshold or environmental conditions required to trigger the memory failure mode. For example, row hammering requires a sufficient number of activation operations on the same memory row within a certain time window to cause adjacent row data to flip. Bit flipping modes refer to the specific way data bits flip when a fault occurs, including single-bit flips, double adjacent bit flips, four-bit burst flips, and error correction code check blind zone flips. Physical location constraints refer to the physical address range to which the fault model applies. For example, row hammering only occurs under a set memory row density, and uncorrectable errors apply to specific memory address regions. The fault semantic model can specify a target address proximity constraint distance, and only physical addresses falling within this distance range can trigger the row hammer effect.

[0021] When performing step S102, based on the hardware platform information obtained in step S101, a fault semantic model matching the hardware platform information is selected from a pre-set fault semantic model library. By loading a fault semantic model matching the hardware platform information, it can be ensured that the injected fault mode is highly consistent with the physical failure modes that may actually occur in the hardware, thereby improving the reliability of the test results.

[0022] S103. Based on the loaded fault semantic model and hardware platform information, establish a multi-dimensional fault space defined by multiple fault dimensions, and initialize the coverage state matrix; the coverage state matrix is ​​used to record the test coverage state of each subspace in the multi-dimensional fault space.

[0023] It should be noted that fault dimensions refer to the independent variables that constitute the fault space, with each dimension representing an attribute aspect of the fault. The coverage state matrix is ​​a data structure used to record whether each subspace in the multidimensional fault space has been tested and covered; its initial state is that all subspaces are uncovered.

[0024] During step S103, the loaded fault semantic model can provide fault parameters such as physical location constraints, and the hardware platform information can provide hardware parameters such as memory node distribution and address range. The loaded fault semantic model and hardware platform information are used together as input to construct a multi-dimensional fault space. This multi-dimensional fault space can include multiple dimensions such as address distribution, time window, bit error mode, error type, and non-uniform memory access nodes. By combining and operating on these multiple dimensions, a discrete set of subspaces is generated, each representing a specific combination of test scenarios. The coverage state matrix is ​​a data structure corresponding to this multi-dimensional fault space, initially set to completely uncovered, and used to record whether each subspace has been tested. By establishing the multi-dimensional fault space and initializing the coverage state matrix, the originally infinite memory fault test scenarios can be transformed into a finite set of computable subspaces, making test sufficiency measurable and verifiable.

[0025] S104. Based on the coverage state matrix, determine the subspace from the multidimensional fault space that matches the physical location constraints of the loaded fault semantic model and whose coverage meets the preset conditions as the next injection target, and use the next injection target to perform memory fault injection operation.

[0026] It should be noted that the injection target refers to the target subspace selected from the multidimensional fault space for performing the next fault injection operation, containing all parameters required for injection, such as physical address, error type, and bit-flipping method. The aforementioned preset conditions may include: the subspace is recorded as an uncovered state in the coverage state matrix, and the subspace matches the loaded fault semantic model in terms of bit-flipping mask conditions.

[0027] In step S104, based on the test coverage of each subspace recorded in the coverage state matrix, the present invention filters out subspaces from the multidimensional fault space that match the physical location constraints of the loaded fault semantic model, and selects subspaces whose coverage meets preset conditions as the next injection target. Then, the memory fault injection operation is performed using the injection target. In this way, by selectively injecting into areas with insufficient coverage in the established fault space, test resources can be focused on fault scenarios that have not yet been fully verified, avoiding repeated testing of covered areas and solving the problems of blind coverage and wasted test resources in traditional fault injection testing.

[0028] In the memory fault injection testing method provided in this embodiment of the invention, the hardware platform information of the electronic device under test is first obtained. A matching fault semantic model is then loaded based on the hardware platform information. This model describes memory failure modes with physical causes as parameterized templates containing triggering conditions, bit flip patterns, and physical location constraints, ensuring that the injected fault matches the actual hardware failure mode. Next, a multi-dimensional fault space defined by multiple fault dimensions is established based on the loaded fault semantic model and hardware platform information, and a coverage state matrix is ​​initialized. This coverage state matrix records the test coverage of each subspace in the multi-dimensional fault space, transforming the abstract fault space into a quantifiable structure. Then, based on the coverage state matrix, subspaces matching physical location constraints and meeting preset coverage conditions are selected as the next injection target and injection is performed. This method can automatically identify areas with insufficient test coverage and perform supplementary injection. This effectively improves the authenticity, controllability, and sufficiency of memory fault injection testing, enabling efficient and reproducible verification of the reliability, availability, and fault tolerance of the memory subsystem without modifying the operating system or application.

[0029] It should be noted that the executing entity of this invention can be a test control device, which can be deployed inside the electronic device under test (DUT) or set up independently of the DUT. When deployed inside the DUT, the test control device can be a test agent module in a controller, such as the test agent module in a Baseboard Management Controller (BMC); when set up independently of the DUT, the test control device can be an external test device that is communicatively connected to the DUT.

[0030] Furthermore, in a specific implementation, in the memory fault injection test method provided in the embodiments of the present invention, step S101, obtaining the hardware platform information of the electronic device under test, may specifically include: obtaining the processor type (such as Intel or AMD) identification information of the electronic device under test; reading the memory controller configuration information of the electronic device under test to determine the memory parameters of the electronic device under test; the memory parameters include the memory generation type and the error correction function enable status; querying the error handling capability status information of the electronic device under test to determine the error handling function support status of the electronic device under test; the error handling function support status includes the machine check error recovery support status, the recoverable address range recovery support status, and the CXL error injection support status; parsing the Advanced Configuration and Power Interface (ACPI) system resource affinity information of the electronic device under test, and extracting the number of Non-Uniform Memory Access (NUMA) nodes and the physical address range corresponding to each node of the electronic device under test.

[0031] In implementation, processor type identification information can be obtained by executing processor identification instructions (such as the Central Processing Unit Identifier, CPUID). This instruction returns the processor's manufacturer identifier and model code, used to distinguish different types of processors, such as Intel and AMD architectures. Memory controller configuration information can be obtained by reading the configuration registers in the memory controller to identify whether the current memory type is Double Data Rate 4 (DDR4) or Double Data Rate 5 (DDR5), and to determine whether Error Correcting Code (ECC), multi-chip error correction protection mechanisms, or Single Device Data Correction (SDDC) capabilities are enabled. Error handling capability status information can be obtained by querying the error handling capability status register to confirm whether the electronic device under test supports machine-checked error recovery mechanisms, recoverable address range mechanisms, or CXL error injection capabilities. ACPI system resource affinity information is a system resource description table provided by the electronic device firmware. Parsing this information can extract the number of NUMA nodes and the physical address range corresponding to each node. This allows for the effective acquisition of the hardware configuration parameters of the electronic device under test, ensuring that the testing method can adaptively adapt to different electronic device hardware platforms.

[0032] Furthermore, in a specific implementation, in the memory fault injection test method provided in the embodiments of the present invention, step S102 loads a matching fault semantic model from a preset fault semantic model library based on hardware platform information. Specifically, this may include: selecting a fault semantic model corresponding to the processor type identification information and memory parameters from the preset fault semantic model library based on the processor type identification information and memory parameters of the electronic device under test; and loading the selected fault semantic model into the runtime injection strategy engine as the parameter configuration basis for the injection operation.

[0033] In implementation, the fault semantic model library pre-loads various fault semantic models for different hardware platforms and physical failure modes. For example, when the processor is Intel architecture and the memory type is DDR4, the row hammer effect model and the single-bit uncorrectable error model are loaded; when the processor is AMD architecture, the memory type is DDR5, and a CXL expansion device is connected, the CXL cache consistency violation model and the persistent memory write failure model are loaded. Each fault semantic model is stored in structured data form and may include the minimum trigger access frequency threshold, the target address proximity constraint distance, the bit flip location mask, and the applicable operating temperature range. The runtime injection strategy engine is a software module deployed in the electronic device under test, responsible for receiving the fault semantic model and driving the execution of the injection operation according to the model parameters. After loading the selected fault semantic model into the runtime injection strategy engine, the injection strategy engine can generate specific injection instructions based on the trigger conditions, bit flip modes, and physical location constraints in the model. This enables the migration and unified execution of test methods and test cases across different hardware platforms without the need to develop separate test scripts for each hardware platform.

[0034] In practical applications, the fault semantic model library can be expanded and updated according to the hardware platform and physical failure mode. For example, when a new generation of memory technology or a new failure mechanism is discovered, the failure mode can be abstracted into a new fault semantic model based on the parameterized template definition method of this invention and added to the model library, thereby maintaining the adaptability of the test method to the hardware platform.

[0035] Furthermore, in specific implementation, in the memory fault injection test method provided in the embodiments of the present invention, step S103 establishes a multi-dimensional fault space defined by multiple fault dimensions and initializes the coverage state matrix. Specifically, this may include: dividing the physical address space into multiple address partitions according to its purpose to form an address distribution dimension; dividing the system operating cycle into multiple time stages according to load characteristics to form a time window dimension; dividing data bit errors into multiple bit error modes according to flip width and position characteristics to form a data bit error mode dimension; dividing errors into multiple error types according to severity and system impact to form an error type dimension; using the memory access node number as the NUMA node dimension; performing a Cartesian product operation on the address distribution dimension, time window dimension, data bit error mode dimension, error type dimension, and NUMA node dimension to generate a discrete set of subspaces as the multi-dimensional fault space; and initializing each item in the coverage state matrix corresponding to each subspace to an uncovered state.

[0036] In implementation, this invention divides the physical address space into multiple address partitions according to their purpose, forming an address distribution dimension. These address partitions can include kernel code areas, user stack areas, memory-mapped input / output (MMIO) areas, and reserved unused areas. The system impact of memory faults differs across address regions, requiring separate testing coverage. The system runtime cycle is divided into multiple time stages based on load characteristics, forming a time window dimension. These time stages can include startup initialization, steady-state operation, high-load pressure, and pre-shutdown cleanup. The memory controller's state differs in different runtime stages, potentially leading to variations in fault detection and recovery behavior. Data bit errors are categorized into various bit error patterns based on flip width and location characteristics, forming a data bit error pattern dimension. These bit error patterns can include single-bit flips, double-adjacent-bit flips, four-bit burst flips, and error correction code check blind zone flips. Different flip patterns have varying degrees of impact on data integrity. Errors are further categorized into various error types based on severity and system impact, forming an error type dimension. This error type can include correctable errors, uncorrectable errors, silent data corruption, cache consistency errors, etc. Different types of errors trigger different system response mechanisms. NUMA node numbers are used as an independent dimension. The NUMA node dimension is independent of the other four dimensions and is used to distinguish local memory and remote memory on different processor nodes.

[0037] This invention can perform Cartesian product operations on the above five dimensions to generate a discrete set of subspaces, which constitute a multi-dimensional fault space. For example, there are 4 types of address partitions, 4 types of time stages, 4 types of bit error modes, 4 types of error types, and N types of NUMA nodes, resulting in a total of 4×4×4×4×N subspaces. Each subspace represents a specific combination of test scenarios; for example, injecting a correctable error of the single-bit flip type in the kernel code section of node 0, during steady-state operation. The items corresponding to each subspace in the coverage state matrix are initialized to an uncovered state, indicating that all test scenarios have not yet been executed. After each injection test is completed, the corresponding subspace is marked as covered. This systematically organizes and manages a large set of test scenarios, making the test coverage state recordable, traceable, and measurable.

[0038] Furthermore, in a specific implementation, in the memory fault injection test method provided in the embodiments of the present invention, step S104, based on the coverage state matrix, determines the subspace from the multidimensional fault space that matches the physical location constraints of the loaded fault semantic model and whose coverage meets the preset conditions as the next injection target. Specifically, this may include: selecting subspaces from the multidimensional fault space that match the physical location constraints of the loaded fault semantic model as candidate subspaces; counting the historical selection count of each candidate subspace; calculating the exploration weight of each candidate subspace; the exploration weight decays exponentially with the increase of the historical selection count; selecting the candidate subspace with the largest exploration weight as the target subspace from the candidate subspaces; extracting physical address parameters, time stage identifiers, bit flip masks, error type parameters, and NUMA node numbers from the target subspace, and combining them to generate the injection parameter set of the next injection target.

[0039] In implementation, this invention first filters out subspaces that have not yet been tested from the coverage state matrix. Then, based on the physical location constraints of the loaded fault semantic model, it filters out subspaces that match the physical location constraints and satisfy the bit-flip mask condition as candidate subspaces. Next, it calculates the exploration weight based on the historical selection count of each candidate subspace and selects the candidate subspace with the highest exploration weight as the target subspace. The historical selection count is used to record the number of times each subspace has been selected as an injection target in previous test iterations, with an initial value of zero. The fewer the historical selection counts, the higher the exploration weight. The exploration weight decays exponentially with the increase in the historical selection count; that is, the more times a subspace is selected, the lower its exploration weight, and vice versa. This ensures that subspaces with lower coverage are selected first each time, achieving optimized allocation of test resources.

[0040] Next, the physical address parameters, time phase identifier, bit-flip mask, error type parameters, and NUMA node number extracted from the target subspace constitute a complete injection parameter set. The physical address parameter specifies the target memory address for the injection operation, the time phase identifier specifies the runtime phase in which the injection is performed, the bit-flip mask specifies the specific method of data bit flipping, the error type parameter specifies the severity of the injection error, and the NUMA node number specifies the processor node where the target address is located. This automatically focuses test resources on areas that are not yet fully covered, avoiding repeated testing of covered areas and accelerating coverage convergence.

[0041] Furthermore, in a specific implementation, in the memory fault injection test method provided in the embodiments of the present invention, step S104 performs a memory fault injection operation using the next injection target, which may specifically include: performing a memory fault injection operation at a specified physical address according to the next injection target through the native error injection interface of the electronic device under test; wherein, the native error injection interface is an interface provided by the hardware platform of the electronic device under test itself for triggering memory errors.

[0042] In implementation, the native error injection interface does not rely on the cooperation of the operating system or application. Instead, it can directly manipulate the error injection control register in the processor or the error injection configuration space in the memory controller to trigger a memory error of a specified type at a specified physical address. The errors triggered by this invention through the native error injection interface are directly generated by the hardware, fully activating the processor's machine error handling process, the memory controller's error correction and retry mechanism, and the platform firmware's error recording and recovery process. This verifies the fault-tolerant behavior of the electronic device under test under real hardware error conditions. During the injection operation, the physical address parameter in the target subspace can be converted to a physical address, the error type parameter can be converted to the corresponding error type code in the hardware register, and the bit-flip mask can be converted to an error injection mask in the register. Then, the error is triggered by writing to the corresponding configuration register.

[0043] Furthermore, in specific implementation, in the above steps, the memory fault injection operation is performed at the specified physical address according to the next injection target through the native error injection interface of the electronic device under test. Specifically, this may include: when the processor of the electronic device under test supports machine-checked error injection, error injection is triggered through machine-checked error injection according to the error type parameters and physical address parameters contained in the next injection target; when the processor of the electronic device under test supports recoverable address range error injection, error injection is triggered through recoverable address range error injection according to the physical address parameters and error type parameters contained in the next injection target; when the electronic device under test is connected to a CXL expansion device, error injection is triggered through the error injection method of the CXL expansion device according to the physical address parameters, bit flip mask, and error type parameters contained in the next injection target; after injection, a cache refresh operation is performed, and a data read operation at the specified physical address is performed to trigger the memory controller's error handling process.

[0044] In implementation, the machine-checked error injection method utilizes the processor's machine-checked exception (MCE) mechanism. MCEs are exceptions triggered when the processor detects a hardware error (such as a memory failure). When the exception is triggered, the processor records the error information to a machine-checked register. The operating system can then handle the exception to perform error recovery or system termination. For example, by writing the error type code and uncorrectable error flag to the corresponding error injection register (e.g., IA32_MCi_STATUS) in the processor, and writing the target physical address to the address register (e.g., IA32_MCi_ADDR), and then setting the valid bit (VALID) to activate the error injection, the processor triggers an error at the specified address according to the written parameters and generates the corresponding machine-checked exception. This method is suitable for processors that support machine-checked error recovery mechanisms (such as Intel architecture processors).

[0045] The Scalable Reliable Address Range (SRAR) error injection method utilizes the processor's SRAR mechanism. This method allows a specified physical address range to be marked as a recoverable region. When a memory error occurs within this region, error recovery operations can be performed without terminating the entire system. For example, an error can be triggered by configuring the address matching register to the target physical address, setting the error attribute register to the specified error type, and setting the injection enable bit. This method is suitable for processors that support the SRAR mechanism (such as AMD architecture processors).

[0046] The CXL extended device error injection method triggers an error by writing an error descriptor containing the target physical address, bit flip mask, and error type into the error injection control register of the CXL extended device.

[0047] After injection, a cache flush instruction (such as CLFLUSH) can be executed first to flush the cache line containing the specified physical address from the processor cache. Then, a memory read instruction (such as MOV) can be executed to read the data at that physical address. By forcibly triggering memory access, the error detection and error handling process of the memory controller can be activated. This can cover electronic device platforms with different processor architectures and different memory configurations, achieving a unified cross-platform fault injection capability.

[0048] Furthermore, in specific implementation, the memory fault injection test method provided in the embodiments of the present invention may further include: collecting response logs generated by the electronic device under test after the memory fault injection operation; parsing the response logs to obtain the actual triggered fault feature information; mapping the fault feature information to the corresponding subspace in the multi-dimensional fault space, and updating the corresponding item in the coverage state matrix to the covered state; determining whether the test termination condition is met based on the updated coverage state matrix; if met, generating a test report containing covered subspace information, uncovered subspace information, and system response behavior information.

[0049] In implementation, the system first collects response logs generated after the injection operation to obtain raw feedback information about the error. Then, the logs are parsed to extract structured fault characteristic information from the unstructured log text, which may include key fields such as error type, error address, and whether it has been corrected. Next, the extracted fault characteristic information is mapped to a multi-dimensional fault space. Based on fields such as physical address, timestamp, and error type, the corresponding subspace is located, and the status of that subspace in the coverage state matrix is ​​updated from uncovered to covered, completing the synchronization of coverage records. Finally, the current coverage is calculated based on the updated coverage state matrix to determine if the preset test termination conditions are met. If met, a test report is generated. This report may specifically include identification information of the covered subspace set, identification information of uncovered high-risk areas, and classification information of the response behavior of the electronic device under test to various types of injected faults. This process automates the testing process, completing the entire test cycle from injection to evaluation without manual intervention, effectively improving testing efficiency.

[0050] The steps described above, including collecting the response logs generated by the electronic device under test (DUT) after the memory fault injection operation, may specifically include: extracting error physical address information, error type identification information, and information on whether the error correction mechanism has corrected the error from the processor exception log of the DUT; extracting the physical location information of the memory module where the error occurred, the cumulative error count information, and information on whether memory page isolation operations were performed from the management controller log of the DUT; extracting the cumulative count values ​​of correctable and uncorrectable errors from the error counting interface of the operating system kernel; and aligning the information extracted from the processor exception log, management controller log, and error counting interface by timestamp and fusing them to generate fault feature information.

[0051] In implementation, response log collection involves multiple data sources, each corresponding to a different level of hardware error handling in the electronic device under test. The processor exception log records the exception information directly generated by the processor when a hardware error is detected. This log may include the physical address of the error, the error type, and the correction results of the error correction mechanism. This log reflects the processor and memory controller's initial perception and preliminary handling of the error. The management controller log records the monitoring results of hardware events by the baseboard management controller as an out-of-band management system. This includes the physical location of the faulty memory module (e.g., memory slot number), the cumulative number of errors, and whether the system performed recovery actions such as memory page isolation. This log reflects the platform management firmware's response and recovery operations to the error. The operating system kernel's error counting interface records the cumulative statistics of correctable and uncorrectable errors. This interface provides global error counting information to the upper layer. These three types of logs are then fused after being aligned by timestamps to generate a fault feature vector. This vector can contain complete information such as the location, type, severity of the fault, and the system's response actions at different levels. This effectively obtains the system's response behavior to injected faults, avoiding misjudgments caused by incomplete information from a single data source.

[0052] In addition, the above steps, determining whether the test termination condition is met based on the updated coverage state matrix, can specifically include: calculating the coverage of the current multidimensional fault space based on the updated coverage state matrix; the coverage is composed of a weighted sum of a coverage ratio term and a coverage diversity term, and the coverage diversity term is calculated based on the distribution information entropy of the covered subspace; if the increase in coverage is less than a preset tolerance threshold in multiple consecutive rounds of testing, or if the coverage reaches a preset convergence threshold, then the test termination condition is determined to be met.

[0053] In implementation, coverage can be weighted by a coverage ratio and a coverage diversity term. The coverage diversity term is calculated based on the distribution information entropy of the covered subspaces, while the coverage ratio term is the proportion of the number of covered subspaces to the total number of subspaces. The coverage increment refers to the increase in coverage between the current test round and the previous test round. Preset tolerance thresholds can be set by the user based on actual testing needs or determined based on experience. For example, coverage convergence can be determined when the coverage increment is less than 1% in five consecutive test rounds, or when the coverage reaches a preset convergence threshold (e.g., 95%), the termination condition can be met. Through quantifiable coverage metrics and convergence determination mechanisms, objective and quantifiable decision-making basis for test termination is provided, avoiding the problem of relying on manual experience to judge the adequacy of testing in traditional testing methods.

[0054] From the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0055] Embodiments of the present invention also provide a memory fault injection testing device. Figure 2 This is a schematic diagram of the memory fault injection testing device provided in an embodiment of the present invention. This embodiment is based on the perspective of functional modules, such as… Figure 2 As shown, the device includes: Platform information acquisition module 10 is used to acquire hardware platform information of the electronic device under test; The fault semantic modeling module 11 is used to load a matching fault semantic model from a pre-set fault semantic model library based on hardware platform information; the fault semantic model is used to describe memory failure modes with physical causes as parameterized templates containing triggering conditions, bit flip patterns and physical location constraints. The fault space construction module 12 is used to establish a multi-dimensional fault space defined by multiple fault dimensions based on the loaded fault semantic model and hardware platform information, and initialize the coverage state matrix; the coverage state matrix is ​​used to record the test coverage state of each subspace in the multi-dimensional fault space. The injection target decision module 13 is used to determine the next injection target from the multidimensional fault space based on the coverage state matrix. The subspace that matches the physical location constraints of the loaded fault semantic model and whose coverage meets the preset conditions is used to perform memory fault injection operation.

[0056] In the memory fault injection testing device provided in this embodiment of the invention, the platform information acquisition module 10 acquires the hardware platform information of the electronic device under test, and the fault semantic modeling module 11 loads a matching fault semantic model. This model describes memory failure modes with physical causes as parameterized templates containing triggering conditions, bit flip modes, and physical location constraints, making the injected fault consistent with the actual hardware failure mode. The fault space construction module 12 establishes a multi-dimensional fault space and initializes the coverage state matrix, transforming the abstract fault space into a quantifiable structure. The injection target decision module 13 determines the subspace that matches the physical location constraints and whose coverage meets the preset conditions as the next injection target based on the coverage state matrix and performs the injection, thus realizing adaptive testing. This device effectively improves the authenticity, controllability, and sufficiency of memory fault injection testing, and can efficiently and reproducibly verify the reliability, availability, and fault tolerance of the memory subsystem without modifying the operating system or application.

[0057] Since the embodiments of the memory fault injection testing device and the memory fault injection testing method correspond to each other, the descriptions of the features in the embodiments corresponding to the memory fault injection testing device can be found in the relevant descriptions of the embodiments corresponding to the memory fault injection testing method, and will not be repeated here. Furthermore, it has the same beneficial effects as the memory fault injection testing method mentioned above.

[0058] Furthermore, in a specific implementation, in the memory fault injection test device provided in the embodiments of the present invention, the platform information acquisition module 10 can be specifically used to acquire the processor type identification information of the electronic device under test; read the memory controller configuration information of the electronic device under test to determine the memory parameters of the electronic device under test; the memory parameters include the memory generation type and the error correction function enable status; query the error handling capability status information of the electronic device under test to determine the error handling function support status of the electronic device under test; the error handling function support status includes the machine check error recovery support status, the recoverable address range recovery support status, and the CXL error injection support status; parse the ACPI system resource affinity information of the electronic device under test, and extract the number of NUMA nodes of the electronic device under test and the physical address range corresponding to each node.

[0059] Furthermore, in a specific implementation, in the memory fault injection test device provided in the embodiments of the present invention, the fault semantic modeling module 11 can be specifically used to select a fault semantic model corresponding to the processor type identification information and memory parameters from a preset fault semantic model library according to the processor type identification information and memory parameters of the electronic device under test; and load the selected fault semantic model into the runtime injection strategy engine as the parameter configuration basis for the injection operation.

[0060] Furthermore, in specific implementation, in the memory fault injection testing device provided in the embodiments of the present invention, the fault space construction module 12 can be specifically used to divide the physical address space into multiple address partitions according to its purpose, forming an address distribution dimension; divide the system operation cycle into multiple time stages according to load characteristics, forming a time window dimension; divide data bit errors into multiple bit error modes according to flip width and position characteristics, forming a data bit error mode dimension; divide errors into multiple error types according to severity and system impact, forming an error type dimension; use the memory access node number as the NUMA node dimension; perform Cartesian product operation on the address distribution dimension, time window dimension, data bit error mode dimension, error type dimension, and NUMA node dimension to generate a discrete set of subspaces as a multidimensional fault space; and initialize each item in the coverage state matrix corresponding to each subspace to an uncovered state.

[0061] Furthermore, in specific implementation, in the memory fault injection testing device provided in the embodiments of the present invention, the injection target decision module 13 can be specifically used to select subspaces that match the physical location constraints from the multi-dimensional fault space according to the physical location constraints of the loaded fault semantic model, as candidate subspaces; count the number of times each candidate subspace has been selected in history; calculate the exploration weight of each candidate subspace; the exploration weight decreases exponentially with the increase of the number of times it has been selected in history; select the candidate subspace with the largest exploration weight from the candidate subspaces as the target subspace; extract physical address parameters, time stage identifiers, bit flip masks, error type parameters and NUMA node numbers from the target subspace, and combine them to generate the injection parameter set of the next injection target; and perform memory fault injection operation at the specified physical address according to the next injection target through the native error injection interface of the electronic device under test; wherein, the native error injection interface is an interface provided by the hardware platform of the electronic device under test itself for triggering memory errors.

[0062] When the processor of the electronic device under test supports machine-checked error injection, error injection is triggered using the machine-checked error injection method according to the error type parameters and physical address parameters contained in the next injection target. When the processor of the electronic device under test supports recoverable address range error injection, error injection is triggered using the recoverable address range error injection method according to the physical address parameters and error type parameters contained in the next injection target. When the electronic device under test is connected to a CXL expansion device, error injection is triggered using the error injection method of the CXL expansion device, according to the physical address parameters, bit flip mask, and error type parameters contained in the next injection target. After injection, a cache refresh operation is performed, and a data read operation at the specified physical address is performed to trigger the error handling process of the memory controller.

[0063] Furthermore, in a specific implementation, the memory fault injection testing device provided in the embodiments of the present invention may further include: a test report generation module, used to collect response logs generated by the electronic device under test after a memory fault injection operation; parse the response logs to obtain the actual triggered fault feature information; map the fault feature information to the corresponding subspace in the multi-dimensional fault space, and update the corresponding item in the coverage state matrix to the covered state; determine whether the test termination condition is met based on the updated coverage state matrix; if met, generate a test report containing covered subspace information, uncovered subspace information, and system response behavior information.

[0064] Specifically, the coverage of the current multidimensional fault space is calculated based on the updated coverage state matrix. The coverage is composed of a weighted sum of the coverage ratio and the coverage diversity, with the coverage diversity calculated based on the distribution information entropy of the covered subspace. If the increase in coverage is less than the preset tolerance threshold in multiple consecutive rounds of testing, or if the coverage reaches the preset convergence threshold, the test termination condition is determined to be met.

[0065] Embodiments of the present invention also provide an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the memory fault injection test method. The present invention can use this electronic device as a hardware carrier for performing the test method.

[0066] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the memory fault injection test method when run.

[0067] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0068] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the memory fault injection test method.

[0069] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described memory fault injection test method embodiments.

[0070] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be performed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microcontroller unit (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.

[0071] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0072] The present invention has provided a detailed description of a memory fault injection testing method and electronic device. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are only intended to aid in understanding the method and core ideas of the invention. It should be noted that those skilled in the art can make various improvements and modifications to the invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A memory fault injection testing method, characterized in that, include: Obtain the hardware platform information of the electronic device under test; Based on the hardware platform information, load the matching fault semantic model from the pre-set fault semantic model library; The fault semantic model is used to describe memory failure modes with physical causes as parameterized templates that include triggering conditions, bit flip patterns, and physical location constraints. Based on the loaded fault semantic model and the hardware platform information, a multi-dimensional fault space defined by multiple fault dimensions is established, and a coverage state matrix is ​​initialized; the coverage state matrix is ​​used to record the test coverage state of each subspace in the multi-dimensional fault space. Based on the coverage state matrix, a subspace that matches the physical location constraints of the loaded fault semantic model and whose coverage meets the preset conditions is determined from the multidimensional fault space as the next injection target, and a memory fault injection operation is performed using the next injection target.

2. The memory fault injection testing method according to claim 1, characterized in that, Obtain the hardware platform information of the electronic device under test, including: Obtain the processor type identification information of the electronic device under test; The memory controller configuration information of the electronic device under test is read to determine the memory parameters of the electronic device under test; the memory parameters include the memory generation type and the error correction function enable status; The error handling capability status information of the electronic device under test is queried to determine the error handling function support status of the electronic device under test; the error handling function support status includes machine check error recovery support status, recoverable address range recovery support status, and high-speed interconnect protocol error injection support status. The system resource affinity information of the advanced configuration and power management interface of the electronic device under test is analyzed to extract the number of non-uniform memory access nodes and the physical address range corresponding to each node.

3. The memory fault injection testing method according to claim 1, characterized in that, Based on the hardware platform information, a matching fault semantic model is loaded from a pre-set fault semantic model library, including: Based on the processor type identification information and memory parameters of the electronic device under test, a fault semantic model corresponding to the processor type identification information and memory parameters is selected from a preset fault semantic model library; The selected fault semantic model is loaded into the runtime injection strategy engine to serve as the basis for parameter configuration of the injection operation.

4. The memory fault injection testing method according to claim 1, characterized in that, Establish a multidimensional fault space defined by multiple fault dimensions, and initialize the coverage state matrix, including: The physical address space is divided into multiple address partitions according to its purpose, forming the address distribution dimension; The system operating cycle is divided into multiple time stages according to load characteristics, forming a time window dimension; Data bit errors are classified into multiple bit error modes according to their flip width and position characteristics, forming a data bit error mode dimension; Errors are categorized into multiple error types based on their severity and system impact, forming an error type dimension. Use the memory access node number as a non-uniform memory access node dimension; Cartesian product operations are performed on the address distribution dimension, the time window dimension, the data bit error mode dimension, the error type dimension, and the non-uniform memory access node dimension to generate a discrete set of subspaces, which serve as a multidimensional fault space. Initialize each item in the covered state matrix corresponding to each subspace to an uncovered state.

5. The memory fault injection testing method according to claim 1, characterized in that, Based on the coverage state matrix, a subspace from the multidimensional fault space that matches the physical location constraints of the loaded fault semantic model and whose coverage meets preset conditions is determined as the next injection target, including: Based on the physical location constraints of the loaded fault semantic model, a subspace that matches the physical location constraints is selected from the multidimensional fault space as a candidate subspace. Count the number of times each candidate subspace has been selected in history; Calculate the exploration weight for each candidate subspace; the exploration weight decays exponentially with the number of times the history has been selected; Select the candidate subspace with the largest exploration weight from the candidate subspaces as the target subspace; The physical address parameters, time stage identifier, bit flip mask, error type parameters, and non-uniform memory access node number are extracted from the target subspace and combined to generate the injection parameter set for the next injection target.

6. The memory fault injection testing method according to claim 1, characterized in that, Performing a memory fault injection operation using the next injection target includes: Through the native error injection interface of the electronic device under test, a memory fault injection operation is performed at the specified physical address according to the next injection target; The native error injection interface is an interface provided by the hardware platform of the electronic device under test, used to trigger memory errors.

7. The memory fault injection test method according to claim 6, characterized in that, Through the native error injection interface of the electronic device under test, a memory fault injection operation is performed at the specified physical address according to the next injection target, including: When the processor of the electronic device under test supports machine-checked error injection, error injection is triggered by the machine-checked error injection method according to the error type parameters and physical address parameters contained in the next injection target. When the processor of the electronic device under test supports the recoverable address range error injection method, error injection is triggered by the recoverable address range error injection method according to the physical address parameters and error type parameters contained in the next injection target. When the electronic device under test is connected to a high-speed interconnect protocol extension device, error injection is triggered by the error injection method of the high-speed interconnect protocol extension device according to the physical address parameters, bit flip mask and error type parameters contained in the next injection target; After the injection is complete, a cache refresh operation is performed, and a data read operation at the specified physical address is performed to trigger the memory controller's error handling process.

8. The memory fault injection test method according to claim 6, characterized in that, Also includes: Collect the response logs generated by the electronic device under test after the memory fault injection operation; The response log is parsed to obtain the actual fault characteristic information that was triggered. The fault feature information is mapped to the corresponding subspace in the multidimensional fault space, and the corresponding item in the coverage state matrix is ​​updated to the covered state. The updated coverage status matrix is ​​used to determine whether the test termination condition is met; if it is met, a test report containing information on covered subspaces, uncovered subspaces, and system response behavior is generated.

9. The memory fault injection test method according to claim 8, characterized in that, Determine whether the test termination condition is met based on the updated coverage state matrix, including: Based on the updated coverage state matrix, the coverage of the current multidimensional fault space is calculated; the coverage is composed of a weighted sum of a coverage ratio term and a coverage diversity term, and the coverage diversity term is calculated based on the distribution information entropy of the covered subspaces. If the increase in coverage is less than a preset tolerance threshold in multiple consecutive rounds of testing, or if the coverage reaches a preset convergence threshold, then the test termination condition is determined to be met.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the memory fault injection test method as described in any one of claims 1 to 9 when executing the computer program.