A method for testing a memory cell
Patent Information
- Application Number
- CN202512007957.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-12-29
AI Technical Summary
[0004]本发明提供一种内存颗粒的测试方法,能够解决现有32位的测试系统只能测试32位的内存颗粒,若需要测试16位的内存颗粒,则需要重新设计相应的16位的测试系统,这样一来会增加测试成本也会降低测试效率的技术问题
[0006]The beneficial effects of this invention are as follows: By employing methods such as channel shielding, code modification, and address mapping, a 32-bit testing system achieves compatibility with two types of memory chips. When testing 32-bit chips, a dual-channel path is enabled; when testing 16-bit chips, stability is ensured through single-channel adaptation and three-layer testing. This compatibility allows a single testing system to cover both types of memory chips without modification, avoiding the need for separate development of testing systems due to different chip types, and reducing equipment procurement and maintenance costs. The solution supports the special scenario of "second-channel access" for 16-bit chips. By adjusting channel mapping and testing procedures, it can adapt to the requirement of "temporarily enabling the second channel when the first channel fails." Simultaneously, the combination of 16MB block segmentation testing and the three-layer testing method can be applied to memory chips of different capacities without adjusting the core logic due to changes in memory capacity, further enhancing the universality and practicality of the solution.
Smart Images

Figure CN121415853B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage technology, and in particular to a testing method for memory chips. Background Technology
[0002] Memory storage (such as DDR) is a common type of memory. Before being officially delivered, memory storage devices need to undergo read / write testing, which often involves covering the data at each bit level. Memory chips are the core components of memory storage devices, and their performance and quality directly determine the capacity, speed, and stability of the entire memory module. Testing of memory storage devices includes testing of the memory chips.
[0003] The existing 32-bit testing system can only test 32-bit memory chips. If it is necessary to test 16-bit memory chips, a corresponding 16-bit testing system needs to be redesigned, which will increase testing costs and reduce testing efficiency. Summary of the Invention
[0004] This invention provides a testing method for memory chips, which solves the technical problem that existing 32-bit testing systems can only test 32-bit memory chips. If 16-bit memory chips need to be tested, a corresponding 16-bit testing system needs to be redesigned, which increases testing costs and reduces testing efficiency.
[0005] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: to provide a testing method for memory chips, the method comprising: Since 32-bit memory chips involve a first channel (16-bit) and a second channel (16-bit), the first channel (16-bit) and the second channel (16-bit) of the 32-bit memory chip correspond to the physical blocks in the 32-bit memory chip. In contrast, 16-bit memory chips only involve the first channel, which corresponds to the physical blocks in the 16-bit memory chip. One channel is equivalent to one interface, meaning that the two interfaces of the 32-bit memory are connected to the two interfaces of the CPU. The program of the 32-bit test system scans all memory space of the 16-bit memory chip. If the program encounters an error, it is determined that the corresponding program involves the second channel. The code involving the second channel in the test system is found and the code corresponding to the second channel is masked. If the program can run, then use the program with the second channel-related code disabled to perform read and write tests on 16-bit memory; If the program can be stopped, modify the code involving the second channel to the first channel. After the program is modified, perform read and write tests on the 16-bit memory.
[0006] The beneficial effects of this invention are as follows: By employing methods such as channel shielding, code modification, and address mapping, a 32-bit testing system achieves compatibility with two types of memory chips. When testing 32-bit chips, a dual-channel path is enabled; when testing 16-bit chips, stability is ensured through single-channel adaptation and three-layer testing. This compatibility allows a single testing system to cover both types of memory chips without modification, avoiding the need for separate development of testing systems due to different chip types, and reducing equipment procurement and maintenance costs. The solution supports the special scenario of "second-channel access" for 16-bit chips. By adjusting channel mapping and testing procedures, it can adapt to the requirement of "temporarily enabling the second channel when the first channel fails." Simultaneously, the combination of 16MB block segmentation testing and the three-layer testing method can be applied to memory chips of different capacities without adjusting the core logic due to changes in memory capacity, further enhancing the universality and practicality of the solution. Attached Figure Description
[0007] Figure 1 This is a flowchart illustrating the testing method for memory chips according to the first embodiment of the present invention.
[0008] Figure 2 yes Figure 1 A flowchart illustrating step 1.
[0009] Figure 3 yes Figure 1 A flowchart illustrating the process after step 4. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0011] The terms "comprising" and "having," and any variations thereof, used in this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0012] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0013] Figure 1 This is a flowchart illustrating the testing method for memory chips according to the first embodiment of the present invention. Figure 1 As shown, the system includes hardware and software components: Step 1: Since 32-bit memory chips involve a first channel (16-bit) and a second channel (16-bit), the first channel (16-bit) and the second channel (16-bit) of the 32-bit memory chip correspond to the physical blocks in the 32-bit memory chip. In 16-bit memory chips, only the first channel is involved, and the first channel corresponds to the physical blocks in the 16-bit memory chip. Here, one channel is equivalent to one interface, that is, the two interfaces of the 32-bit memory are connected to the two interfaces of the CPU. Step 2: Scan all memory spaces of the 16-bit memory chips using the program of the 32-bit test system. If the program encounters an error, determine that the corresponding program involves the second channel, find the code in the test system that involves the second channel, and mask the code corresponding to the second channel. Step 3: If the program can run, perform read and write tests on 16-bit memory using the program with the second channel-related code disabled. Step 4: If the program can be stopped, modify the code involving the second channel to the first channel. After the program is modified, perform read and write tests on the 16-bit memory.
[0014] Step 1: Correspondence between channel architecture and physical block of memory chips (basic premise) Clarify the essential difference between 32-bit and 16-bit memory chips in the correspondence between "channel (interface) - physical block". The essence of the channel is the data transmission interface between memory and CPU, which determines the bit width and path of data interaction.
[0015] The channel architecture of 32-bit memory chips: A 32-bit memory chip has a 32-bit bit width, and its data transmission is split into two independent 16-bit channels (Channel 1 and Channel 2). These two channels each correspond to different physical storage blocks within the memory chip—essentially dividing the 32-bit storage capacity and data transmission capability into two separate 16-bit "sub-interfaces." Simultaneously, these two channels directly connect to the corresponding two 16-bit interfaces on the CPU, forming a "dual-interface" data transmission link to ensure complete data interaction across the 32-bit bit width.
[0016] The channel architecture of 16-bit memory chips: A 16-bit memory chip has a bit width of only 16 bits. The hardware only designs a first channel; there is no physical structure for a second channel. Its single first channel directly corresponds to all the physical storage blocks inside the chip. In this case, the 16-bit memory chip only needs to connect to a single 16-bit interface of the CPU to complete data transmission; the second channel is "missing" at the hardware level. The program in the 32-bit memory testing system is based on the dual-channel hardware design of 32-bit chips and will by default use both the first and second channels. However, 16-bit chips do not have a second channel, and running the program directly will cause problems due to hardware incompatibility.
[0017] Step 2: Scanning and locating code in the test system that involves the second channel and then masking it. The purpose is to find the code segment in the test system that calls the "non-existent second channel" and temporarily eliminate hardware mismatch errors by masking it. The specific logic is as follows: When scanning memory space and handling errors, a program from a 32-bit test system scans the entire memory space of a 16-bit memory chip. The program runs according to the dual-channel logic of the 32-bit chip, attempting to interact with memory simultaneously through both the first and second channels. However, the 16-bit chip lacks a physical interface and corresponding physical block for the second channel. Read / write commands and address requests sent by the program to the second channel will trigger errors due to "hardware unresponsiveness" (such as addressing failure, data transmission interruption, or program crash). Once such an error occurs, it can be determined that the erroneous program code segment is calling code from the second channel—because the first channel is compatible with the 16-bit chip hardware, and only operations on the second channel will trigger hardware-level incompatibility.
[0018] The purpose of disabling the second-channel code is to find the code that involves the second channel and then "disable" it (such as commenting out the code, disabling related function calls, blocking the instruction sending path of the second channel, etc.). Essentially, it makes the test program temporarily "ignore" the existence of the second channel and only retain the calling logic of the first channel.
[0019] Step 3: Based on the 16-bit memory read / write test of the disabled program, confirm whether the program with "only the first channel retained" can interact normally with the 16-bit memory chip. The specific logic is as follows: The prerequisite for the test is that the program has completed the masking of the second channel code. At this time, the program's running logic is forcibly switched from "dual channel" to "single channel", and data is only exchanged with the 16-bit memory chip through the first channel to match the hardware architecture of the 16-bit chip.
[0020] The core purpose of the test is to perform read and write tests (such as writing specific data to a specified physical block of memory, then reading to verify data consistency, traversing memory addresses to perform batch read and write, etc.). It mainly verifies two key points: first, whether the program can run stably without hardware incompatibility errors after the second channel is blocked; second, whether the code logic of the first channel can correctly interact with the physical block of the 16-bit chip and whether the data read and write is accurate.
[0021] The significance of the test results is as follows: if the read and write tests are normal, it means that the logic of "using only the first channel" is fully compatible with 16-bit memory chips, and only a permanent logical modification to the second channel code is needed in the future; if the test still results in an error, it is necessary to go back to step 2 to check whether there is any unmasked second channel code, or whether there is a logical problem in the code of the first channel itself (such as address addressing errors, improper data width processing, etc.).
[0022] Step 4: Modify the second channel code to the first channel and complete the final test. Unlike the "temporary masking" in Step 2, this step is a permanent conversion from dual-channel to single-channel from the code logic. The specific logic is as follows: The premise of "the program can stop running" does not mean that the program crashes, but that the program can start normally, complete read and write tests, and terminate normally according to the instructions after the second channel is blocked. This means that the interaction between the program's first channel logic and 16-bit memory is completely stable, the temporary adaptation solution is effective, and the conditions for permanent code modification are met.
[0023] The core logic of the code modification is to change the "code involving the second channel" found in step 2 from "masking" to "logical replacement": This is not a simple deletion or commenting, but rather mapping all instructions, addresses, and functions that originally called the second channel to the first channel. For example, memory addresses that originally pointed to the second channel are now addressed to the first channel; read / write instructions that were originally sent to the second channel are now called to read / write functions in the first channel; and 16-bit data transfer tasks originally allocated to the second channel are merged into the first channel's transfer logic. This ensures that all functions originally designed for the second channel are implemented through the first channel, completely eliminating the dependency on the second channel while fully utilizing the first channel's 16-bit width to cover the entire physical block of the 16-bit memory chip.
[0024] In the final read / write test, after the code modifications are complete, a full read / write test is performed on the 16-bit memory again to verify whether the modified program can stably and accurately complete all memory interaction operations. If the test passes, it indicates that the program of the 32-bit test system has achieved permanent adaptation to the 16-bit memory chips; if the test encounters errors, it is necessary to check whether the code modifications are thorough (e.g., whether there is any missing second-channel logic), or whether the load on the first channel is abnormal due to the merging of the second-channel tasks (e.g., data transfer rate, address conflicts, etc.).
[0025] Figure 2 yes Figure 1 The flowchart for step 1 is as follows: Figure 2 As shown, step 1 includes: Step 101: The physical block mapping of the 32-bit memory chip is a set of independent storage units inside each 16-bit channel corresponding memory chip. In this case, the 32-bit memory chip may actually be packaged from two 16-bit memory chips, which are bound to different channels respectively. The CPU's memory controller connects to the first channel and the second channel through two independent 16-bit interfaces respectively. Step 102: When the 16-bit memory chip uses only the first channel, all physical blocks are accessed through the first channel. Step 103: When the 16-bit memory chip uses only the second channel, all physical blocks are accessed through the second channel.
[0026] Step 101: Physical Block Structure and Channel Binding of 32-bit Memory Chips (Hardware-Level Supplement), revealing the physical structure of 32-bit memory chips and the specific implementation of physical block and channel binding, making the hardware essence of the dual-channel architecture clearer: From a hardware manufacturing perspective, a 32-bit memory chip's "dual 16-bit package" structure is not a single 32-bit storage core. Instead, it often uses a "two 16-bit memory chips packaged together" approach—that is, two independent 16-bit memory chips (each with a complete 16-bit memory bus width) are packaged within the same 32-bit chip casing, forming a "1+1=32-bit" width combination. This packaging method utilizes mature 16-bit chip manufacturing processes and can quickly meet the product requirements of 32-bit width, making it a common hardware design solution in the industry.
[0027] The "one-to-one binding" rule between physical blocks and channels means that the two packaged 16-bit chips are each bound to one of the two channels of the 32-bit chip: all physical storage units (i.e., all physical blocks) of the first 16-bit chip uniquely correspond to the first channel, and data can only be read and written through the first channel; all physical storage units (all physical blocks) of the second 16-bit chip uniquely correspond to the second channel, and data can only be read and written through the second channel. This binding is a fixed design at the hardware level and cannot be modified by software, ensuring that the storage resources of the two channels are independent of each other and avoiding access conflicts.
[0028] The CPU employs a "dual independent interface" connection logic with the memory channels. The CPU's memory controller is specifically designed with two independent 16-bit interfaces, directly connecting to the first and second channels of the 32-bit memory chips respectively. The first interface is solely responsible for communication with the first channel, sending addressing, read, and write instructions for the first 16-bit chip; the second interface is solely responsible for communication with the second channel, sending addressing, read, and write instructions for the second 16-bit chip. This "dual interface-dual channel" connection essentially allows the CPU to access two 16-bit chips simultaneously in parallel, ultimately achieving high-speed data transmission with a 32-bit width (parallel transmission through two 16-bit channels, totaling 32 bits). This is the hardware basis for the higher transmission efficiency of 32-bit chips compared to 16-bit chips.
[0029] Step 102: Physical block mapping rules for 16-bit memory chips' "single-channel (first channel) access" clarify the access association between physical blocks and channels when the 16-bit chip uses only the first channel, and improve the hardware interaction rules for single-channel scenarios: The core logic of the "all-physical-block unified channel" is that the 16-bit memory chip itself only has a 16-bit wide physical storage unit (i.e., a complete physical block), and the hardware is only designed with an interface that matches the first channel. When only the first channel is used, all physical blocks of the chip (all storage units from the first address to the last address) will be "normalized" to the first channel—that is, no matter which physical block is accessed, instructions and data must be sent through the first channel, and there is no situation where "some physical blocks belong to other channels".
[0030] Regarding the CPU interface compatibility, the CPU's memory controller only needs to enable the 16-bit interface corresponding to the first channel. By connecting to the first channel of the 16-bit chip through this interface, access to all physical blocks can be covered. The CPU's second channel interface will be in an "idle state" because the 16-bit chip does not have corresponding second channel hardware. Even if it is enabled, a valid connection cannot be established. This further explains why the 32-bit test program triggers an error when calling the second channel.
[0031] Step 103: Physical block mapping rules for 16-bit memory chip "single-channel (second-channel) access" (supplementary for special scenarios), revealing the possibility of "reverse channel access" and improving the full-scenario coverage of single-channel access: The "hardware interface compatibility" allows for reverse channel usage. Although 16-bit memory chips are designed by default to adapt to the first channel, some 16-bit chips have "channel compatibility" in their hardware interface. This means that the signal pins of their physical blocks can be matched not only with the CPU's first channel interface but also with the CPU's second channel interface through hardware jumpers, BIOS settings, or test system configuration. This design is mainly used for testing, repair, or special hardware combination scenarios (such as temporarily using the second channel when the first channel interface fails).
[0032] The "full physical block reverse mapping" access rule means that when the 16-bit chip uses only the second channel, the access paths for all its physical blocks will be "reversed": all physical blocks that were originally accessed through the first channel will be remapped to the second channel, and can only be addressed and read / written through the second channel. At this time, the CPU needs to enable the 16-bit interface corresponding to the second channel, and the first channel interface is idle. The access logic is completely symmetrical to "using the first channel", the only difference being the reversal of the channel numbers.
[0033] The relevance of compatibility with 32-bit testing systems is illustrated in this scenario: if the 16-bit NAND flash memory supports second-channel access, it can be bound to the second channel during testing. In this case, the 32-bit testing program needs to disable the first-channel code and modify it to call the second-channel code, with the adaptation logic consistent with "using the first channel." This also demonstrates that the channel access of 16-bit NAND flash memory possesses "two-way flexibility," rather than being limited to the first channel, providing a hardware foundation for multi-scenario adaptation of the testing system.
[0034] Step 101 reveals the physical structure of the 32-bit chip's "dual 16-bit package," transforming the logic of "dual channels corresponding to dual physical block sets" from an "abstract architecture" into a "concrete hardware implementation," explaining the fundamental reason for the independence of the dual channels. Regarding the original steps 2-4 (program adaptation): Steps 102 and 103 clarify the rule of "all physical blocks unified into one channel" when accessing a single channel of the 16-bit chip, further verifying the rationality of "32-bit programs will encounter errors when calling unused channels"—if the 16-bit chip uses the first channel, calling the second channel will result in no corresponding physical block; if the second channel is used, calling the first channel will also result in no corresponding physical block, reinforcing the core basis of "channel-physical block binding" for program adaptation. The overall logical loop: from "dual 16-bit package of 32-bit chips → channel-physical block binding → single-channel full physical block access of 16-bit chips," a complete chain of "hardware structure → channel rules → access scenarios" is formed, making the underlying hardware logic of the 32-bit test system more robust for adapting to 16-bit chips, and providing more accurate hardware references for subsequent program modifications and troubleshooting.
[0035] Following step 1, the method further includes: Step 104: Initialize memory by writing all 1s (0xFFFFFFFF) to ensure that all storage units are in a known state, which is used to detect whether the storage units can correctly retain data; Step 105: In memory testing, fixed-mode writing is used to detect "stuck bits"; Step 106: If the values read after writing are inconsistent, it indicates a data bus failure or a damaged storage unit.
[0036] Step 104: The core logic and function of initializing memory in all-1 mode (0xFFFFFFFF) is to establish a "known baseline state" for subsequent fault detection through unified all-1 data writing. The core logic and value are as follows: The underlying principle of "establishing a known state" is that when memory chip storage cells are not initialized, they may be in a "random state" (uncertain data values) due to factors such as power outage residue and manufacturing process deviations. Direct testing makes it difficult to distinguish between "storage cell failure" and "initial random data interference." By writing all 1s (0xFFFFFFFF, 32-bit data; 0xFFFF for 16-bit chips) to all storage cells (regardless of which channel or chip type), all storage cells can be forced into a "unified known state"—that is, the theoretical value of each storage cell is all 1s. Subsequent readings only require comparing whether the "actual value equals all 1s" to directly determine whether the storage cell can correctly retain the data.
[0037] For channel architecture adaptation details, for 32-bit memory chips: it is necessary to write all 1 data to the corresponding two 16-bit chip storage units through two channels respectively—the first channel writes to all physical blocks of the first 16-bit chip, and the second channel writes to all physical blocks of the second 16-bit chip, ensuring that all storage units of the 32-bit chip are initialized; for 16-bit memory chips (regardless of whether the first or second channel is used): it is only necessary to write all 1 data (0xFFFF in the 16-bit scenario) to all physical blocks through the currently enabled channel. Idle channels do not need to be operated to avoid meaningless instruction sending.
[0038] The core function of "data retention capability detection" is to read all storage units after initialization, typically at intervals of a certain time (e.g., milliseconds or seconds, simulating data retention scenarios in actual use). If the read values are still all 1s, it indicates that the storage unit can stably retain data; if some bits become 0 or other values, it can be preliminarily determined that these storage units have "data retention failures" (such as data loss due to leakage), laying the foundation for subsequent accurate fault type identification.
[0039] Step 105: The principle and scenario of fixed-mode write detection of "stuck bits". When a storage unit fails, the faulty bit is accurately identified through repeated writes and reads in a fixed mode. The specific logic is as follows: The definition and hazards of a "stuck bit": A "stuck bit" refers to a memory storage unit where one (or more) bits are damaged by hardware (such as transistor breakdown or short circuit), causing their data value to remain "fixed"—regardless of whether 0 or 1 is written, the read value is always a fixed value (such as always 0 or always 1). This type of fault leads to data writing distortion and is easily missed by conventional random data testing, requiring specialized detection through fixed-pattern writing.
[0040] The selection and testing logic of fixed modes commonly include "all 0 mode (0x00000000)," "all 1 mode (0xFFFFFFFF)," and "chessboard mode (alternating between 0x55555555 and 0xAAAAAAAA)." The core logic is to repeatedly cover the storage unit with "fixed data with clear characteristics" to amplify the fault characteristics of the stuck bits. For example, first write all 0 mode to the storage unit, read and record all bits whose "actual value ≠ 0" (which may be stuck bits); then write all 1 mode to the same storage unit, read and record all bits whose "actual value ≠ 1." If a certain bit does not change with the written value in both writes (e.g., it is 1 when writing 0 and still 1 when writing 1), then it can be determined that the bit is a "stuck bit," and the stuck value is 1. Similarly, bits with a stuck value of 0 can be identified.
[0041] In coordination with channel access rules, testing must strictly adhere to the "channel-physical block" binding rule: fixed-mode writes must be performed through the channel currently enabled by the chip (dual-channel path for 32-bit chips, single-channel path for 16-bit chips) to ensure that write commands accurately reach the target storage unit. Simultaneously, each channel's corresponding storage unit must be tested independently to avoid cross-channel interference—for example, the stuck bit of the storage unit corresponding to the first channel of a 32-bit chip must be identified through write and read operations on the first channel, and is unrelated to the second channel.
[0042] Step 106: Fault type determination logic for inconsistent write and read values. Based on the phenomenon of "write value ≠ read value", combined with the previous test results, it distinguishes between "data bus failure" and "storage unit damage" to avoid misjudgment of faults. The specific logic is as follows: The commonalities and differences of "inconsistency": After writing to a certain pattern (e.g., all 1s, all 0s), inconsistency between the read and written values is a typical manifestation of memory failure. However, the "inconsistency characteristics" caused by different failure types differ, requiring comprehensive judgment based on previous tests (e.g., initialization in step 104, fixed-pattern testing in step 105). "Inconsistency" exhibits "bit randomness": that is, the inconsistent bits in different storage units are irregular (e.g., some bits are 0, some are 1, and the distribution has no fixed pattern), and the position of the inconsistent bits changes when repeatedly writing to the same pattern—this is most likely a data bus failure. Because the data bus is responsible for transmitting written data, if a line on the bus has poor contact or signal attenuation, some bits will flip during transmission, and the flipped bits will change with fluctuations in the transmitted signal, exhibiting randomness. If the "inconsistency" exhibits "bit fixation": that is, the position of the inconsistent bit is fixed (e.g., the 3rd bit of several memory cells is always inconsistent), and the state of the inconsistent bit remains unchanged when repeatedly written to the same mode (e.g., it is always 0) – it is highly likely that the memory cell is damaged. Because memory cell damage (e.g., stuck bit, leakage) will cause its data value to be fixed or unable to be maintained, the position and state of the faulty bit are relatively stable, which is significantly different from the randomness of bus transmission.
[0043] In conjunction with the three-layer testing method, the preliminary judgment in step 106 can be further verified by the three-layer testing method: if the preliminary judgment is a data bus fault, the first-layer data bus test (such as Walking1 mode) can be performed to verify whether there is a bit flip on the bus; if the preliminary judgment is a memory cell damage, the third-layer device test (such as long-term residence test) can be performed to confirm the degree of damage to the memory cell, and finally realize the fault location closed loop of "preliminary judgment → accurate verification".
[0044] The fault diagnosis differs for different types of NAND flash memory. For 32-bit NAND flash memory: if only the memory cell corresponding to a certain channel is inconsistent, while the other channel is normal, the bus of that channel should be checked first (e.g., if the bus failure of the first channel only affects the first 16-bit NAND flash memory); if both channels are inconsistent, their respective buses or corresponding memory cells should be checked separately. For 16-bit NAND flash memory: only the bus and all memory cells of the currently active channel need to be checked. Idle channels have no corresponding hardware and do not need to be included in the diagnosis scope to avoid expanding the scope of fault diagnosis.
[0045] Steps 104-106 clarify the test operation rules for different particle types and channel scenarios, ensuring that the "channel-physical block" binding relationship is implemented in specific test execution, avoiding test failures due to improper channel operation. When the 32-bit test program is adapted to 16-bit particles, steps 104-106 can serve as the core test method after "disabling / modifying channel code"—for example, after disabling the second channel, the memory unit corresponding to the first channel is verified to be normal through all-1 initialization and fixed-mode writing, ensuring that the program can not only run after adaptation but also accurately detect hardware faults.
[0046] Step 105 includes: Step 1051: Divide the memory into 16MB blocks and test each block in a loop to isolate address line errors; Step 1052: If the block test passes, the block belongs to the first channel, as it can be accessed normally with only a 16-bit interface. Step 1053: If the block test fails, it is due to a fault in the second channel, and the read data does not match the expected value.
[0047] Step 1051: Memory testing logic for 16MB block segmentation (address line error isolation core). By dividing the memory into independent test blocks of a fixed size (16MB), the impact range of address line errors is accurately located, avoiding "generalization" of fault location. The specific logic is as follows: The underlying principle and operation of "block partitioning" is that the memory address space is contiguous. However, address line errors (such as a faulty address line or signal interference) often prevent the normal access of memory units within a "specific address range." For example, if an address line corresponds to the addressing control of a "16MB address segment," a failure in that address line will cause addressing errors for all addresses within that 16MB range. Therefore, partitioning memory in 16MB increments essentially breaks down the contiguous address space into independent units that match the "address line control range," allowing each test block to correspond to a set of address line control areas.
[0048] In practice, the number of blocks needs to be calculated based on the total memory capacity (e.g., 1GB of memory can be divided into 64 16MB blocks), and each block needs to be assigned an independent "block number + address range" (e.g., block 1 corresponds to address 0x00000000-0x00FFFFFF, block 2 corresponds to 0x01000000-0x01FFFFFF, etc.) to ensure that the address range of each block does not overlap or omit anything, and covers the entire memory space.
[0049] The execution logic of "loop testing" and its value in isolating address line errors lies in its execution logic. Loop testing refers to performing a complete read / write test sequentially on each 16MB block according to its block number (e.g., writing data in a fixed pattern first, then reading and comparing). During the test, operations are only performed on the address range of the current block, without involving other blocks. This method can isolate address line errors through "inter-block differences in test results." If a few consecutive or specifically numbered blocks fail the test, while other blocks are normal, it means that the faulty address line corresponds to the address range of these failed blocks. For example, if the test fails for blocks 3-5, it means that there is an error in a certain address line that controls the address range of these three blocks, rather than all address lines being faulty. If all blocks fail the test, then it is necessary to check the global address lines (such as the core control lines of the address bus) or other common problems (such as data bus failure).
[0050] Compared to a "one-time full memory test", a 16MB block segmentation test can narrow down the location of address line errors from "the entire memory" to "a 16MB interval", significantly reducing the complexity of subsequent troubleshooting.
[0051] Regarding the adaptation details with the channel architecture, the segmentation test must strictly adhere to the "channel-address space" binding rule: For 32-bit memory chips, the two channels correspond to independent address spaces (the first channel corresponds to one portion of the address space, and the second channel corresponds to another portion). Segmentation must be performed according to the address range of each channel—for example, if the first channel corresponds to 512MB of memory, it should be segmented into 32 16MB blocks; the second channel also corresponds to 512MB of memory and should also be segmented into 32 16MB blocks to avoid test interference caused by cross-channel segmentation. In contrast, 16-bit memory chips correspond to the address space of only a single channel and can be directly segmented into 16MB blocks based on the total capacity.
[0052] Step 1052: Test the "first channel attribution" determination logic of the passed block (single channel adaptation verification). Based on the block test results, clarify the correspondence between the "normal access block" and the first channel, and verify the access validity of the single channel (16-bit interface). The specific logic is as follows: The core criterion for "test passed" is that a 16MB block passes the test. This means that after performing a fixed-mode write (such as all 1s, all 0s, or checkerboard pattern) within the address range of the block, the read values of all memory cells are completely consistent with the expected values, and the results are stable after repeated tests. This indicates that the block's "address line addressing is normal" (it can accurately locate each memory cell), "data bus transmission is normal" (data does not flip), and "memory cell read and write is normal" (no dead bits, data is well maintained), and the overall hardware link is fault-free.
[0053] The underlying logic of "belonging to the first channel" combines the hardware architecture of "dual 16-bit packaging + dual-channel binding" for 32-bit memory chips (step 101) and the rule of "single-channel access" for 16-bit memory chips (step 102). The core basis for "blocks that pass the test belong to the first channel" is that the first channel corresponds to a 16-bit wide hardware interface, and the test of a 16MB block is based on a 16-bit wide read / write logic design (such as parallel transmission of 16-bit data and addressing control of 16-bit addresses). Only the 16-bit interface of the first channel is needed to complete all access operations - stable read and write can be achieved without calling the interface of the second channel. If a block needs to rely on the second channel to be accessed, its test must be based on the second channel being normal. However, in the judgment scenario of step 1052, the test passed without involving the second channel operation, indicating that the hardware link of the block is fully compatible with the first channel and does not require the participation of the second channel. Therefore, it belongs to the first channel.
[0054] The significance of verifying compatibility with 16-bit interface access is that the result of this step can directly verify the "access coverage capability of the 16-bit interface"—if most of the 16MB blocks pass the test and belong to the first channel, it means that the 16-bit interface of the first channel can stably cover the address space of these blocks. In subsequent 16-bit memory particle test scenarios (only the first channel is enabled), these blocks can be selected as "benchmark test blocks" to reduce test errors caused by hardware incompatibility.
[0055] Step 1053: Attribution logic for the "second channel failure" of the test failure block (fault correlation analysis). For the phenomenon of block test failure, clarify its correlation with the second channel failure and explain the core reason for "data mismatch". The specific logic is as follows: The typical manifestation of "test failure" is related to the second channel. The core manifestation of block test failure is "data read does not match the expected value". The premise of attributing it to "second channel failure" is that the hardware link of the test block is bound to the second channel - that is, in the 32-bit memory chip, the 16MB block belongs to the 16-bit chip storage unit corresponding to the second channel ("the second 16-bit chip is bound to the second channel" in step 101). Its access must rely on the interface of the second channel. If the second channel has a fault (such as a short circuit on a line of the data bus, attenuation of the address line signal, or poor contact between the second channel and the CPU interface), it will cause the access link of the block to be interrupted: when writing data, the second channel cannot accurately transmit the data to the storage unit; when reading data, the second channel cannot accurately transmit the data from the storage unit back, ultimately resulting in "read value ≠ expected value", and test failure.
[0056] The troubleshooting and verification logic for "second channel failure" requires confirming that the failed test block was indeed caused by a second channel failure. This necessitates performing "fault isolation verification": Step 1: Disable the code involving the second channel in the test program (refer to step 2) and attempt to access the failed block only through the first channel interface. If access is still unsuccessful or the read data does not match, the failure may not be due to the second channel, and the address lines or the storage unit itself need to be checked. Step 2: Repair the second channel failure (e.g., reconnect the interface, replace the faulty second channel line), and test the failed block again. If the test passes, it confirms that the previous failure was indeed caused by a second channel failure, verifying the accuracy of the attribution.
[0057] The significance of adapting to 32-bit testing systems, and the attribution logic of this step, further improves the troubleshooting system for 32-bit testing systems adapted to 16-bit memory chips: In scenarios where 32-bit testing programs default to dual-channel operation, if a 16MB block fails to test, it is not necessary to directly determine that the storage unit is damaged. Instead, the second channel fault should be investigated first. Especially in scenarios where only the first channel is used for 16-bit memory chips, a second channel fault will cause the block it is bound to to be inaccessible. In this case, it is necessary to disable the second channel code, repair the second channel, etc., so that the testing program can focus only on the block corresponding to the first channel and avoid misjudgment of faults.
[0058] The block segmentation test provides a more precise range for the "address line error determination" in step 106. If a 16MB block fails the test due to an address line error, step 106 can quickly locate the faulty address line by combining the address range of the block, avoiding blind troubleshooting across the entire memory range. For steps 1052 (first channel attribution) and 102 (16-bit chip single-channel access): the determination result of step 1052 verifies the feasibility of "16-bit chip full physical block access through the first channel" in step 102. The blocks that pass the test are attributable to the first channel, indicating that the 16-bit interface of the first channel can cover the access requirements of these blocks and is fully compatible with the single-channel hardware architecture of the 16-bit chip. For steps 1053 (second channel fault attribution) and 2 (second channel code masking): if the block test failure is attributed to a second channel fault, the operation in step 2 can be directly referred to to mask the second channel code, avoiding interference from the faulty channel to the test, and providing a clear fault indication for subsequent second channel repair, reducing invalid operations.
[0059] After step 1053, the method further includes: In step 1054, if the program can run after the relevant code of the second channel is blocked, then the error only exists in the second channel; In step 1055, the address space of the second channel is mapped to the first channel, so that the access code of the second channel is redirected to the first channel, and the 16-bit memory chip communicates only through the first interface.
[0060] Step 1054: Error location logic that allows the program to run after disabling the second-channel code. By reverse-engineering the program's running state after disabling the second-channel code, the unique source of the error can be confirmed, avoiding misjudgment of the fault range. The specific logic is as follows: The core criterion for determining whether a program is "runnable" is that after disabling the code related to the second channel (such as commenting out the initialization function of the second channel, blocking the instruction sending link of the second channel, disabling the address addressing logic of the second channel, etc.), the program can start normally and complete the core test process—including dividing memory into 16MB blocks, performing fixed-pattern read and write tests on each block, comparing the read data with the expected value, etc.—without errors such as "addressing failure," "data transmission interruption," or "program crash." "Runnable" here not only means that the program does not report errors, but also requires that the test functions are complete—for example, it must be able to cover all 16MB blocks corresponding to the first channel, and the test results must be stable, eliminating the invalid state of "program starting but test functions missing."
[0061] The deduction that "the error only exists in the second channel," combined with the previous testing process (steps 1051-1053), shows that the program can run after disabling the second channel code. The source of the error can be deduced through the process of elimination: First, the core links on which the program depends (the address lines, data bus, and memory units of the first channel, as well as the CPU's first channel 16-bit interface) are all normal. If these links have errors, even if the second channel is disabled, the program will still report an error due to abnormal access to the first channel. However, the current program can run, indicating that the first channel and its associated hardware are not faulty. Second, the root cause of the original program error is the call to the second channel. After disabling the second channel code, the program no longer interacts with the second channel, and the error disappears. This suggests that the original error was triggered by operations related to the second channel. Finally, the possibility of "cross-fault between the first and second channels" is ruled out. If a cross-fault exists (such as the first channel relying on the signal of the second channel to work normally), disabling the second channel would cause the first channel to also fail to run. However, the current program can use the first channel normally, indicating that the hardware links of the two channels are independent of each other, and the faults have no cross-effect.
[0062] In summary, it can be determined that the error exists only in the second channel, eliminating the need to investigate other hardware or software logic and significantly narrowing down the scope of fault repair.
[0063] The practical significance of adapting to 32-bit test systems is that this step provides a "fault location basis" for adapting 16-bit memory chips to 32-bit test systems. In scenarios where only the first channel of the 16-bit memory chip is enabled, if the original 32-bit program reports an error, by disabling the second channel code to verify whether the program can run, it is possible to quickly determine whether the incompatibility is caused by the second channel call, avoid misjudging "software channel call error" as "hardware failure", and reduce unnecessary hardware troubleshooting operations.
[0064] Step 1055: The implementation logic of mapping the second channel address space to the first channel. This step is the core adaptation method of single-channel covering dual-channel address space. Through address mapping, the access requirements of the second channel are redirected to the first channel, realizing the goal of 16-bit memory chips communicating only through the first interface. The specific logic is as follows: The core definition and goal of "address space mapping" is that the address space of the second channel refers to the address range corresponding to the 16-bit chip bound to the second channel in a 32-bit memory chip (e.g., 0x80000000-0xFFFFFFFF, the specific range depends on the hardware design). "Mapping the second channel address space to the first channel" essentially redirects "address requests" originally pointing to the second channel address space to the first channel address space through software logic or hardware configuration—making the program mistakenly believe it is accessing the second channel, while all operations are actually completed through the first channel, ultimately achieving "a single-channel interface covering the address requirements of a dual-channel system."
[0065] For example, if the address of a 16MB block in the second channel is 0x81000000-0x81FFFFFF, after mapping, the program's access to that address will be converted to access to the first channel address 0x01000000-0x01FFFFFF. Data reading and writing are both completed through the 16-bit interface of the first channel.
[0066] The specific implementation of mapping depends on both the hardware architecture and software logic. There are two common implementation methods: Software-level address offset mapping: Add an "address translation function" to the test program. When the program generates the address of the second channel, automatically subtract the "second channel address start value" and add the "first channel address offset value" to convert it into a valid address of the first channel. For example, if the second channel address start value is 0x80000000 and the first channel address offset value is 0x00000000, then the second channel address 0x81000000 will be converted to 0x01000000 (0x81000000 - 0x80000000 + 0x00000000 = 0x01000000). Simultaneously, ensure that the converted first channel address does not exceed its actual address range to avoid address overflow. Hardware-level address redirection: This involves setting "redirection rules for the second-channel address space" through the memory controller's configuration registers (such as the CPU's memory mapping register). This establishes a one-to-one mapping between the address ranges of the second channel and the address ranges of the first channel, with the hardware automatically performing the address translation without software intervention. The advantage of this method is its fast response time and avoidance of latency caused by software translation, making it suitable for scenarios with high access efficiency requirements.
[0067] The adaptation result of "16-bit memory chips communicating only through the first interface" is that after mapping, the communication link of the 16-bit memory chips is "completely unified": all access code in the program involving the second channel (such as address addressing, data read and write instructions) will be converted into operations of the first channel through mapping, and finally the 16-bit interface of the first channel will interact with the memory chip; the 16-bit interface of the second channel will always be in an "idle state" and do not need to be enabled, completely avoiding incompatibility issues caused by the lack of second channel hardware; from the program's perspective, it can still "fully access" the original dual-channel address space, with no missing functions, but in reality, all communication is completed through the first interface at the hardware level, perfectly adapting to the single-channel architecture of the 16-bit memory chips.
[0068] In terms of logical connection with the previous steps, this step extends the solution to "the error is only in the second channel" in step 1054. Since there is a fault or hardware deficiency in the second channel, its access requirements can be transferred to the normal first channel through address mapping. This not only preserves the program's ability to access the complete address space, but also eliminates the need to repair or enable the second channel. It is a "low-cost and high-efficiency" adaptation solution in the 16-bit memory chip testing scenario. At the same time, it also improves the entire process logic of "fault location → solution → functional verification".
[0069] Step 1054, through "masking verification," transforms the "attribution hypothesis" of Step 1053 into a "determined conclusion," avoiding incorrect repair directions due to attribution bias. Step 1055's address mapping is a technical upgrade to Step 4, "modifying the second channel code to the first channel"—upgrading from "modifying code logic" to "address space redirection," offering greater flexibility and preserving the program's access to the complete address space, resulting in more comprehensive functionality. These two steps combine to form an adaptation path of "problem verification → solution"—first, Step 1054 confirms the error is only in the second channel, then Step 1055 implements address mapping, ensuring that the 16-bit chip can complete all tests using only the first interface, completely resolving the incompatibility issue between the dual-channel calls of the 32-bit program and the single-channel hardware of the 16-bit chip.
[0070] Figure 3 yes Figure 1 A flowchart following step 4 is shown below. Figure 3 As shown, after step 4, the method further includes: Step 5: Determine channel faults through data consistency, but if data bus errors, address bus errors, or memory unit errors are not distinguished, verify the fault through the three-layer test method; Step 6, First Layer: Verify the continuity and transmission accuracy of the data cable individually to eliminate faults at the data bus level; Step 7, Second Layer: Verify the uniqueness and accuracy of address line addressing based on address line overlap detection; Step 8: The third layer simulates actual business scenarios and performs full, multi-mode read and write tests on the memory storage unit; Step 9: Integrate the three-layer testing method into the previous process of adapting the 32-bit testing system to 16-bit memory chips.
[0071] Step 5: Determine the correlation between data consistency and channel failure triggering of the three-layer testing method. When channel failure is determined by data consistency (whether it equals 0xFFFFFFFF) but the error type cannot be further subdivided, the three-layer testing method needs to be activated for precise verification. The specific logic is as follows: The limitations of data consistency assessment, relying solely on whether a read value equals 0xFFFFFFFF to determine channel faults, can only provide a preliminary indication that a fault exists. It cannot distinguish whether the fault originates from the data bus (e.g., data transmission bit flipping), the address bus (e.g., addressing errors leading to reading incorrect memory cell data), or the memory cell itself (e.g., the memory cell cannot maintain all 1s). For example, a read value of 0xFFFFFFFE (only one bit different from 0xFFFFFFFF) could be caused by a fault on a data bus line leading to bit flipping during transmission, a stuck bit in the memory cell, or an addressing error reading non-all 1 data from another memory cell. Data consistency alone cannot pinpoint the exact cause.
[0072] The three-layer testing method is triggered immediately when a situation arises where "data consistency mismatch and fault type are ambiguous." Triggering scenarios primarily include: a data inconsistency error occurring when the 32-bit test program scans 16-bit memory chips (scenario 2); data inconsistency persisting in read / write tests even after disabling the second-channel code (scenario 3); and data anomalies appearing in tests after modifying the second-channel code (scenario 4). Activating the three-layer testing method at this point allows for gradual narrowing of the fault scope through layered verification, preventing misjudgments that could lead to deviations in the repair direction.
[0073] Step 6: Specific application of the first layer (data bus test) in the adaptation process. In the process of adapting a 32-bit test system to 16-bit memory chips, a data bus test needs to be performed based on the characteristics of the channel architecture. The specific application method is as follows: The testing scope is channel-dependent. For 32-bit memory chips: if a channel (such as the second channel) is suspected of being faulty, the 16-bit data bus of that channel must be tested separately. For example, in Walking1 mode, send 16 bits of data (such as 0x0001, 0x0002, 0x0004...0x8000) bit by bit to the memory cell corresponding to the second channel. After reading and comparing the data, if a bit cannot be accurately returned, the corresponding line of the data bus of that channel is determined to be faulty. At the same time, the data bus of the first channel must also be tested to ensure that it is normal, so as to provide a benchmark for subsequent adaptation.
[0074] For 16-bit memory chips: Only the 16-bit data bus of the first channel (or the enabled second channel) needs to be tested; idle channels do not need to be involved. During the test, send 16-bit wide Walking1 data through this channel, covering all data lines. If the test passes, the data bus fault is ruled out, and the fault is identified as the address bus or memory unit. If the test fails, the data bus problem is directly located, and there is no need to proceed to subsequent tests.
[0075] Integrating with the nodes of the adaptation process, this layer of testing can be incorporated into steps 2 (scanning and locating the second-channel code) and 3 (read-write test after masking): In step 2, if the program scans the 16-bit memory chip and finds data inconsistency, the first-layer test is executed first. If the data bus is fault-free, it is then determined to be a problem with the second-channel code call; In step 3, if data inconsistency still exists after masking the second-channel code, the first-layer test is used to rule out data bus faults first, and then other problems are investigated to improve fault location efficiency.
[0076] Step 7: The specific application of the second layer (address bus test) in the adaptation process, combining the address space characteristics of 32-bit and 16-bit memory chips, verifies the addressing accuracy through address line overlap detection. The specific application method is as follows: The test design matches the channel address space. Address line overlap detection needs to be performed independently for the address spaces of different channels: For the first channel of a 32-bit memory chip, the address space is usually a low address segment (e.g., 0x00000000-0x7FFFFFFF). Write unique identifier data to each address in this range (e.g., write 0x00000000 to address 0x0000000, write 0x00000001 to address 0x00000001, etc.), and then read and verify address by address. Perform the same operation on the high address segment corresponding to the second channel (e.g., 0x80000000-0xFFFFFFFF). If the address read data and identifier do not match in a certain channel, it is determined that the address bus of that channel has an overlap or failure problem.
[0077] For 16-bit memory chips, tests are only performed on the address space of the enabled channel (such as 0x00000000-0x7FFFFFFF of the first channel). If an address line fault occurs, the address bus must be repaired first (such as adjusting the address line connection or repairing the memory controller) before continuing the adaptation process to avoid subsequent tests being invalid due to addressing errors.
[0078] Integrating with the nodes of the adaptation process, this layer of testing can be incorporated into step 4 (testing after code modification) and step 1053 (attributing block test failures): In step 4, after modifying the second channel code to the first channel, if inconsistent data is found during testing, the second layer test is executed. If the address bus is fault-free, the code modification is then checked to ensure it is thorough. In step 1053, if the 16MB block test fails and is attributed to a second channel fault, the second layer test is executed first. If the second channel address bus is faulty, the address bus is repaired first, and then the test is repeated to avoid misjudging it as another fault.
[0079] Step 8: The specific application of the third layer (storage unit testing) in the adaptation process. Simulating real-world business scenarios, multi-mode testing of storage units is conducted to ensure their reliability during the adaptation process. The specific application methods are as follows: The test mode is combined with the adaptation scenario. For the scenario of adapting a 32-bit test system to 16-bit memory chips, the third-level test needs to cover all storage units of the 16-bit memory chip and adopts multi-mode verification: Perform an all-0 / all-1 test (write 0x00000000 and 0xFFFFFFFF and read) to verify the basic retention capability of the storage unit; perform a checkerboard test (alternately write 0x55555555 and 0xAAAAAAAA) to verify crosstalk between storage units; perform a long-term residence test (write data and read after 1 hour) to verify the long-term data retention capability. The test must be performed through an enabled channel (such as the first channel of the 16-bit chip). If data inconsistency occurs in a certain mode test, the storage unit is determined to be damaged and the memory chip needs to be replaced; if all mode tests pass, it means that the storage unit is not faulty and the adaptation process can continue.
[0080] Integrating with the nodes of the adaptation process, this layer of testing can be incorporated into step 3 (read / write test after masking) and step 1052 (block test pass determination): In step 3, after masking the second channel code, the third layer test is executed to ensure that the storage unit corresponding to the first channel is fault-free, providing a reliable hardware foundation for subsequent code modifications; In step 1052, after the 16MB block test passes, the third layer test is executed to further verify the stability of the block storage unit, avoiding subsequent test anomalies due to hidden faults in the storage unit.
[0081] Step 9: The overall framework for integrating the three-layer testing method into the adaptation process, determining the application order and role of each layer of testing at key nodes in the adaptation process, and forming a complete fault diagnosis and verification system. The specific integration framework is as follows: The adaptation process nodes correspond to the three-layer test. Step 2 (Scanning and locating the second-channel code): First, execute the first layer (data bus test) to rule out data bus faults; then execute the second layer (address bus test) to rule out address bus faults. If both layers are fault-free, the problem is identified as a second-channel code call issue, and the code masking phase begins. Step 3 (Read / Write Test after Masking): First, execute the first-layer test to verify the first-channel data bus; then execute the second-layer test to verify the first-channel address bus; finally, execute the third-layer test to verify the first-channel memory unit. If all three layers pass, the program adaptation after masking is effective, and the code modification phase begins. Step 4 (Testing after Code Modification): Execute the tests in the order of "first layer → second layer → third layer". If all three layers pass, the code modification is successful, and the adaptation is complete. If a certain layer test fails, perform targeted repairs (such as repairing the circuitry for data bus faults, adjusting the addressing logic for address bus faults, and replacing the memory chip for memory unit faults) and retest. Steps 1051-1053 (16MB block partitioning test): For blocks that fail the test, first perform the first-level test to eliminate data bus faults, then perform the second-level test to eliminate address bus faults, and finally perform the third-level test to verify the storage unit; for blocks that pass the test, perform the third-level test to confirm the stability of the storage unit and ensure accurate block ownership determination.
[0082] The optimized adaptation process, particularly the integration of the three-layer testing method, elevates the adaptation process of a 32-bit testing system to 16-bit memory chips from "single fault determination" to "layered, precise verification." This resolves the issues of ambiguous fault types and low troubleshooting efficiency in the original process. For example, in the original process, program errors might require checking both hardware links and software code one by one. However, with the integration of three-layer testing, fault types can be quickly located and targeted repairs can be made, significantly shortening the adaptation cycle. Simultaneously, layered testing ensures that each adaptation operation is based on "fault-free hardware," avoiding repeated adjustments to software adaptation due to hardware issues and improving adaptation reliability.
[0083] This embodiment achieves compatibility between the 32-bit test system and two types of NAND flash memory through channel shielding, code modification, and address mapping: when testing 32-bit NAND flash memory, a dual-channel path is enabled; when testing 16-bit NAND flash memory, stability is ensured through single-channel adaptation and three-layer testing. This compatibility allows a single test system to cover both bit width NAND flash memory without modification, avoiding the need to develop separate test systems for different NAND flash memory types and reducing equipment procurement and maintenance costs.
[0084] The solution supports the special scenario of "second channel access" for 16-bit memory chips (step 103). By adjusting the channel mapping and testing process, it can adapt to the requirement of "temporarily activating the second channel when the first channel fails". At the same time, the combination of 16MB block segmentation test and three-layer test method can be applied to memory chips of different capacities (such as 1GB and 2GB) without adjusting the core logic due to changes in memory capacity, further improving the universality and practicality of the solution.
[0085] Leveraging the "single-channel full physical block access" feature of 16-bit chips (steps 102-103), and combining it with address space mapping (step 1055), the first channel covers the original dual-channel address space, eliminating the need to activate the idle second channel interface. Simultaneously, a three-layer testing method is used to verify the carrying capacity of the first channel (e.g., data bus transmission efficiency, address space coverage), ensuring that the single channel can meet all the access requirements of the 16-bit chips while avoiding hardware resource waste (e.g., the CPU second channel interface and memory second channel lines remain idle and do not require modification).
[0086] Because this solution can accurately pinpoint faults (such as a faulty address line or a failed data line), it eliminates the need to replace the entire memory chip or CPU interface; only targeted repairs are required (such as adjusting the connection of a specific address line or repairing a data line in a particular channel). Furthermore, it eliminates the need for additional hardware to adapt to 16-bit memory chips (e.g., no need to add a second channel interface). Compared to traditional "complete hardware replacement" adaptation solutions, this reduces hardware investment costs by over 60%, making it particularly suitable for large-scale 16-bit memory chip testing scenarios.
[0087] Breaking away from the limitations of the original process that only "verifies data consistency," this approach achieves end-to-end testing: At the data bus level, Walking1 mode covers the continuity of each data line; at the address bus level, address line overlap detection and block segmentation ensure addressing uniqueness; at the storage unit level, all-0 / all-1, checkerboard, and long-term residency tests verify data retention capabilities and anti-crosstalk performance. For example, the third-layer long-term residency test (such as reading after 1 hour) can detect hidden faults such as "normal short-term testing, but data loss in long-term use" in advance, avoiding data deviations in the test system during actual application and ensuring the reliability of test results.
[0088] By employing "code modification (step 4) + address mapping (step 1055)," the 32-bit test program is ensured to be compatible with both 16-bit single-channel architectures and to retain access to the complete address space (e.g., mapping the second channel address to the first channel). Simultaneously, the three-layer testing method verifies the stability of the adapted program in different scenarios (e.g., random read / write, long-term operation), avoiding repeated adjustments due to program compatibility issues and reducing maintenance workload by over 90%.
[0089] By using "channel masking verification (steps 2 and 1054)" and "address space mapping (step 1055)," hardware faults and software logic problems are clearly distinguished: if the program runs normally after masking the second channel code, it is determined to be a channel call error at the software level; if errors still occur, the hardware link is investigated using a three-layer testing method. This isolation mechanism avoids "misjudging software compatibility issues as hardware damage" (e.g., no need to replace normal memory chips) or "misjudging hardware faults as software vulnerabilities" (e.g., no need to repeatedly debug code), significantly reducing the dual costs of hardware repair and software debugging.
[0090] For core adaptation issues (such as missing second-channel hardware and 16-bit interface coverage), clear solutions are provided: In the event of a second-channel failure, temporary adaptation is achieved by masking the code, followed by permanent coverage through address space mapping; when only a single 16-bit granular element exists, single-channel link stability is ensured through layer-3 testing. This "problem-as-it-is" response mechanism eliminates the need for additional adaptation path exploration. For example, if a 16MB block test fails, the failure is directly attributed to the second channel via step 1053, followed by address mapping in step 1055, avoiding adaptation stagnation due to uncertain solutions.
[0091] A standardized process has been established, which includes "hardware architecture clarification (steps 101-103) → fault location (steps 2, 1051-1053) → temporary shielding verification (steps 3, 1054) → permanent adaptation (steps 4, 1055) → three-layer test verification (steps 5-9)". For example, when a 32-bit test system is adapted to a 16-bit chip, the hardware logic of "dual 16-bit packaging" is first clarified through step 101. Then, the second channel code is scanned and located according to step 2. After shielding, it is verified by step 3. After modifying the code, it is confirmed through three-layer testing. Each step has clear guidance, avoiding the disorderly trial and error of "repeatedly debugging code and blindly checking hardware" in the original process, and compressing the adaptation cycle from "several days" to "hours".
[0092] By combining a "three-layer testing method + 16MB block segmentation", the first layer, Walking1 mode, is used to identify data bus faults. The second layer, address line overlap detection, is used to locate address bus problems. The third layer, multi-mode testing, is used to verify the reliability of the memory unit. Then, 16MB block segmentation is used to accurately locate the fault area (such as overlapping address lines or failure of a data line in a certain channel). This completely solves the pain point of "ambiguous fault type and unknown scope of impact", avoids ineffective troubleshooting caused by misjudgment (such as misjudging an address bus fault as a memory unit failure), and saves more than 70% of the fault location time.
[0093] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0094] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0095] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for testing memory chips, characterized in that, The method includes: The 32-bit memory chip involves a first channel with a 16-bit width and a second channel with a 16-bit width, with the first channel and the second channel each corresponding to different physical blocks; The 16-bit memory chip only involves the first channel, which corresponds to the physical block in the 16-bit memory chip. Among them, the 32-bit memory chip is the first memory chip, and the 16-bit memory chip is the second memory chip. The program of the 32-bit test system scans all memory space of the 16-bit memory chip. If the program encounters an error, it is determined that the corresponding program involves the second channel. The code involving the second channel in the test system is found and the code corresponding to the second channel is masked. If the program can run, then use the program with the second channel related code disabled to perform read and write tests on the 16-bit memory chip; If the program can start normally, complete the read and write test, and terminate normally according to the instructions after disabling the code related to the second channel, then modify the code in the program involving the second channel to the first channel. After the program is modified, perform read and write tests on the 16-bit memory chip. The physical block mapping of a 32-bit memory chip is a set of independent storage units inside each 16-bit channel corresponding memory chip. The 32-bit memory chip is packaged from two 16-bit memory chips, which are bound to different channels respectively. The CPU's memory controller is connected to the first channel and the second channel through two independent 16-bit interfaces respectively. If the first 16-bit memory chip in the first memory chip uses only the first channel, then all physical blocks are accessed through the first channel; If the second 16-bit memory chip in the first memory chip uses only the second channel, then all physical blocks are accessed through the second channel.
2. The testing method for memory chips according to claim 1, characterized in that, The method further includes: Memory is initialized by writing to all 1s to ensure that all storage units are in a known state, which is used to detect whether the storage units can correctly retain the data. In memory testing, fixed-pattern writing is used to detect "stuck bits"; If the values read after writing are inconsistent, it indicates a data bus failure or a damaged storage unit.
3. The testing method for memory chips according to claim 2, characterized in that, In memory testing, the fixed-pattern write procedure for detecting "stuck bits" includes: The memory is divided into 16MB blocks, and each block is tested in a loop to isolate address line errors; If the block test passes, the block belongs to the first channel, as it can be accessed normally with only a 16-bit interface. If the block test fails, it is due to a fault in the second channel, resulting in a mismatch between the read data and the expected value.
4. The testing method for memory chips according to claim 3, characterized in that, Following the step that if the block test fails due to a second channel malfunction and the read data does not match the expected value, the method further includes: If the program can run after disabling the code related to the second channel, then the error only exists in the second channel; The address space of the second channel is mapped to the first channel, causing the access code of the second channel to be redirected to the first channel, and the second memory particle communicates only through the first interface.
5. The testing method for memory chips according to claim 1, characterized in that, If the program can start normally, complete read / write tests, and terminate normally according to instructions after disabling the code related to the second channel, then the code in the program involving the second channel will be modified to correspond to the first channel. After the program is modified, read / write tests will be performed on the 16-bit memory chip. The method further includes: When channel faults are identified through data consistency but the data bus error, address bus error, or memory unit error is not distinguished, the fault is verified through a three-layer test method. First layer: Verify the continuity and transmission accuracy of the data cable individually to eliminate faults at the data bus level; The second layer: Verify the uniqueness and accuracy of address line addressing based on address line overlap detection; The third layer: simulates actual business scenarios and performs full, multi-mode read and write tests on the memory storage unit; The three-layer testing method was integrated into the process of adapting the previous 32-bit testing system to 16-bit memory chips.
Citation Information
Patent Citations
Flash memory particle debugging method, flash memory particle debugging equipment and readable storage medium
CN115509442A
Verification processing method and device for multichannel chip performance, equipment and medium
CN118607466A