A memory testing method, system, device and storage medium
By initializing the memory and repeatedly writing test data, combined with data verification of other storage units on the same row, the data jump problem caused by the difficulty of detecting memory hardware problems in the prior art is solved, and effective detection and screening of memory defective products is realized.
Patent Information
- Application Number
- CN202510113142.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art is difficult to detect data jumps caused by memory hardware problems, making it difficult to detect memory defective products.
By initializing the memory to be tested, and repeatedly writing the test data according to the preset number of writes, combined with data verification of other memory units in the same row, the memory unit is configured as a good unit or a fault unit, and finally the memory is determined as a good product or a fault memory.
Effectively detect defective products with hardware problems in the memory, reduce the subsequent troubleshooting and repair costs caused by hardware problems, and ensure the quality of memory products.
Smart Images

Figure CN119580816B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a memory testing method, system, device and storage medium. Background Art
[0002] Currently, when testing a produced memory, a read and write test is usually performed on a target storage location, thereby testing the read and write performance of the target storage location of the memory.
[0003] Data jumps caused by hardware problems of the memory itself cannot be tested, making it difficult to detect hardware problems of the produced memory and resulting in a large number of defective memories. Summary of the invention
[0004] In view of this, the purpose of the embodiments of the present invention is to provide a memory testing method, system, device and storage medium, which can test data jumps caused by memory hardware problems, find memories with hardware problems, and reduce defective memories.
[0005] In a first aspect, an embodiment of the present invention provides a memory testing method, comprising:
[0006] After initializing the memory to be tested, an initial memory is obtained, wherein data of all storage units in the initial memory are initialized to preset data;
[0007] Repeatedly write the test data into the storage unit in the nth row and the mth column of the initial memory according to a preset number of write times, and perform data verification on other storage units in the nth row to obtain a verification result each time the test data is written, where n and m both represent positive integers greater than or equal to 1, n is less than or equal to a maximum number of rows N of the storage unit, and m is less than or equal to a maximum number of columns M of the storage unit;
[0008] If the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data, the storage cells in the nth row and the mth column are configured as good cells; otherwise, the storage cells in the nth row and the mth column are configured as faulty cells;
[0009] In the case that all the storage units in the initial memory are not faulty units, the initial memory is configured as a good memory; otherwise, the initial memory is configured as a faulty memory.
[0010] In some optional embodiments, after configuring the storage unit in the nth row and the mth column as a faulty unit, the method further includes:
[0011] Acquire the storage unit in the nth row where data jump occurs;
[0012] Associating the storage unit in the nth row and the mth column with the storage unit where data jump occurs to obtain a hardware association group;
[0013] Detecting the circuits within the hardware association group to obtain a detection result;
[0014] A hardware fault type of the hardware association group is determined according to the detection result.
[0015] In some optional embodiments, determining the hardware fault type of the hardware association group according to the detection result includes:
[0016] When the detection result indicates that a short circuit or open circuit occurs in a line within the hardware association group, configuring the hardware fault type as a line fault;
[0017] When the detection result indicates that no short circuit or open circuit occurs in the line within the hardware association group, the hardware fault type is configured as a stability fault.
[0018] In some optional embodiments, the method further includes:
[0019] In the case where the hardware failure type is configured as the stability failure, replacing the first storage address of the spare unit in the initial memory with the second storage address of the storage unit where the data jump occurs, and deleting the storage address of the storage unit where the data jump occurs;
[0020] When the hardware fault type is configured as a line fault, a line image of the hardware association group is acquired through a scanning electron microscope, a fault location is determined according to the line image, and line repair is performed on the fault location.
[0021] In some optional embodiments, the setting of the preset number of write times includes:
[0022] Obtaining the historical writing times of the test data of a preset number of tested memories in the same batch when a faulty unit is detected;
[0023] Convert the historical write times of a preset number of tested memories into a coordinate system to obtain a set of coordinate points;
[0024] Obtaining the coordinate neighborhood of each coordinate point in the coordinate point set;
[0025] Calculate the average number of writes for each coordinate neighborhood, where the average number of writes represents the average value of the number of writes represented by other coordinate points in the coordinate neighborhood;
[0026] Calculate the standard deviation according to the average number of writes;
[0027] Determine the deviation coordinate point according to the average number of write times, the standard deviation and a preset standard deviation multiple;
[0028] Deleting the deviated coordinate points to obtain a plurality of standard coordinate points;
[0029] Performing linear fitting on a plurality of the standard coordinate points to obtain a fitting curve;
[0030] The number of write times corresponding to the fitting curve is configured as the preset number of write times.
[0031] In some optional embodiments, repeatedly writing the test data into the storage unit in the nth row and the mth column of the initial memory according to a preset number of write times includes:
[0032] Obtaining a storage capacity of the storage unit in the nth row and the mth column, wherein the storage capacity represents a maximum amount of data stored;
[0033] When the amount of the test data is less than or equal to the storage capacity, repeatedly writing the test data into the storage unit in the nth row and the mth column of the initial memory for a preset number of write times;
[0034] When the amount of the test data is greater than the storage capacity, the test data is split into first data and second data, and the first data is repeatedly written into the storage unit in the nth row and the mth column of the initial memory with a preset number of writes, and the amount of the first data is less than or equal to the storage capacity.
[0035] In some optional embodiments, determining the deviation coordinate point according to the average number of write times, the standard deviation, and a preset standard deviation multiple includes:
[0036] Obtaining the number of deviations according to the standard deviation and the preset standard deviation multiple;
[0037] Obtaining the average number of write times corresponding to each point in the coordinate point set;
[0038] The point where the average number of write times is greater than the number of deviation times is configured as the deviation coordinate point.
[0039] In a second aspect, an embodiment of the present invention provides a memory testing system, including:
[0040] The first module is used to obtain an initial memory after initializing the memory to be tested, wherein the data of all storage units in the initial memory are initialized to preset data;
[0041] The second module is used to repeatedly write the test data into the storage unit of the nth row and the mth column of the initial memory according to a preset number of write times, and perform data verification on other storage units of the nth row to obtain a verification result each time the test data is written, where n and m both represent positive integers greater than or equal to 1, n is less than or equal to the maximum number of rows N of the storage unit, and m is less than or equal to the maximum number of columns M of the storage unit;
[0042] A third module is configured to configure the storage cell in the nth row and the mth column as a good cell if the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data; otherwise, configure the storage cell in the nth row and the mth column as a faulty cell;
[0043] The fourth module is used to configure the initial memory as a good memory if all the storage units in the initial memory are not faulty units, and otherwise configure the initial memory as a faulty memory.
[0044] In a third aspect, an embodiment of the present invention provides a memory testing device, which is applied to a smart card, and the device includes:
[0045] at least one processor;
[0046] at least one memory for storing at least one program;
[0047] When the at least one program is executed by the at least one processor, the at least one processor implements the method as described above.
[0048] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a program executable by a processor is stored. When the program executable by the processor is executed by the processor, it is used to perform the method as described above.
[0049] The implementation of the embodiment of the present invention includes the following beneficial effects: The embodiment of the present invention provides a memory test method, including: obtaining an initial memory after initializing the memory to be tested, wherein the data of all storage cells in the initial memory are initialized to preset data; repeatedly writing the test data into the storage cells in the nth row and the mth column of the initial memory according to the preset number of writes, and performing data verification on the other storage cells in the nth row to obtain a verification result each time the test data is written, wherein n and m both represent positive integers greater than or equal to 1, n is less than or equal to the maximum number of rows N of the storage cells, and m is less than or equal to the maximum number of columns M of the storage cells; when the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data, the storage cells in the nth row and the mth column are configured as good cells, otherwise the storage cells in the nth row and the mth column are configured as faulty cells; when all the storage cells in the initial memory are not faulty cells, the initial memory is configured as a good memory, otherwise the initial memory is configured as a faulty memory. By implementing the above technical means, defective products with hardware problems in the memory can be found. This avoids serious problems such as data errors and crashes after the memory products with hardware problems are put into use, which affects the stability and reliability of the entire device or system. The present application can effectively make up for the shortcomings of existing testing methods. By comprehensively and specifically detecting the data changes of other memory cells in the same row when a certain memory cell in the entire memory writes data, data jumps caused by hardware failures can be discovered in advance, thereby screening out defective products, ensuring the quality of the memory products finally delivered for use, and reducing the subsequent troubleshooting and maintenance costs caused by hardware problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of the steps of a memory testing method provided by an embodiment of the present invention;
[0051] Figure 2 is a schematic diagram of a storage unit of a memory provided by an embodiment of the present invention;
[0052] Figure 3 is a structural block diagram of a memory test system provided by an embodiment of the present invention;
[0053] Figure 4 It is a structural block diagram of a memory testing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims or the above drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0056] An embodiment of the present invention provides a memory testing method, comprising: initializing a memory to be tested to obtain an initial memory, wherein data of all storage cells in the initial memory are initialized to preset data; repeatedly writing test data into the storage cells in the nth row and the mth column of the initial memory according to a preset number of writes, and performing data verification on other storage cells in the nth row to obtain a verification result each time the test data is written, wherein n and m both represent positive integers greater than or equal to 1, n is less than or equal to a maximum number of rows N of the storage cells, and m is less than or equal to a maximum number of columns M of the storage cells; if the verification result indicates that the storage data of other storage cells in the nth row are equal to the preset data, configuring the storage cells in the nth row and the mth column as good cells, otherwise configuring the storage cells in the nth row and the mth column as faulty cells; if all the storage cells in the initial memory are not faulty cells, configuring the initial memory as a good memory, otherwise configuring the initial memory as a faulty memory. In the technical solution of this embodiment, when a certain storage unit in the entire memory writes data, the data changes of other storage units in the same row are detected, so that the data jumps caused by the hardware problems of the memory can be tested, the memory with hardware problems can be found, and the defective memory can be reduced.
[0057] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.
[0058] like Figure 1 As shown, an embodiment of the present invention provides a memory testing method, which includes the steps shown below.
[0059] S100, initializing a memory to be tested to obtain an initial memory, wherein data of all storage units in the initial memory are initialized to preset data.
[0060] Specifically, before the memory to be tested begins to be tested, all storage units of the entire memory to be tested are filled with preset data to provide a unified initial state for subsequent targeted writing and detection. By filling with preset data, each storage unit in the memory can have an initial setting value, which is convenient for subsequent observation of changes in data of other storage units under specific write operations, that is, when the storage unit of the target address is written, if the data of the storage units of other addresses in the same row changes, it can be easily found by comparing with the preset unified initial value, so as to more accurately determine whether there is a data jump caused by a hardware problem; when new data is written to the target address, if the data of other addresses in the same row also changes, this may mean that there is crosstalk or other hardware problems between the storage units, such as line short circuit, signal interference, etc., which are difficult to find in the traditional test method that only writes data to the target position. For example, the setting of specific preset data can be unified to set the preset data to all "0" or all "1" and other simple and easy to identify and compare numerical modes, or other set data. The preset data written to each storage unit can be the same or different, and the specific setting is not limited here.
[0061] S200. Repeatedly write the test data into the storage cells in the nth row and the mth column of the initial memory according to a preset number of write times, and perform data verification on other storage cells in the nth row to obtain a verification result each time the test data is written, where n and m both represent positive integers greater than or equal to 1, n is less than or equal to the maximum number of rows N of the storage cells, and m is less than or equal to the maximum number of columns M of the storage cells.
[0062] Specifically, refer to Figure 2, according to the preset number of writes, that is, the preset number of writes, the test data is written multiple times into the storage cells of the nth row and the mth column specified in the initial memory; wherein the nth row represents any number of rows between the maximum number of rows N and the minimum number of rows 1 of the storage cell, and the mth column represents any number of columns between the maximum number of columns M and the minimum number of columns 1 of the storage cell. It should be noted that the preset number of writes is usually determined based on past experience, test standards, or the need to fully detect the performance of the memory; for example, in order to fully simulate the frequent data writing scenarios in actual use, a relatively large number of writes will be set, such as hundreds of thousands or even millions of times, so as to fully test whether the data of other storage cells in the same row of the storage cell jumps under repeated write operations. In the physical structure of the memory, the storage cells in different rows and columns have their own circuit connections and mutual association methods, and the circuits of the storage cells in the same row are associated. By writing test data to the storage cells at a specific position and detecting whether the data of other storage cells in the same row jump, it is possible to more specifically troubleshoot possible local hardware problems. Specifically, all storage units of the initialized memory are written with test data for a preset number of times, and when each storage unit is written with test data, other storage units in the same row are detected to see whether data jumps occur, and the mutual influence between the storage units is analyzed.
[0063] In some optional embodiments, after configuring the storage unit in the nth row and the mth column as a faulty unit, the method further includes: acquiring the storage unit in the nth row where data jump occurs; associating the storage unit in the nth row and the mth column with the storage unit in which data jump occurs to obtain a hardware association group; detecting the lines within the hardware association group to obtain a detection result; and determining the hardware fault type of the hardware association group based on the detection result.
[0064] Specifically, each time the test data is written to the designated storage unit of the nth row and the mth column, the data of other storage units in the nth row are verified at the same time. This is because the storage units in the same row often share some circuit lines inside the memory, and there is a potential mutual influence relationship between them. For example, when the target storage unit is written, if there are hardware problems such as short circuits and signal interference in the line, it is likely to cause abnormal changes in the data of other storage units in the same row. A specific verification algorithm or comparison mechanism is used to determine whether the data of other storage units in the same row has changed, and the judgment result is recorded as the verification result. The verification method can be to simply compare whether the current data of the storage unit is consistent with the known data (or preset data) after the last write operation, or to use more complex verification codes (such as parity checks, CRC checks, etc.) to check the integrity and accuracy of the data. Therefore, a hardware association group is obtained by associating the storage unit where data jump occurs with the storage unit where test data is being written, and detecting whether the lines between the hardware association group are short-circuited, open-circuited, etc., thereby generating a corresponding hardware fault type. The hardware fault type means that a hardware fault exists between the hardware association groups, and subsequent fault processing is performed to eliminate the fault, improve the yield rate of the memory, and avoid directly discarding the memory to cause waste and increase production costs.
[0065] In some optional embodiments, determining the hardware fault type of the hardware association group based on the detection result includes: when the detection result indicates that a short circuit or open circuit occurs in a line within the hardware association group, configuring the hardware fault type as a line fault; when the detection result indicates that a short circuit or open circuit does not occur in a line within the hardware association group, configuring the hardware fault type as a stability fault.
[0066] Specifically, when the detection results show that a short circuit or open circuit occurs in the line within the hardware association group, the hardware fault type is determined to be a line fault. A short circuit means that there is an abnormal direct connection between lines that should not be connected, resulting in abnormal increase in current, signal interference and other problems; an open circuit means that the line is interrupted, making it impossible to transmit the signal normally, affecting operations such as reading and writing data in related storage units. Whether it is a short circuit or an open circuit, it indicates that there is a hardware problem at the line level, so it is classified as a line fault.
[0067] If the test results show that the circuits within the hardware association group are not short-circuited or open-circuited, but other data anomalies occur; for example, the storage unit with data jumps in row 3 is storage unit A, and the storage unit in row 3 that is writing test data is storage unit B. During the normal reading and writing process of storage unit B, data jumps of storage unit A frequently occur, resulting in data being unable to be stably stored. This means that the problem between storage unit A and storage unit B is not the connection problem, but the stability of storage unit A or the entire hardware association group is defective. In this case, the hardware fault type is configured as a stability fault. This fault is caused by quality problems of the storage unit itself, imperfect storage control mechanism, external interference and other factors, resulting in the stability of data storage and reading and writing being affected.
[0068] By clearly configuring the fault type, the root cause of the hardware fault can be more accurately located. This has important guiding significance for subsequent maintenance, improvement, and quality control. If it is determined to be a line fault, the maintenance personnel can focus on troubleshooting line connections, wiring, and whether related electronic components are damaged, and other line-related issues; if it is determined to be a stability fault, the cause will be found from factors that affect stability, such as storage unit performance, storage control logic, and electromagnetic interference protection, to improve the efficiency of troubleshooting and resolution.
[0069] In some optional embodiments, the method further includes: when the hardware fault type is configured as the stability fault, replacing the first storage address of the backup unit in the initial memory with the second storage address of the storage unit where the data jump occurs, and deleting the storage address of the storage unit where the data jump occurs; when the hardware fault type is configured as a line fault, acquiring a line image of the hardware association group through a scanning electron microscope, determining the fault location based on the line image, and performing line repair on the fault location.
[0070] Specifically, when the hardware failure type is configured as a stability failure, it means that the data storage stability of the storage unit itself has problems, which manifests as abnormal situations such as data jumps. At this time, replacing the first storage address of the spare unit in the initial memory with the second storage address of the storage unit where the data jump occurs is a strategy to ensure the normal function of the entire memory through redundant design. The spare unit is originally in an unactivated state and is used to replace the problematic storage unit when such a failure occurs, so that the memory can continue to store and read and write data stably, and minimize the impact of the instability of individual storage units on the entire memory. Specifically, the second storage address of the storage unit where the data jump occurs is first obtained, and then the first storage address corresponding to the spare unit is configured to the logical position of the storage unit originally corresponding to the second storage address in the entire storage system through the control logic of the hardware, so that subsequent data read and write operations can correctly point to the replaced spare unit, achieve seamless connection, and ensure that the normal data processing flow is not disturbed. After the replacement of the storage unit is completed, the storage address of the storage unit where the data jump occurs is deleted. This step is mainly to prevent the system from continuing to try to operate the storage unit with stability problems, prevent data errors or interfere with the operation of other normal storage units. By removing its storage address from the storage management system, the problematic storage unit is logically isolated, further ensuring the overall stability and reliability of the storage.
[0071] The scanning electron microscope (SEM) has high-resolution imaging capabilities and can clearly present the microstructure, morphology, and connection status of the internal circuits of the hardware association group. When the hardware fault type is determined to be a line fault, the line image of the hardware association group is obtained through SEM, and it can be detected whether there is physical damage to the line, such as line breakage, abnormal connection at the short circuit point, corrosion on the surface of the line, etc., to provide a reliable basis for accurately determining the fault location. When performing scanning electron microscope testing, the hardware association group sample containing the faulty line is properly prepared to meet the observation requirements of the SEM, such as necessary cleaning, fixation, etc., to ensure that it can be clearly imaged under the microscope. Then place the sample on the sample stage of the scanning electron microscope, and obtain a high-quality line image by adjusting the scanning parameters of the electron beam, such as acceleration voltage, scanning range, etc. According to the obtained image, the integrity and connection status of the line are detected to find out the abnormal characteristics related to the fault.
[0072] Based on the line image obtained by the scanning electron microscope, the host computer equipment connected to the scanning electron microscope is used to analyze and determine the specific location of the line fault. For example, if the image shows an obvious break in the middle of a certain line section, then this location is the circuit breaker point; if it is found that there is excess conductive material between two lines that should not be connected to form a short circuit path, then this is the short circuit fault location. Accurately determining the fault location is the key prerequisite for subsequent effective repairs.
[0073] Line repair: Take appropriate line repair measures for the determined fault location. For open circuit problems, micro-welding, metal wire bridging and other technologies can be used to reconnect the disconnected lines; if it is a short circuit fault, it is necessary to carefully remove the excess conductive material at the short circuit point to restore the normal insulation state between the lines. The repair is automatically carried out through specific repair equipment. During the repair process, the repair specifications of precision electronic circuits are followed to ensure that the repaired lines can meet the requirements of normal operation of the memory in terms of electrical performance, mechanical stability, etc., while avoiding the introduction of new fault hazards due to the repair operation.
[0074] In some optional embodiments, the setting of the preset number of write times includes: obtaining the historical number of writes of the test data of a preset number of tested memories in the same batch when a faulty unit is detected; converting the historical number of writes of the preset number of tested memories into a coordinate system to obtain a set of coordinate points; obtaining the coordinate neighborhood of each coordinate point in the set of coordinate points; calculating the average number of writes for each of the coordinate neighborhoods, the average number of writes representing the average number of writes represented by other coordinate points in the coordinate neighborhood; calculating the standard deviation based on the average number of writes; determining the deviated coordinate points based on the average number of writes, the standard deviation and a preset standard deviation multiple; deleting the deviated coordinate points to obtain a plurality of standard coordinate points; performing linear fitting on the plurality of standard coordinate points to obtain a fitting curve; and configuring the preset number of writes corresponding to the fitting curve.
[0075] In some optional embodiments, determining the deviated coordinate point based on the average number of writes, the standard deviation and the preset standard deviation multiple includes: obtaining the number of deviations based on the standard deviation and the preset standard deviation multiple; obtaining the average number of writes corresponding to each point in the coordinate point set; and configuring the point whose average number of writes is greater than the deviation number as the deviated coordinate point.
[0076] Specifically, when analyzing the tested memories of the same batch, understanding the historical write times of the test data of each tested memory when a faulty unit is detected can provide basic data for subsequent exploration of the overall performance of the batch of memories, fault occurrence patterns, etc. These historical write times reflect the number of writes that different individual memories can withstand when a fault occurs in actual testing. Different write times indicate that there are differences in quality, stability, etc. between individual memories; while the memories of the same batch have similar quality and stability. The historical write times of the test data of the tested memory when a faulty unit is detected can determine the number of writes when other memories of the same batch are subsequently tested, which is the preset write times; thereby avoiding excessive write times leading to an increase in test time and improving test efficiency. The purpose of converting the historical write times data into a coordinate system to form a set of coordinate points is to present the abstract data in an intuitive geometric form, so as to facilitate the subsequent use of mathematical methods to analyze and mine the inherent laws between the data. In this coordinate system, each storage unit of the tested memory can usually be regarded as an independent coordinate point (when the storage unit writes test data, other storage units in the same row have data jumps, and the number of times the corresponding test data is written when the data jump occurs is the historical number of writes of the storage unit), and its horizontal and vertical coordinates can be set according to specific analysis requirements. For example, the horizontal coordinate is the number of the storage unit in the memory or other identification information, and the vertical coordinate is the corresponding historical number of writes. According to the selected coordinate system rules, the historical number of writes corresponding to several storage units of each tested memory is matched one by one with the corresponding identification information, and converted into specific points on the coordinate plane. Many such points together constitute a set of coordinate points.
[0077] The coordinate neighborhood of each coordinate point is determined based on the fact that in data analysis, the data around a data point often has a certain correlation or similarity with it. By defining the coordinate neighborhood, we can focus on the data around each coordinate point, and then analyze the local data characteristics and change trends, which is helpful to discover local anomalies or patterns in the data. The range of the coordinate neighborhood is determined according to a certain distance metric (such as Euclidean distance, etc.). For each coordinate point in the coordinate system, the neighborhood range is defined according to a pre-set distance threshold (for example, a circular area with a fixed length as the radius or a rectangular range, etc.) with the point as the center, and the other coordinate points falling within this range constitute the coordinate neighborhood of the coordinate point. Calculating the average number of writes for each coordinate neighborhood is to comprehensively consider the number of writes represented by all relevant coordinate points in the coordinate neighborhood, and use an average value to characterize the approximate level of writes when the memory fails in this local area; this average value can smooth out the fluctuations or errors that may exist in individual data points in the local area, and can better reflect the characteristics of the number of writes with certain commonalities. Sum the number of writes corresponding to all coordinate points in the coordinate neighborhood, and then divide it by the number of coordinate points in the coordinate neighborhood. The result is the average number of writes for the coordinate neighborhood.
[0078] The standard deviation is an important statistical indicator to measure the degree of dispersion of a set of data. By calculating the standard deviation of the average number of writes for each coordinate neighborhood, we can understand the overall distribution dispersion of these average write times, that is, the difference in the average number of writes between different coordinate neighborhoods. A larger standard deviation means that the data distribution is more dispersed, and the number of writes to the storage in different areas when a failure occurs is more different; a smaller standard deviation means that the data is relatively concentrated, and the situation in each area is more similar. According to the standard deviation calculation formula in statistics, the average number of writes for all coordinate neighborhoods is calculated. First, the square of the difference between each average number of writes and the average of all average number of writes is calculated, and these square values are summed and divided by the number of coordinate neighborhoods (or adjusted according to the specific degree of freedom requirements). Finally, the square root is taken to obtain the standard deviation.
[0079] The deviation coordinate points are determined based on the average number of writes, standard deviation, and preset standard deviation multiples, aiming to find those data points that are obviously inconsistent with the overall data distribution law and have a large degree of deviation. These deviation coordinate points may be caused by abnormal conditions during the test, special defects of individual memories, or other accidental factors. Their existence will interfere with the accurate grasp of the overall data law, so they need to be identified and eliminated. By setting a reasonable preset standard deviation multiple (for example, set to 2 times, 3 times the standard deviation, etc. based on experience or data analysis requirements), those coordinate points whose difference from the average number of writes exceeds the preset standard deviation multiple multiplied by the standard deviation are determined as deviation coordinate points.
[0080] The multiple standard coordinate points obtained after deleting the deviated coordinate points can more accurately reflect the real distribution law and internal trend of the number of test data writes of the batch of memories when a fault occurs. Linear fitting of these standard coordinate points is to find a straight line that best fits the distribution trend of these data points, so that these standard coordinate points are distributed as evenly as possible near this straight line, so as to characterize the general law of the batch of memories through this fitting curve, and provide a scientific basis for the subsequent configuration of the preset number of writes. The standard coordinate points are fitted using a mathematical linear fitting algorithm (such as the least squares method, etc.) to obtain a fitting curve. The number of writes corresponding to this fitting curve is configured as the preset number of writes, which means that when the untested memories of the same batch or similar memories are tested in the future, the operation can be performed based on this representative number of writes analyzed and summarized from the historical data, making the test process more scientific and reasonable, more in line with the actual performance characteristics of the batch of memories, and helping to more accurately detect potential faults and improve test efficiency and quality.
[0081] In some optional embodiments, repeatedly writing the test data into the storage cells in the nth row and the mth column of the initial memory according to a preset number of writes includes: obtaining the storage capacity of the storage cells in the nth row and the mth column, the storage capacity representing the maximum data storage capacity; when the amount of the test data is less than or equal to the storage capacity, repeatedly writing the test data into the storage cells in the nth row and the mth column of the initial memory with a preset number of writes; when the amount of the test data is greater than the storage capacity, splitting the test data into first data and second data, and repeatedly writing the first data into the storage cells in the nth row and the mth column of the initial memory with a preset number of writes, the amount of the first data being less than or equal to the storage capacity.
[0082] Specifically, when the amount of test data is less than or equal to the storage capacity of the storage unit, it means that the test data can be completely stored in the designated storage unit of the nth row and the mth column at one time. At this time, repeatedly writing the test data into the storage unit according to the preset number of writes can simulate the scenario of multiple and stable data writes to the storage unit within the normal storage capacity range, which is convenient for observing the performance of other storage units in the same row under repeated operations, such as whether hardware-related problems such as data jumps and storage errors will occur, and then effectively testing the quality and stability of the memory. For example, if the storage capacity of a storage unit is 1024 bytes, and the amount of test data is 512 bytes, when the preset number of writes is set to 100,000 times, the 512-byte test data can be directly written into the storage unit repeatedly according to the prescribed 100,000 times, and the data written each time can be completely stored therein, and then the subsequent data verification of other storage units in the same row and other operations can be performed to determine whether there are hidden dangers of failure.
[0083] When the amount of test data is larger than the storage capacity of the storage unit, if you try to write the entire test data directly, it will inevitably lead to the inability to store the data completely. The excess data may be lost or overwrite the data in other storage areas, destroying the data integrity of the entire storage and not conducive to accurate subsequent fault detection. Therefore, it is necessary to split the test data into the first data and the second data to ensure that the amount of data written each time is within the acceptable range of the storage unit.
[0084] The writing operation of the first data: the first data whose data volume is less than or equal to the storage capacity is repeatedly written into the storage unit of the nth row and the mth column of the initial memory with a preset number of writes. The purpose of doing this is also to simulate the processing of the storage unit when facing a large amount of data in actual use, observe the reaction of the storage unit during multiple writes and whether it will cause abnormal phenomena such as data changes in other storage units in the same row, so as to detect the hardware correlation and stability of the memory. For example, the total amount of test data is 2048 bytes, and the storage capacity of the storage unit is 1024 bytes, then the test data can be split into two 1024-byte data (first data and second data), first select the first data and write it repeatedly into the storage unit according to the preset number of writes (such as 100,000 times); then decide to delete or alternately write the second data according to the specific situation and perform corresponding detection and analysis.
[0085] S300. When the verification result indicates that the storage data of other storage cells in the nth row are equal to the preset data, configure the storage cells in the nth row and the mth column as good cells; otherwise, configure the storage cells in the nth row and the mth column as faulty cells.
[0086] Specifically, when the verification result shows that the storage data of other storage cells in the nth row is equal to the preset data, it means that in the process of performing the test data writing operation on the storage cells in the nth row and the mth column, the other storage cells in the same row are not disturbed, and the data stored therein remains in the preset data state set initially. This indicates that under the current write operation and hardware environment, there is no abnormal data change caused by hardware failure (such as line short circuit, crosstalk between storage cells, etc.) between the storage cell (the storage cell in the nth row and the mth column) and other storage cells in the same row. From this perspective, the working state of the storage cell is relatively normal, so it can be configured as a good cell. This configuration method provides a clear basis for the subsequent evaluation of the overall quality of the memory and its further use and management, and helps to screen out storage cells with reliable performance.
[0087] For example, the entire memory is pre-written with preset data of all "0", and then the test data is written to the storage cell in the 3rd row and 5th column according to the preset write times. The data of other storage cells in the 3rd row are checked each time a write is made. If the check finds that the data of other storage cells always remain all "0", then the storage cell in the 3rd row and 5th column can be determined to be a good cell, indicating that it has no adverse effects on other storage cells in the same row when performing data write operations, and its own storage function also functions normally.
[0088] On the contrary, if the verification result shows that the storage data of other storage units in the nth row is not equal to the preset data, it means that when the storage unit in the nth row and the mth column is written, the data of other storage units in the same row has changed abnormally and is no longer the preset data originally set. This situation is likely due to hardware problems in the storage unit itself, such as a fault in its internal circuit, which causes abnormal signal interference when writing data, thereby affecting the data storage of other storage units in the same row; or the storage unit with data jump has hardware defects; or the entire hardware association group (including the line of the row, storage control logic, etc.) has defects, so that the data cannot be stored stably and correctly. In view of the existence of these hardware-related factors that may cause data anomalies, it is necessary to configure this storage unit in the nth row and the mth column as a faulty unit so that it can be further troubleshooted, repaired or marked as a defective product in the future, so as to avoid more serious data errors caused by this problematic storage unit in actual use.
[0089] For example: Taking the example of pre-writing all "0" preset data, during the process of writing test data to the storage cells in the 4th row and the 7th column, it is found that the data of other storage cells in the 4th row have partially changed to "1" or other situations that do not conform to the all "0" preset. This indicates that the storage cell in the 4th row and the 7th column and / or the storage cell with data jump in the hardware association group associated with it has failed and needs to be marked as a faulty cell. Then, further detection means (such as checking the circuit, analyzing the electrical performance of the storage cell itself, etc.) can be used to determine the specific cause of the failure.
[0090] S400: If all the storage units in the initial storage unit are not faulty units, configure the initial storage unit as a good storage unit; otherwise, configure the initial storage unit as a faulty storage unit.
[0091] Specifically, after testing all storage units in the initial memory (such as repeatedly writing test data, verifying data in other storage units in the same row, etc.), it is determined that all storage units in the initial memory are not faulty units, which means that during the entire detection process, each storage unit can write and store data normally, and will not cause adverse data impact on other storage units in the same row, and there are no data anomalies caused by hardware failures. Overall, the initial memory performs well under various test indicators, its hardware structure is complete and stable, and it has reliable data storage capabilities that can meet the use requirements in actual applications. Therefore, it is configured as a good memory to ensure the quality of factory products.
[0092] For example: suppose an initial memory with 100×100 (rows×columns) memory cells is tested. According to the established test method, the corresponding test data is written to each memory cell and the data of the memory cells in the same row are verified. Finally, it is found that all the 10,000 memory cells can keep the data stable during each test, and the verification results are in line with expectations. There is no situation where any unit is judged to be a faulty unit. In this case, this initial memory is configured as a good memory and can be put into normal use.
[0093] During the detection process, as long as one storage unit in the initial memory is configured as a faulty unit, it indicates that there is a hardware problem with the memory. Even if only one storage unit fails, it may cause a series of serious consequences such as data errors and system instability in subsequent actual use. Because there is a correlation between storage units, a faulty unit may affect the operation of other normal storage units through lines and signal interference, or cause abnormal data reading and writing operations of the entire memory. Therefore, as long as there is a faulty unit, the entire initial memory needs to be configured as a faulty memory. This can prevent memories with potential problems from entering the market or application scenarios, ensure that the memories used have reliable performance, and facilitate further analysis and repair of the faulty memory, find the cause of the failure, improve the production process or product design, and improve the overall product qualification rate.
[0094] For example: For the initial memory of the above 100×100 storage units, during the detection process, if the storage unit at the 30th row and 40th column is found to be a faulty unit after verification, even if the other 9999 storage units are normal, this initial memory needs to be configured as a faulty memory according to the rules, and it needs to be processed accordingly later, such as checking whether it is a problem with the storage unit itself or there are hidden dangers of failure in the hardware circuits related to its row and column.
[0095] The implementation of the embodiment of the present invention includes the following beneficial effects: The embodiment of the present invention provides a memory test method, including: obtaining an initial memory after initializing the memory to be tested, wherein the data of all storage cells in the initial memory are initialized to preset data; repeatedly writing the test data into the storage cells in the nth row and the mth column of the initial memory according to the preset number of writes, and performing data verification on the other storage cells in the nth row to obtain a verification result each time the test data is written, wherein n and m both represent positive integers greater than or equal to 1, n is less than or equal to the maximum number of rows N of the storage cells, and m is less than or equal to the maximum number of columns M of the storage cells; when the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data, the storage cells in the nth row and the mth column are configured as good cells, otherwise the storage cells in the nth row and the mth column are configured as faulty cells; when all the storage cells in the initial memory are not faulty cells, the initial memory is configured as a good memory, otherwise the initial memory is configured as a faulty memory. By implementing the above technical means, defective products with hardware problems in the memory can be found. This avoids serious problems such as data errors and crashes after the memory products with hardware problems are put into use, which affects the stability and reliability of the entire device or system. The present application can effectively make up for the shortcomings of existing testing methods. By comprehensively and specifically detecting the data changes of other memory cells in the same row when a certain memory cell in the entire memory writes data, data jumps caused by hardware failures can be discovered in advance, thereby screening out defective products, ensuring the quality of the memory products finally delivered for use, and reducing the subsequent troubleshooting and maintenance costs caused by hardware problems.
[0096] Second, refer to Figure 3 , an embodiment of the present invention provides a memory testing system, comprising:
[0097] The first module is used to obtain an initial memory after initializing the memory to be tested, wherein the data of all storage units in the initial memory are initialized to preset data;
[0098] The second module is used to repeatedly write the test data into the storage unit of the nth row and the mth column of the initial memory according to a preset number of write times, and perform data verification on other storage units of the nth row to obtain a verification result each time the test data is written, where n and m both represent positive integers greater than or equal to 1, n is less than or equal to the maximum number of rows N of the storage unit, and m is less than or equal to the maximum number of columns M of the storage unit;
[0099] A third module is configured to configure the storage cell in the nth row and the mth column as a good cell if the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data; otherwise, configure the storage cell in the nth row and the mth column as a faulty cell;
[0100] The fourth module is used to configure the initial memory as a good memory if all the storage units in the initial memory are not faulty units, and otherwise configure the initial memory as a faulty memory.
[0101] It can be seen that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0102] Thirdly, refer to Figure 4 , an embodiment of the present invention provides a memory testing device, comprising:
[0103] at least one processor;
[0104] at least one memory for storing at least one program;
[0105] When at least one program is executed by at least one processor, the at least one processor implements the above method.
[0106] It can be seen that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0107] In a fourth aspect, in addition, the embodiments of the present application further disclose a computer program product or a computer program, which is stored in a computer storable medium. The processor of a computer device can read the computer program from a computer readable storage medium, and the processor executes the computer program so that the computer device executes the above method or the above system. Similarly, the contents in the above method embodiments are all applicable to the storage medium embodiments, and the functions specifically implemented by the storage medium embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0108] It is understood that all or some steps and systems in the disclosed method above can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital information processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media generally contain computer readable instructions, data structures, program modules, or other data in modulated data information such as carrier waves or other transport mechanisms, and may include any information delivery media.
[0109] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in the technical field without departing from the purpose of the present invention.
Claims
1. A memory testing method, characterized in that: include: After initializing the memory to be tested, an initial memory is obtained, wherein data of all storage units in the initial memory are initialized to preset data; Repeatedly write the test data into the storage unit in the nth row and the mth column of the initial memory according to a preset number of write times, and perform data verification on other storage units in the nth row to obtain a verification result each time the test data is written, where n and m both represent positive integers greater than or equal to 1, n is less than or equal to a maximum number of rows N of the storage unit, and m is less than or equal to a maximum number of columns M of the storage unit; The data verification means that when data is written to the storage unit in the nth row and the mth column, data jump detection is performed on other storage units in the nth row; If the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data, the storage cells in the nth row and the mth column are configured as good cells; otherwise, the storage cells in the nth row and the mth column are configured as faulty cells; In the case that all the storage units in the initial memory are not faulty units, the initial memory is configured as a good memory; otherwise, the initial memory is configured as a faulty memory.
2. The method according to claim 1, characterized in that After configuring the storage unit in the nth row and the mth column as a faulty unit, the method further includes: Acquire the storage unit in the nth row where data jump occurs; Associating the storage unit in the nth row and the mth column with the storage unit where data jump occurs to obtain a hardware association group; Detecting the circuits within the hardware association group to obtain a detection result; A hardware fault type of the hardware association group is determined according to the detection result.
3. The method according to claim 2, characterized in that The determining the hardware fault type of the hardware association group according to the detection result includes: When the detection result indicates that a short circuit or open circuit occurs in a line within the hardware association group, configuring the hardware fault type as a line fault; When the detection result indicates that no short circuit or open circuit occurs in the line within the hardware association group, the hardware fault type is configured as a stability fault.
4. The method according to claim 3, characterized in that The method further comprises: In the case where the hardware failure type is configured as the stability failure, replacing the first storage address of the spare unit in the initial memory with the second storage address of the storage unit where the data jump occurs, and deleting the storage address of the storage unit where the data jump occurs; When the hardware fault type is configured as a line fault, a line image of the hardware association group is acquired through a scanning electron microscope, a fault location is determined according to the line image, and line repair is performed on the fault location.
5. The method according to claim 1, characterized in that The setting of the preset number of write times includes: Obtaining the historical writing times of the test data of a preset number of tested memories in the same batch when a faulty unit is detected; Convert the historical write times of a preset number of tested memories into a coordinate system to obtain a set of coordinate points; Obtaining the coordinate neighborhood of each coordinate point in the coordinate point set; Calculate the average number of writes for each coordinate neighborhood, where the average number of writes represents the average value of the number of writes represented by other coordinate points in the coordinate neighborhood; Calculate the standard deviation according to the average number of writes; Determine the deviation coordinate point according to the average number of write times, the standard deviation and a preset standard deviation multiple; Deleting the deviated coordinate points to obtain a plurality of standard coordinate points; Performing linear fitting on a plurality of the standard coordinate points to obtain a fitting curve; The number of write times corresponding to the fitting curve is configured as the preset number of write times.
6. The method according to claim 1, characterized in that The step of repeatedly writing the test data into the storage unit in the nth row and the mth column of the initial memory according to a preset number of write times includes: Obtaining a storage capacity of the storage unit in the nth row and the mth column, wherein the storage capacity represents a maximum amount of data stored; When the amount of the test data is less than or equal to the storage capacity, repeatedly writing the test data into the storage unit in the nth row and the mth column of the initial memory for a preset number of write times; When the amount of the test data is greater than the storage capacity, the test data is split into first data and second data, and the first data is repeatedly written into the storage unit in the nth row and the mth column of the initial memory with a preset number of writes, and the amount of the first data is less than or equal to the storage capacity.
7. The method according to claim 5, characterized in that The step of determining the deviation coordinate point according to the average number of write times, the standard deviation and a preset standard deviation multiple includes: Obtaining the number of deviations according to the standard deviation and the preset standard deviation multiple; Obtaining the average number of write times corresponding to each point in the coordinate point set; The point where the average number of write times is greater than the number of deviation times is configured as the deviation coordinate point.
8. A memory testing system, characterized in that: include: The first module is used to obtain an initial memory after initializing the memory to be tested, wherein the data of all storage units in the initial memory are initialized to preset data; The second module is used to repeatedly write the test data into the storage unit of the nth row and the mth column of the initial memory according to a preset number of write times, and perform data verification on other storage units of the nth row to obtain a verification result each time the test data is written, where n and m both represent positive integers greater than or equal to 1, n is less than or equal to the maximum number of rows N of the storage unit, and m is less than or equal to the maximum number of columns M of the storage unit; The data verification means that when data is written to the storage unit in the nth row and the mth column, data jump detection is performed on other storage units in the nth row; A third module is configured to configure the storage cell in the nth row and the mth column as a good cell if the verification result indicates that the storage data of the other storage cells in the nth row are equal to the preset data; otherwise, configure the storage cell in the nth row and the mth column as a faulty cell; The fourth module is used to configure the initial memory as a good memory if all the storage units in the initial memory are not faulty units, and otherwise configure the initial memory as a faulty memory.
9. A memory testing device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Method for detecting memory
CN103700408A
Test method and test system of memory
CN117012255A
Hard disk test method and device, electronic equipment and storage medium
CN119201573A