Data loss detection method and system for storage chip
By obtaining address mapping tables in the memory chip, determining the fault coupling strength, establishing a fault propagation model and generating differentiated detection paths, the problems of long detection time, large damage and neglecting the fault propagation effect in the prior art are solved, and more efficient and accurate data loss detection is achieved.
Patent Information
- Application Number
- CN202510690728.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Data loss detection methods of existing memory chips usually adopt global traversal, resulting in long detection time and large damage, and ignore the fault propagation effect between various memory cells in the memory chip, making it impossible to accurately identify the fault chain and coupling impact.
By obtaining the address mapping table of the memory chip, the fault coupling strength between the memory cells is determined, the fault propagation model is established, the steady-state solution is calculated, and the potential fault region is determined based on the information entropy, a differentiated detection path is generated, and the detection and repair are carried out according to the path.
Improves the accuracy and efficiency of detection, accurately identifying the fault chain and coupling impact, reduces detection time and damage, and ensures more efficient, less damage and more accurate memory chip data loss detection.
Smart Images

Figure CN120216406A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of storage chips, and particularly to a method and system for detecting data loss of a storage chip. Background Art
[0002] A storage chip is a hardware component used to store data in an electronic device and is widely used in devices such as computers, mobile phones, and servers. It is mainly divided into volatile storage chips and non-volatile storage chips. Non-volatile storage chips, such as NAND Flash, EEPROM, and FRAM, can retain the stored data even after power-off and are commonly used in solid-state drives (SSDs), USB flash drives, and mobile devices. These chips save information through charge storage or other physical mechanisms and do not require power to maintain data, possessing persistence.
[0003] Non-volatile storage chips are a core component of modern electronic devices and undertake the long-term storage of critical data. However, due to reasons such as multiple erasures, manufacturing defects, and environmental factors (such as high temperature and humidity), the storage chip may malfunction, resulting in data loss or damage. For such storage chips, it is crucial to detect their faults in a timely manner. Through data loss detection, potential storage cell errors can be discovered, fault areas can be identified and repaired, thereby avoiding the loss of important data and ensuring the reliability of the system and data security, especially in fields with high reliability requirements such as data centers and medical devices.
[0004] However, existing data loss detection schemes usually adopt a global traversal method for fault detection, which causes great damage to the storage chip, has a long detection time, and often ignores the fault propagation effect between storage units in the storage chip, resulting in the inability to accurately identify the fault chain and coupling effect, thereby missing some potential fault areas and reducing the accuracy and efficiency of detection. Summary of the Invention
[0005] In view of the above deficiencies of the prior art, the purpose of the embodiments of the present invention is to provide a method for detecting data loss of a storage chip, which can solve the technical problems existing in the prior art that usually adopt a global traversal method for fault detection, causing great damage to the storage chip, having a long detection time, and often ignoring the fault propagation effect between storage units in the storage chip, resulting in the inability to accurately identify the fault chain and coupling effect, thereby missing some potential fault areas and reducing the accuracy and efficiency of detection.
[0006] In the first aspect of the embodiments of the present invention, a method for detecting data loss of a storage chip is proposed, including: S1: Obtain the address mapping table of the storage chip; S2: Combine the address mapping table to determine the fault coupling strength between two storage units in the storage chip; S3: Establish a fault propagation model between different memory cells related to the fault probability of memory cells in combination with the fault coupling strength; S4: Calculate the steady-state solution of the fault propagation model, and determine the potential fault area of the memory chip in combination with information entropy; S5: Generate a differential detection path for the potential fault area; S6: Detect the memory chip according to the differential detection path; S7: Repair the detected faulty memory cells in the order of detection; S8: Output the data loss detection result of the memory chip according to the number of faulty memory cells that failed to be repaired.
[0007] In the second aspect of the embodiments of the present invention, a data loss detection system for a memory chip is proposed, including: a processor and a memory;
[0008] The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the data loss detection method for the memory chip in the first aspect are implemented.
[0009] In the third aspect of the embodiments of the present invention, a readable storage medium is proposed. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the data loss detection method for the memory chip in the first aspect are implemented.
[0010] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: In the embodiments of the present invention, by introducing the fault coupling strength and the fault propagation model, the traditional fault detection method that only relies on global traversal in the prior art is solved, and the inefficiency problem caused by too wide a scanning range and too long a detection time is avoided. First, by obtaining the address mapping table of the memory chip and determining the fault coupling strength between memory cells, the potential influence between cells can be accurately identified, providing a basis for subsequent fault propagation modeling. Second, by calculating the steady-state solution in combination with the fault propagation model and using information entropy to determine the fault area, the high-risk area can be accurately identified, so as to generate a differential detection path targeted, avoiding blind full scanning. Finally, repair is carried out in the order of the detected faulty memory cells, and the effect of fault repair is output in combination with the detection result. This method not only improves the detection accuracy, but also greatly improves the detection efficiency, solves the problem of easy omission of fault chains and coupling effects, and ensures more efficient, less damaging and more accurate data loss detection of memory chips. Description of the Drawings
[0011] The accompanying drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of a method for detecting data loss of a storage chip provided by an embodiment of the present invention;
[0013] Figure 2 It is a schematic structural diagram of a system for detecting data loss of a storage chip provided by an embodiment of the present invention. Detailed implementation manners
[0014] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] The method for detecting data loss of a storage chip provided by an embodiment of the present invention will be described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0016] Refer to the attached drawings of the specification Figure 1 which shows a schematic flowchart of a method for detecting data loss of a storage chip provided by an embodiment of the present invention.
[0017] An embodiment of the present invention provides a method for detecting data loss of a storage chip, which may include the following steps: S1: Obtain the address mapping table of the storage chip.
[0018] Among them, the address mapping table is a table structure in the storage chip used to describe the relationship between the logical address and the physical address. In a non-volatile storage chip, the logical address is the address used by the user or the operating system to access data, while the physical address is the physical location where the storage unit actually stores data. The address mapping table is usually maintained by the controller of the storage chip. By mapping the logical address to the physical address, it ensures that data can be correctly stored and retrieved. By obtaining this mapping table, the layout and data distribution of the storage unit can be accurately identified, providing the necessary basic data for subsequent fault detection and the establishment of a fault propagation model.
[0019] S2: Determine the fault coupling strength between pairs of memory cells in the memory chip in combination with the address mapping table.
[0020] Among them, the fault coupling strength refers to the degree of fault correlation generated between two memory cells in the memory chip due to physical adjacency or logical dependence, that is, when one memory cell fails, how likely it is to affect the failure of another cell.
[0021] It should be noted that determining the fault coupling strength between memory cells in the memory chip in combination with the address mapping table can accurately identify the mutual influence between different memory cells. This method avoids the inefficiency of traditional global scanning, helps to accurately locate the fault propagation path, improves the accuracy and efficiency of fault detection, and especially effectively reduces the missed detection rate in complex fault scenarios.
[0022] In a possible implementation manner, S2 specifically includes: S201: Obtain the physical architecture information of the memory chip, where the physical architecture information includes the physical diffusion scale that describes the influence range of the unit charge in a single memory cell.
[0023] Among them, the physical diffusion scale is related to the type of memory chip production process, that is, how many memory cells can be affected by a unit charge determined in advance based on the production process.
[0024] Optionally, the physical diffusion scale can also be directly calculated in real time according to the formula. The specific calculation method is: ; Among them, represents the Boltzmann constant, represents the current temperature of the memory chip, represents the distance between adjacent memory cells, represents the thickness of the insulating layer of the memory chip, represents the pi, represents the material viscosity coefficient, represents the physical diffusion scale.
[0025] It should be noted that the material viscosity coefficient specifically represents the resistance characteristics of the memory cell insulating layer (such as SiO2, SiN, etc.) to charge diffusion. The insulating layer (such as SiO2, HfO2) is the core structure that isolates memory cells, and its thickness directly affects the charge tunneling probability.
[0026] S202: Based on the physical architecture information, determine the physical three-dimensional coordinates of each memory cell in the memory chip.
[0027] Among them, the physical architecture information refers to the data describing the internal structure and characteristics of the storage chip, mainly including the layout of storage units, the spacing between storage units, the thickness of the isolation layer, electrical characteristics, etc. It helps us understand the physical characteristics of the storage chip and its distribution within the chip. A storage unit is the basic unit in the storage chip for storing data. Each storage unit has a unique physical address indicating its actual position in the chip. The physical three-dimensional coordinates are a coordinate system used to represent the spatial position of the storage unit in the storage chip, usually represented by three dimensions: x, y, and z. These coordinates define the physical position of each storage unit within the chip, thus helping to identify the relative distance between units and their possible mutual influences. By obtaining the physical architecture information of the storage chip, the physical three-dimensional coordinates of the storage unit can be accurately determined. This information helps to understand the spatial layout of each storage unit within the chip, providing important data such as the relative position between storage units and possible electrical interference, and serving as the basis for subsequent fault propagation models and detection paths.
[0028] S203: Establish a directed acyclic graph of the storage chip based on the address mapping table. Among them, the nodes of the directed acyclic graph are the logical addresses of the storage units reflected in the address mapping table, the edges of the directed acyclic graph are the mapping relationships between different storage units in the address mapping table, and the edge attribute is the historical interaction times between two storage units.
[0029] Among them, the historical interaction times are recorded through the address mapping table and the access log system of the storage chip controller, representing the interaction frequency between logical units.
[0030] S204: Calculate the fault coupling strength according to the physical diffusion scale of the storage chip and the directed acyclic graph.
[0031] The specific calculation method of the fault coupling strength is as follows: ; Among them, represents the fault coupling strength between the i-th storage unit and the j-th storage unit, e represents the natural constant, and respectively represent the physical three-dimensional coordinates of the i-th storage unit and the j-th storage unit, represents the physical diffusion scale, represents the historical mapping times between the i-th storage unit and the j-th storage unit, that is, the historical interaction times, represents the maximum logical jump length of the storage chip, log represents the logarithmic function, represents the square of the Euclidean distance.
[0032] It should be noted that the calculation formula of the fault coupling strength takes into account the physical distance between two storage units, the number of historical interactions, and the logical jump length, which have an impact on fault propagation. The first term in the formula calculates the square of the physical distance between storage units to consider the influence of physical adjacency on fault coupling. The closer the distance, the greater the coupling strength. The second term reflects the enhancement effect of logical dependence and access times on fault propagation based on the historical interaction frequency between storage units. The higher the interaction frequency, the greater the coupling strength. The overall calculation quantifies the potential of fault propagation between storage units by combining these two terms.
[0033] Specifically, by obtaining the physical architecture information of the storage chip, the mutual influence between storage units is described in detail. First, the physical diffusion scale is calculated, which determines the influence range of a unit charge on adjacent storage units and is calculated in real time based on factors such as the chip's manufacturing process, temperature, and insulation layer thickness. Then, by obtaining the physical three-dimensional coordinates of each storage unit and constructing a directed acyclic graph (DAG) based on the address mapping table, this graph shows the logical relationship between storage units and the number of their historical interactions, reflecting the access frequency between logical units. By combining the physical diffusion scale and the directed acyclic graph, the fault coupling strength is calculated, and this parameter quantifies the potential of fault propagation between two storage units. Finally, the system can accurately identify the mutual influence and fault propagation path between storage units, improving the accuracy and efficiency of fault detection.
[0034] S3: Combine the fault coupling strength to establish a fault propagation model between different storage units related to the fault probability of storage units.
[0035] Among them, the fault probability of a storage unit refers to the possibility of a certain storage unit failing under specific conditions. It is usually calculated based on the historical behavior of the storage unit, service life, environmental factors (such as temperature, humidity), and the mutual influence with other units. In this context, the fault probability of a storage unit is closely related to the fault coupling strength. By calculating the fault probability of a storage unit, it can provide a quantitative input for the fault propagation model, helping to more accurately predict which areas may have faults, thus optimizing the detection and repair strategies. By combining the fault coupling strength to establish a fault propagation model, it can accurately describe the fault propagation path and probability between different storage units. This method can dynamically reflect the cascading effect and multi-level influence of faults, avoiding the over-simplification of traditional methods. By combining the fault propagation model with the unit fault probability, the system can identify potential fault areas and prioritize the detection of high-risk areas, thereby improving the detection accuracy, reducing the missed detection rate, and enhancing the overall detection efficiency.
[0036] In a possible implementation manner, S3 specifically includes: S301: Calculate the three-dimensional Laplace operator of the storage unit failure in the physical space to describe the physical layer diffusion intensity of the storage unit failure in the physical space.
[0037] Among them, the three-dimensional Laplace operator describes the curvature (i.e., the diffusion rate) of the storage unit failure in the physical space. The specific calculation method of the physical layer diffusion intensity is as follows: ; ; Among them, represents the physical diffusion scale, represents the three-dimensional Laplace operator, represents the physical three-dimensional coordinates of the storage unit, represents taking the partial derivative, represents the storage unit failure probability of the i-th storage unit in the storage chip, represents the physical layer diffusion intensity.
[0038] It should be noted that the diffusion intensity of the storage unit failure in the physical space is described by calculating the three-dimensional Laplace operator. The Laplace operator measures the rate of change of the failure probability in space and reflects the degree of propagation of the failure between surrounding units. The physical diffusion intensity evaluates the spatial diffusion ability of the failure by calculating the curvature of the storage unit failure probability in the three-dimensional space, helping to predict the propagation path of the failure in the storage chip.
[0039] S302: Determine the logical layer diffusion intensity of the storage unit failure in the logical space, i.e., the logical mapping table, based on the failure coupling intensity.
[0040] Among them, the specific calculation method of the logical layer diffusion intensity is as follows: ; ; Among them, represents the logical failure driving factor, represents the failure coupling intensity between the i-th storage unit and the k-th storage unit, represents the number of logical mapping operations per unit time, represents the maximum mapping operation frequency of the storage chip, represents the data retention duration, represents the storage unit failure probability describing the k-th storage unit in the storage chip, represents the logical layer diffusion intensity.
[0041] It should be noted that by combining the fault coupling strength and the fault probability of the storage unit, the diffusion strength of the logic layer is calculated. The diffusion strength of the logic layer measures the propagation of faults in the logical space, taking into account the interaction frequency between logical units and the fault driving factor. By calculating the coupling strength and fault probability between each storage unit and other units, and combining the number of logical mapping operations and the storage retention duration, the propagation effect of faults in the logical mapping table can be accurately described.
[0042] S303: Determine the self-repair strength of the storage chip in combination with the error correction ability of the storage chip. The specific formula for the self-repair strength is: ; ; Among them, represents the ECC error correction ability of the storage chip, represents the self-repair coefficient of the storage chip, represents the error correction response speed of the storage chip, represents the number of redundant blocks of the storage chip, represents the average repair delay of the storage chip, represents the self-repair strength.
[0043] Among them, the unit of the ECC error correction ability is bit / page.
[0044] It should be noted that by combining the ECC error correction ability, the number of redundant blocks and the repair delay of the storage chip, the self-repair strength of the storage chip is calculated. The self-repair strength reflects the ability of the chip to rely on ECC error correction and redundant resources for repair when a fault occurs. By evaluating the effectiveness of the chip's self-repair, the fault repair strategy can be optimized.
[0045] S304: Establish a fault propagation model by combining the physical layer diffusion strength, the logic layer diffusion strength and the self-repair strength. The specific expression of the fault propagation model is: ; Among them, t represents the time variable.
[0046] Specifically, by combining multiple factors such as the physical layer, logical layer, and self-repair strength, the propagation process of storage unit failures can be simulated more precisely. By calculating the diffusion strength of the physical layer and the diffusion strength of the logical layer, the system can capture the transfer laws of faults in the physical space and logical space. The introduction of self-repair strength further considers the error correction ability of the storage chip and the influence of redundant resources, enhancing the feasibility of fault repair. By comprehensively considering these factors, the overall model not only improves the accuracy of the fault propagation model but also effectively predicts and identifies high-risk areas, optimizes the detection path, thereby achieving more accurate and efficient fault detection and repair, and avoiding misjudgment and undetected problems in traditional methods.
[0047] S4: Calculate the steady-state solution of the fault propagation model and determine the potential fault areas of the storage chip in combination with information entropy.
[0048] Among them, the steady-state solution refers to the stable state that the system finally reaches in a dynamic system as time goes by. In the fault propagation model, the steady-state solution represents the final stable value of the fault probabilities of all storage units during the fault propagation process. The steady-state solution no longer changes with time and reflects the final state of the storage units after long-term operation or multiple fault propagations. Information entropy is a measure of information uncertainty, representing the randomness or chaos degree in the system. In the fault propagation model, information entropy is used to quantify the uncertainty of the fault probabilities of each storage unit in the storage chip. The higher the entropy value, the more uniform the fault distribution in the system, and the more difficult it is to identify the fault areas. The lower the entropy value, the more concentrated the fault areas, and the easier it is to identify. The potential fault areas refer to the areas in the storage chip where faults may occur. These areas have relatively high fault probabilities or are identified as areas with greater risks due to the calculation of the fault propagation model and information entropy. The potential fault areas are usually high-risk areas judged based on the steady-state solution and information entropy and may be key monitored and repaired in subsequent detections.
[0049] It should be noted that by calculating the steady-state solution of the fault propagation model and combining information entropy, the potential fault areas in the storage chip can be accurately identified. Through the steady-state solution, the system can predict the final state of the faults of each storage unit, while information entropy helps to evaluate the uncertainty and concentration degree of the fault areas, thereby improving the identification accuracy of the fault areas. This method can more specifically focus on high-risk areas, reduce blind detections, and improve the detection efficiency and accuracy.
[0050] In a possible implementation, S4 specifically includes: S401: Obtain the faulty storage units in the storage chip.
[0051] S402: Set the storage unit fault probabilities of the faulty storage units to 1 respectively and substitute them into the fault propagation model to update the fault propagation model.
[0052] S403: Set the updated fault propagation model to zero to obtain the steady-state solutions corresponding to the fault storage units, where each of the steady-state solutions forms a steady-state distribution.
[0053] S404: Combine the steady-state solutions and determine the discrimination threshold for the fault area of the storage chip through information entropy. The specific calculation method of the discrimination threshold for the fault area is as follows: ; ; ; where represents the discrimination threshold for the fault area, represents the total number of storage units in the storage chip, represents the information entropy quantifying the uncertainty of the fault probability of the storage unit, represents the steady-state solution corresponding to the i-th storage unit, represents the maximum information entropy related to the total number of storage units.
[0054] It should be noted that by combining the steady-state solutions and information entropy to determine the discrimination threshold for the fault area, the uncertainty of the fault probability in the storage chip can be quantified. Information entropy reflects the uniformity of the fault probability distribution. The lower the entropy value, the more concentrated the fault area and the easier it is to identify. Calculating the threshold helps to dynamically determine the discrimination criteria for high-risk areas, avoiding misjudgment or missed detection caused by traditional fixed thresholds, and improving the recognition accuracy and detection efficiency of the fault area.
[0055] S405: Retain the target steady-state solutions greater than the discrimination threshold for the fault area, and use the connected domain formed by the storage units corresponding to the target steady-state solutions as the potential fault area.
[0056] Among them, the connected domain refers to a region formed by all storage units with a fault probability greater than the discrimination threshold for the fault area in space. In such a region, there is a strong correlation or fault propagation possibility among the storage units.
[0057] Specifically, this process accurately identifies the potential fault area of the storage chip by gradually updating and calculating the fault propagation model. First, obtain the fault storage units in the storage chip, then set the fault probability of these fault units to 1 and substitute them into the fault propagation model for updating. Next, calculate the steady-state solutions through the updated model to obtain the final fault probability of each storage unit. By combining information entropy, calculate the discrimination threshold for the fault area to determine which areas have a higher fault probability. Finally, retain the steady-state solutions greater than the discrimination threshold, and define the areas where these fault units are located as potential fault areas through connected domain analysis. This process can effectively identify and focus on the high-risk areas in the storage chip, improving the accuracy and efficiency of fault detection.
[0058] S5: Generate a differential detection path for the potential fault area.
[0059] Among them, the differential detection path refers to a targeted detection route designed according to the different risk levels, fault coupling strengths, and fault propagation models of the fault area. Different from the traditional global traversal, the differential detection path preferentially scans high-risk areas and potential fault areas, reduces unnecessary detection work, optimizes the detection order, and avoids wasting time and resources.
[0060] In a possible implementation manner, the expression of the differential repair path is specifically: ; Among them, represents the differential detection path, represents the set of candidate detection paths when taking the maximum value of the function , represents the detection duration from the detection timestamp of the storage unit p to the storage unit q in the potential fault area to the current timestamp, represents the time decay factor, represents the fault coupling strength between p and q.
[0061] It should be noted that this differential repair path preferentially repairs areas with strong fault coupling and long undetected time by considering the fault coupling strength and time decay factor between storage units. This method can dynamically adjust the repair order according to the mutual influence between storage units, avoid ineffective repairs, ensure that the most critical fault areas are repaired in a timely manner, thereby improving the repair efficiency, reducing resource waste, and enhancing the reliability and repair ability of the system.
[0062] Optionally, if it is a stable environment for periodic review, the time decay factor can be set to , if it is a high-temperature / high-load scenario with fast fault diffusion, the time decay factor can be set to , if it is a cold storage environment with high long-term stability requirements, the time decay factor can be set to .
[0063] S6: Detect the storage chip according to the differential detection path.
[0064] It can be understood that by detecting the storage chip according to the differential detection path, it is ensured that areas with higher fault probabilities and greater risks are preferentially covered. Compared with the traditional full-disk scanning method, this method is more accurate and efficient, can timely identify and locate potential faults, reduce unnecessary repeated scans, improve the speed and accuracy of fault detection, and thus effectively save detection time and resources.
[0065] In a possible implementation, S6 is specifically as follows: Use the detection data to detect the storage chip according to a differentiated detection path, where the detection data includes: interleaved bit pattern data, all-zero or all-one data, de-interleaved bit pattern data, line-by-line inversion data pattern, and pseudo-random fill data.
[0066] Among them, the interleaved bit pattern data is a fixed pattern in which 0 and 1 are alternately arranged in each row or column, and is used to detect structural faults such as bit coupling, programming disturbance, bit line interference, etc., such as 01010101 and 10101010. The all-zero or all-one data is a static pattern filled with all 0s or all 1s, and is used to detect stuck-at (fixed to 0 or 1) type faults, such as 00000000 or 11111111. The de-interleaved bit pattern data is an interleaved bit pattern arranged alternately between adjacent rows, with odd and even rows inverted, forming a checkerboard-like alternating structure, such as 0101, 1010, 0101, 1010 or 10101010 and 01010101. The line-by-line inversion data pattern is that starting from the first row, the data of each row is inverted bit by bit relative to the previous row, that is, 0 becomes 1 and 1 becomes 0, such as 00000000 and 11111111. The pseudo-random fill data is data generated using a pseudo-random number generator (such as an LFSR), and the data pattern approximates the real workload with higher coverage, such as 11001100 and 10111001.
[0067] It can be understood that by using different types of detection data (such as interleaved bit patterns, all-zero or all-one data, etc.) according to the differentiated detection path, various fault types of the storage chip can be more comprehensively covered. This method can select an appropriate detection mode according to the characteristics of the fault area, improve the accuracy of fault detection, avoid ineffective scans, and effectively reduce the detection time and resource consumption, thereby improving the overall detection efficiency and fault location accuracy.
[0068] S7: Repair the detected faulty storage units in the order of detection.
[0069] It should be noted that by repairing in the detection order, the possibility of fault spread can be minimized, and the normal function of the storage chip can be gradually restored according to the severity of the fault and the repair priority, thereby improving the repair efficiency and ensuring the integrity of the data and the stability of the system.
[0070] In a possible implementation, S7 is specifically as follows: Use the ECC of the storage chip to repair the detected faulty storage units in the order of detection.
[0071] Among them, ECC (Error Correction Code) is a coding technology used to detect and repair errors in stored data. In a storage chip, ECC detects and corrects bit errors by adding redundant bits to the data. It can automatically repair some correctable errors to ensure the accuracy of the stored data. Common ECC technologies include BCH codes, LDPC codes, etc.
[0072] By using ECC to repair faulty storage units in sequence according to the detection order during the repair process, the risk of fault expansion can be minimized, and the high-priority fault areas can be repaired first. ECC can effectively repair smaller errors, avoid data loss, and at the same time reduce the impact on other parts of the system, improving the reliability and repair efficiency of the storage chip.
[0073] S8: Output the data loss detection result of the storage chip according to the number of faulty storage units that failed to be repaired.
[0074] It should be noted that the health status of the storage chip is judged by the number of faulty storage units that failed to be repaired. If the number of failed units exceeds the preset threshold, the system will determine that data loss has occurred; otherwise, it indicates that the storage chip is normal. This helps to quickly judge whether the storage chip has suffered a serious fault or data loss, ensuring the reliability of the system and data security.
[0075] In a possible implementation manner, S8 specifically includes: S801: When the number of faulty storage units exceeds the preset number of faulty storage units, determine that the storage chip has data loss; otherwise, determine that the storage chip is normal.
[0076] It should be noted that those skilled in the art can set the size of the preset number of faulty storage units according to actual needs, and the present invention does not limit this here.
[0077] Specifically, the preset number of faulty storage units can be set to the maximum number of bad blocks allowed by the storage chip controller, that is, the maximum number of faulty storage units.
[0078] S802: Output the data loss of the storage chip and the normal state of the storage chip as the data loss detection result.
[0079] It should be noted that by setting a threshold for the preset number of faulty storage units, it is possible to effectively judge whether the storage chip has suffered a serious fault or data loss. When the number of faulty units exceeds the preset threshold, the system automatically determines that data loss has occurred, thus ensuring the timely discovery of irreparable faults. This method avoids misjudgment caused by individual small faults by reasonably setting the threshold, improves the accuracy and reliability of the judgment, helps to quickly handle faults, and ensures the stable operation of the system.
[0080] In the actual application process, first, by obtaining the address mapping table of the storage chip, the relationship between storage units is determined, and the fault coupling strength between them is calculated. Then, combining this information, a fault propagation model is established to evaluate the fault probability of the storage units. By calculating the steady-state solution of the fault propagation model and combining it with information entropy, potential fault areas are identified. After generating a differentiated detection path, priority detection is carried out according to this path, and high-risk areas are given key attention. After detecting a faulty storage unit, repairs are carried out in the order of detection of the faults. Finally, the health status of the storage chip is judged according to the number of storage units with repair failures. If the number of failed units exceeds a preset threshold, it is judged as data loss. This process ensures the accurate identification of fault areas, optimizes the detection efficiency and maximally avoids data loss.
[0081] In the embodiment of the present invention, by introducing the fault coupling strength and the fault propagation model, the traditional fault detection method that only relies on global traversal in the prior art is solved, and the inefficiency problem caused by too wide a scanning range and too long a detection time is avoided. First, by obtaining the address mapping table of the storage chip and determining the fault coupling strength between storage units, the potential influence between units can be accurately identified, providing a basis for subsequent fault propagation modeling. Secondly, by calculating the steady-state solution in combination with the fault propagation model and using information entropy to determine the fault area, high-risk areas can be accurately identified, so as to generate a differentiated detection path targeted and avoid blind full scanning. Finally, repairs are carried out in the order of the detected faulty storage units, and the effect of fault repair is output in combination with the detection result. This method not only improves the detection accuracy, but also greatly improves the detection efficiency, solves the problem of easy omission of fault chains and coupling effects, and ensures more efficient, less damaging and more accurate detection of data loss in the storage chip.
[0082] Refer to the attached Figure 2 description, which shows a schematic structural diagram of a data loss detection system for a storage chip provided by an embodiment of the present invention.
[0083] An embodiment of the present invention provides a data loss detection system 20 for a storage chip, including: a processor 201 and a memory 202; The memory 202 stores programs or instructions that can be run on the processor 201. When the programs or instructions are executed by the processor 201, the steps of the above-mentioned data loss detection method for the storage chip are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not elaborate further.
[0084] It should be understood that the processor 201 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0085] It should also be understood that the memory 202 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0086] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0087] It should be understood that in various embodiments of the present invention, the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0088] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0089] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0090] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0091] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0093] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0094] The embodiment of the present invention provides a readable storage medium including: a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the data loss detection method of the above-mentioned storage chip are implemented, and the same technical effect can be achieved. To avoid repetition, the present invention will not be described in detail again.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for detecting data loss in a storage chip, characterized in that, Including: S1: Obtain the address mapping table of the storage chip; S2: Combine the address mapping table to determine the fault coupling strength between pairwise storage units in the storage chip; S3: Combine the fault coupling strength to establish a fault propagation model between different storage units related to the storage unit fault probability; S4: Calculate the steady-state solution of the fault propagation model, and combine the information entropy to determine the potential fault area of the storage chip; S5: Generate a differential detection path for the potential fault area; S6: Detect the storage chip according to the differential detection path; S7: Repair in sequence according to the detection sequence of the detected faulty storage units; S8: Output the data loss detection result of the storage chip according to the number of faulty storage units that failed to be repaired.
2. The data loss detection method for a storage chip according to claim 1, wherein The specific content of S2 includes: S201: Obtain the physical architecture information of the storage chip, where the physical architecture information includes the physical diffusion scale describing the influence range of the unit charge in a single storage unit; S202: Based on the physical architecture information, determine the physical three-dimensional coordinates of each storage unit in the storage chip; S203: Based on the address mapping table, establish a directed acyclic graph of the storage chip, where the nodes of the directed acyclic graph are the logical addresses of the storage units reflected in the address mapping table, the edges of the directed acyclic graph are the mapping relationships between different storage units in the address mapping table, and the edge attribute is the number of historical interactions between two storage units; S204: Calculate the fault coupling strength according to the physical diffusion scale of the storage chip and the directed acyclic graph.
3. The data loss detection method of the storage chip according to claim 1, characterized in that The specific content of S3 includes: S301: Calculate the three-dimensional Laplace operator of the storage unit fault in the physical space to describe the physical layer diffusion strength of the storage unit fault in the physical space; S302: Based on the fault coupling strength, determine the logical layer diffusion strength of the storage unit fault in the logical space, i.e., the logical mapping table; S303: Combine the error correction ability of the storage chip to determine the self-repair strength of the storage chip; S304: Combine the physical layer diffusion strength, the logical layer diffusion strength and the self-repair strength to establish the fault propagation model.
4. The data loss detection method for a storage chip according to claim 1, wherein The specific content of S4 includes: S401: Obtain the faulty storage units in the storage chip; S402: Respectively set the storage unit fault probability of the faulty storage units to 1 and substitute them into the fault propagation model to update the fault propagation model; S403: Let the updated fault propagation model be equal to zero to obtain the steady-state solution corresponding to the faulty storage unit, where each of the steady-state solutions forms a steady-state distribution; S404: Combine the steady-state solution to determine the fault area discrimination threshold of the storage chip through information entropy; S405: Retain the target steady-state solutions greater than the fault area discrimination threshold, and use the connected domain formed by the storage units corresponding to the target steady-state solutions as the potential fault area.
5. The data loss detection method of the storage chip according to claim 1, characterized in that The expression of the differential detection path is specifically: ; Among them, represents the differential detection path, represents the set of candidate detection paths when the function takes the maximum value , represents the detection duration from the detection timestamp of the storage unit p to the storage unit q in the potential failure area to the current timestamp, represents the time decay factor, represents the fault coupling strength between p and q.
6. The data loss detection method for a storage chip according to claim 1, wherein The specific content of S6 is: Detect the storage chip according to the differential detection path by using detection data, wherein the detection data includes: interleaved bit pattern data, zero-one data, de-interleaved bit pattern data, line-by-line inversion data pattern, and pseudo-random fill data.
7. The method for detecting data loss of a storage chip according to claim 1, wherein The specific content of S7 is as follows: Repair the detected faulty storage units in sequence according to the detection sequence by using the ECC of the storage chip.
8. The method for detecting data loss of a storage chip according to claim 1, wherein The specific content of S8 includes: S801: When the number of the faulty storage units exceeds the preset number of faulty storage units, determine that data of the storage chip is lost; otherwise, determine that the storage chip is normal. S802: Output the data loss of the storage chip and the normality of the storage chip as the data loss detection result.
9. A data loss detection system for a storage chip, characterized in that, It includes: A processor and a memory; The memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the data loss detection method of the storage chip according to any one of claims 1 to 8 are implemented.
10. A readable storage medium, characterized in that, Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by the processor, the steps of the data loss detection method of the storage chip according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
DRAM test method and device, readable storage medium and electronic equipment
CN112599178A
Energy storage equipment fault monitoring platform under remote identification
CN119891557A
Data storage device which can be controlled remotely and remote control system
US20210318831A1