A memory chip fault self-detection method and system
By using CNN in the memory chip for self-detection of faults and generating repair configuration information in combination with genetic algorithms, the problem of limited ability of BIST technology to identify complex or atypical faults is solved, and the accuracy and efficiency of fault recognition and repair are improved.
Patent Information
- Application Number
- CN202510354762.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Existing BIST technologies have limited identification capabilities in the face of complex or atypical failures, and their self-test accuracy is limited.
Convolutional neural network (CNN) is used for self-detection, identify and classify different types of failures, and dynamically generate repair configuration information in combination with genetic algorithms to adapt to complex or atypical failures.
It improves the accuracy of fault identification and chip repair capabilities, can effectively face complex or atypical failures, and improves the reliability of memory chips.
Smart Images

Figure CN119862066B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a storage chip fault self-detection method and system. Background Art
[0002] Memory chips are key components of many computing and embedded systems, and their reliability directly affects the overall stability of the system. Through fault self-checking, the system can detect potential memory chip failures at an early stage, so that corresponding measures can be taken to repair or replace them, avoiding system crashes or data loss, and significantly improving system reliability.
[0003] Currently, the main method used for fault self-checking is built-in self-test (BIST). The memory is embedded with test circuits, which can automatically execute test sequences when the chip is running to detect and locate faults.
[0004] However, the test patterns of BIST technology are usually predefined in the design stage and are used to detect these specific types of faults when the chip is running. This results in BIST technology being generally effective for predefined common faults, but limited in its ability to identify complex or atypical faults, with significant limitations and limited self-test accuracy. Summary of the invention
[0005] In order to solve the technical problem that the current BIST technology is generally effective for predefined common faults, but has limited ability to identify complex or atypical faults, has significant limitations, and has limited self-test accuracy, the present invention provides a storage chip fault self-test method and system.
[0006] The technical solution provided by the embodiment of the present invention is as follows:
[0007] First aspect:
[0008] A memory chip fault self-checking method provided by an embodiment of the present invention is applied to a memory, comprising:
[0009] S1: Obtain the number of row cells and column cells with defects, and count the defect locations;
[0010] S2: Enable the spare row and spare column to quickly repair the fault; determine whether the repair is complete; if so, end the self-check; otherwise, enter S3;
[0011] S3: Use convolutional neural network to perform self-detection to determine whether it is a typical fault; if so, enter S4; otherwise, enter S5;
[0012] S4: directly perform repair according to the repair method corresponding to the typical fault stored in the fault repair library; determine whether the repair is completed; if so, end the self-check; otherwise, enter S5;
[0013] S5: Generate repair configuration information using a genetic algorithm, download the repair configuration information, and repair the fault; determine whether the repair is complete; if so, end the self-check; otherwise, scrap the memory chip.
[0014] Second aspect:
[0015] An embodiment of the present invention provides a storage chip fault self-checking system, comprising:
[0016] processor;
[0017] A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the memory chip fault self-detection method as described in the first aspect is implemented.
[0018] The third aspect:
[0019] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method for self-checking a storage chip fault as described in the first aspect is implemented.
[0020] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0021] (1) In the present invention, CNN is used for self-detection, which can intelligently identify and classify different types of faults. With the continuous training and updating of CNN, its fault recognition ability can be continuously enhanced to adapt to new fault modes. It has good recognition ability when facing complex or atypical faults, thereby improving the accuracy of self-detection.
[0022] (2) In the present invention, repair strategies from simple to complex are tried in sequence, which not only ensures the rapid repair of simple faults, but also provides a deep repair path for complex faults. In particular, in the case of complex faults, genetic algorithms can be used to dynamically generate repair configuration information, which can adaptively respond to different fault scenarios and improve the chip repair capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A schematic diagram of a flow chart of a storage chip fault self-checking method provided by an embodiment of the present invention;
[0025] Figure 2 A schematic diagram of the structure of a storage chip fault self-checking system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0027] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0028] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0029] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0030] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0031] Reference Manual Attached Figure 1 , shows a flow chart of a storage chip fault self-detection method provided by an embodiment of the present invention.
[0032] The embodiment of the present invention provides a processing flow of a storage chip fault self-check method, which may include the following steps:
[0033] S1: Obtain the number of row cells and column cells with defects, and count the defect locations.
[0034] It should be noted that in the storage unit, by reading and comparing the stored data, it is possible to detect whether there is a bit flip error (that is, the stored bit value is inconsistent with the expected value). This process is usually implemented through memory test routines such as read-write-read tests. Faults in storage cells can also be automatically detected and located by using codes such as parity checks, Hamming codes, or more complex error detection and correction codes (ECC). These methods can not only detect faults, but also provide the specific location of the fault.
[0035] Specifically, counters can be set for row units and column units respectively, with an initial value of zero. All rows and columns of the memory are traversed, and when a fault is detected, the specific row number and column number where the fault occurs are recorded, and the counters of the corresponding rows and columns are increased accordingly. The row number and column number of each memory cell where a fault is detected are recorded to form a defect location list or matrix.
[0036] S2: Enable the spare row and spare column to quickly repair the fault. Determine whether the repair is complete. If so, end the self-test. Otherwise, enter S3.
[0037] In the present invention, by preferentially enabling spare rows and spare columns, the system can immediately repair the detected fault, can respond quickly, reduce the duration of the fault in the system, and avoid greater impact or system downtime.
[0038] In a possible implementation manner, enabling the spare row unit and the spare column unit in S2 to quickly repair the fault specifically includes sub-steps S201 and S202:
[0039] S201: Determine whether the number of defective row units is greater than the defect threshold. If so, prioritize starting the spare row to repair the fault. Otherwise, prioritize allocating the spare column to repair the fault.
[0040] Among them, those skilled in the art can set the size of the defect threshold according to actual conditions, and the present invention does not limit it.
[0041] It should be noted that by judging whether the number of defective row units exceeds the threshold, the system can reasonably allocate limited spare row and column resources. If there are more defects in a certain dimension (for example, more defects in row units), spare rows are used for repair first to ensure that the most serious faults are handled first. This strategy can maximize the use of limited hardware resources and avoid unnecessary waste of resources.
[0042] S202: When a spare row is allocated to repair a fault, it is determined whether there are remaining defective row units and column units. If so, the spare row is activated to repair the fault.
[0043] In the present invention, enabling spare rows or columns can effectively isolate the fault area and prevent the fault from spreading to other areas. This local repair method can improve the overall reliability of the system and ensure that the normal operation of other parts is not affected.
[0044] S3: Use the convolutional neural network to perform self-detection to determine whether it is a typical fault. If so, go to S4. Otherwise, go to S5.
[0045] Among them, convolutional neural network (CNN) is a deep learning model specially designed for processing data with grid structure, such as images, videos, audio and time series. The core idea of CNN is to extract local features in the input data through convolution operation and combine these features layer by layer to achieve efficient classification and recognition of input data.
[0046] It should be noted that the use of CNN for self-detection can intelligently identify and classify different types of faults. With the continuous training and updating of CNN, its fault identification ability can be continuously enhanced to adapt to new fault modes. It has good identification ability when facing complex or atypical faults, thereby improving the accuracy of self-detection.
[0047] In a possible implementation, using a convolutional neural network in S3 to determine whether it is a typical fault specifically includes sub-steps S301 to S308:
[0048] S301: constructing a storage unit state matrix according to the layout structure of the storage unit.
[0049] It should be noted that since CNN is better at processing image data, a binary state map, a frequency domain map, and a fault density map are constructed based on the unit state matrix.
[0050] S302: Setting the values of the storage cells with faults in the storage cell state matrix to 1, and setting the values of the storage cells without faults in the storage cell state matrix to 0, to construct a binary state diagram.
[0051] It should be noted that the binary state diagram is obtained by marking the faulty storage unit as 1 and the non-faulty unit as 0, thereby obtaining an intuitive fault distribution diagram. This image can directly and clearly show the fault condition of the storage unit, which is helpful for subsequent feature extraction.
[0052] S304: Perform a two-dimensional discrete Fourier transform on the binary state graph to obtain a frequency domain graph.
[0053] It should be noted that by performing a two-dimensional discrete Fourier transform (2D-DFT) on the binary state diagram, the image can be converted from the spatial domain to the frequency domain. The frequency domain diagram can reveal the periodicity and spatial frequency characteristics of the fault distribution in the storage unit, which may not be obvious in the spatial domain, but can be clearly seen in the frequency domain.
[0054] S303: Using a 5×5 detection window, count the number of storage units with faults in each detection window in the binary state diagram as the fault density at the center position, and construct a fault density map.
[0055] It should be noted that the fault density map reflects the fault concentration in each local area by counting the number of faults in the local area. This image can help CNN identify concentrated areas or hot spots of faults, for example, when large-scale faults occur in certain areas.
[0056] S304: In the input layer of the convolutional neural network, the binary state diagram, the frequency domain diagram, and the fault density diagram are input through three channels.
[0057] S305: In the three channels of the convolutional neural network, the state feature graph, the frequency domain feature graph, and the fault density feature graph of the binary state graph are independently extracted through the convolution layer and the pooling layer respectively.
[0058] Among them, the convolution layer is mainly used to extract local features from the input data. The convolution layer performs convolution operations on the input data through a set of convolution kernels (also called filters) to generate feature maps.
[0059] Among them, the pooling layer is another key layer in the convolutional neural network, which is used to reduce the spatial size of the feature map while retaining the most important features. The pooling layer usually follows the convolutional layer. It reduces the dimension of the data through downsampling operations, thereby reducing the computational complexity of the model and preventing overfitting.
[0060] In a possible implementation manner, after S305 and before S306, the following steps are further included:
[0061] S309: Introduce a channel attention mechanism to add attention weights to each channel, highlight the target area that contributes to fault detection, and suppress irrelevant areas that do not contribute to fault detection.
[0062] It should be noted that by introducing the channel attention mechanism, the network can automatically learn which features in which channels are most important for fault detection and give these channels higher weights. This means that the network can pay more attention to the feature areas that contribute to fault detection and ignore those areas that have little impact or are irrelevant to the detection results.
[0063] Optionally, S309 specifically includes:
[0064] S3091: compress the feature map:
[0065]
[0066] Among them, Y represents the feature map after compression processing, S represents the compression function, X represents the input feature map, including the state feature map, frequency domain feature map and fault density feature map, (i, j) represents the coordinates of the similarity point, H represents the height of the feature map, and W represents the width of the feature map.
[0067] It should be noted that compressing the feature map can capture the global trend of the entire feature map and help the network understand the overall characteristics of the input data instead of just focusing on local information.
[0068] S3092: Introduce the channel attention mechanism, add attention weights to each channel, and obtain the attention weight function:
[0069] E(Y)= Sigmoid [ W s ( ReLU ( W r (Y)))]
[0070] Among them, E represents the attention weight function, Y represents the feature map after compression processing, and W s represents the reconstruction of the fully connected layer, W r Represents a compressed fully connected layer, Sigmoid represents the Sigmoid activation function, and ReLU represents the ReLU activation function.
[0071] It should be noted that the input feature map Y contains the global features of each channel, but does not consider the relationship between channels. The compressed fully connected layer further reduces the dimension of the feature map Y and performs nonlinear mapping to reduce the feature dimension and enhance the expressiveness of the model by introducing nonlinear activation. While reducing the dimension, the compressed fully connected layer helps remove redundant information, allowing the network to focus more on the most important global features and prepare to provide more concise and meaningful input for the subsequent reconstruction layer. The role of the reconstructed fully connected layer is to restore the reduced-dimensional features to the original number of channels through the fully connected layer and generate weights for each channel. This step utilizes the previously compressed global features, but after reconstruction, the network can learn how to assign appropriate attention weights to each channel.
[0072] S3093: Use the attention weight function to multiply the feature map to highlight the target area that contributes to fault detection and suppress irrelevant areas that do not contribute to fault detection.
[0073] In the present invention, by multiplying the attention weights with the feature maps, the network can significantly enhance the features that contribute to fault detection. This means that key fault features are amplified, while irrelevant or noise features are suppressed, making the network more accurate and sensitive when dealing with complex or minor faults.
[0074] S306: In the fully connected layer of the convolutional neural network, the state feature map, the frequency domain feature map, and the fault density feature map are weightedly fused to obtain a fused feature map:
[0075]
[0076] Among them, H represents the fusion feature map, h 1 represents the state characteristic diagram, ω 1 Represents the weight coefficient of the state feature graph, h 2 represents the frequency domain feature map, ω 2 Represents the weight coefficient of the frequency domain feature map, h 3 represents the fault density characteristic diagram, ω 3 Represents the weight coefficient of the fault density feature map.
[0077] Among them, those skilled in the art can set the weight coefficient ω of the state characteristic diagram according to actual conditions. 1 , the weight coefficient ω of the frequency domain feature map 2 And the weight coefficient ω of the fault density feature map 3 The present invention does not limit the size.
[0078] It should be noted that the state feature map, frequency domain feature map and fault density feature map extract relevant information of the fault from different angles. By weighted fusion of these feature maps, information of different dimensions can be integrated together, so that the final fusion feature map can more comprehensively reflect the characteristics of the input data.
[0079] In the present invention, by performing weighted fusion on the state feature map, frequency domain feature map and fault density feature map in the fully connected layer of the convolutional neural network, it is possible to integrate multi-dimensional information, adaptively adjust the relative importance of features, and improve the discrimination and generalization capabilities of the model.
[0080] S307: In the classification layer of the convolutional neural network, the probability that the current fault belongs to each fault type is calculated based on the fused feature graph:
[0081] ,
[0082] ,
[0083] Among them, P represents the probability that the current fault belongs to each fault type, pi represents the probability that the current fault belongs to the i-th fault type, Softmax represents the Softmax activation function, W represents the weight matrix, H represents the fused feature map, and b represents the bias term.
[0084] S308: Determine whether the probability that the current fault belongs to a certain fault type is greater than a preset probability value. If so, determine that the current fault belongs to a typical fault, and use the fault type with the largest probability value as the fault type of the current fault.
[0085] Among them, those skilled in the art can set the size of the preset probability value according to actual conditions, and the present invention does not limit it.
[0086] In the present invention, by calculating the probability of each fault type, the network can clearly determine which category the input fault is more likely to belong to, and judge whether it is a typical fault by the size of the probability value. By selecting the fault type with the largest probability value as the final classification result, the system can make the most reasonable classification decision with the greatest possibility.
[0087] S4: Perform repair directly according to the repair method corresponding to the typical fault stored in the fault repair library. Determine whether the repair is completed. If so, end the self-check. Otherwise, enter S5.
[0088] In the present invention, the fault repair library contains pre-stored repair solutions for known typical faults. By directly applying these solutions, the system can quickly handle the faults, avoiding complex analysis and calculation processes, thereby greatly shortening the repair time.
[0089] S5: Generate repair configuration information using genetic algorithm, download the repair configuration information, and repair the fault. Determine whether the repair is complete. If so, end the self-check. Otherwise, scrap the memory chip.
[0090] Among them, Genetic Algorithm (GA) is an optimization algorithm based on natural selection and genetic mechanism. It imitates the biological evolution process and gradually optimizes the solution of the problem by simulating operations such as inheritance, mutation, selection and crossover. Genetic algorithms are widely used in complex optimization and search problems, especially when the search space of the problem is huge and it is difficult to effectively solve it through traditional methods, genetic algorithms show strong advantages.
[0091] In a possible implementation manner, generating the repair configuration information by using a genetic algorithm in S5 specifically includes sub-steps S501 to S508:
[0092] S501: Using virtual reconfigurable circuit technology, virtualize the circuit structure of the memory chip.
[0093] Among them, Virtual Reconfigurable Circuit (VRC) technology is a method of virtualizing the hardware circuit structure through software simulation, so that the circuit can be dynamically reconfigured and optimized in a virtual environment. Virtual Reconfigurable Circuit technology is a mature existing technology and will not be described in detail in the present invention.
[0094] S502: Using Cartesian genetic programming technology to encode the circuit structure of the memory chip.
[0095] Among them, Cartesian Genetic Programming (CGP) is a variant of Genetic Programming (GP), which represents the problem solution as a directed acyclic graph (DAG) instead of a traditional tree structure. CGP is widely used in the fields of evolutionary computing and automatic programming due to its flexible encoding method and efficient search capability. Cartesian Genetic Programming is a mature prior art and will not be described in detail in the present invention.
[0096] S503: Constructing a fitness function of a genetic algorithm with the goal of reducing the number of defective storage cells and the actual execution time of the circuit.
[0097] ,
[0098] Wherein, f represents the fitness function, Z represents the repair configuration information, C represents the number of defective memory cells, T represents the actual execution time of the circuit, and μ represents the weight coefficient of the number of defective memory cells.
[0099] S504: constructing an initial population using the repair configuration information stored in the fault repair library, wherein the population includes a plurality of individuals, and each individual represents a type of repair configuration information.
[0100] S505: Adopt the elite retention strategy to the initial population Q 1 The individuals in the selection process are processed.
[0101] Specifically, the individuals are arranged from large to small according to the objective function value, and the 1 / 4 individuals with the worst objective function value are replaced with new individuals to form a new population Q 2 .
[0102] In the present invention, the elite retention strategy ensures that the optimal or near-optimal solution will not be lost during the evolution process by retaining the individuals with the highest fitness values. This helps to gradually improve the overall quality of the population and ensure that the algorithm evolves towards the optimal solution.
[0103] S506: Perform crossover processing on the population after the selection processing.
[0104] Specifically, from the population Q 2 Randomly select two sets of solution vectors as parents, generate a random number, and compare the random number with the crossover probability Compare the size, if the random number is less than the crossover probability , then the parent is cross-operated to generate new individuals to form a population Q 3 , new individual y 1 ,y 2 The generation method is as follows:
[0105] ,
[0106] ,
[0107] Among them, x 1 、x 2 Represents the parent, and rand represents a random number between 0 and 1.
[0108] In the present invention, the genes of two different parents are recombined through crossover processing, and the generated new individuals contain different characteristics of the parents. This gene recombination helps to introduce diversity into the population, prevent individuals in the population from being too similar, thereby expanding the search space and increasing the possibility of finding the global optimal solution.
[0109] Optionally, the crossover probability P e Specifically:
[0110]
[0111] Among them, P e represents the crossover probability, P e,max represents the maximum crossover probability, P e,min represents the minimum crossover probability, δ represents the fitness value of the current individual, δ avg represents the average fitness value of the population, δ max Represents the maximum fitness value among individuals.
[0112] In the present invention, by adjusting the crossover probability with the change of individual fitness, the algorithm can adaptively select the appropriate crossover operation probability. Individuals with higher fitness can use a lower crossover probability to retain the current high-quality gene combination; while individuals with lower fitness use a higher crossover probability to increase the chance of generating a better solution.
[0113] S507: Perform mutation processing on the population after the crossover processing.
[0114] Specifically, from the population Q 3 Randomly select a parent x from 3, generate a random number and compare the random number with the mutation probability P m Compare the size. If the random number is less than the mutation probability P, m , then for the parent x 3 Perform mutation operation to generate new individual y 3 Replace the original individuals to form a population Q 4 , new individual y 3 The generation method is as follows:
[0115]
[0116] Among them, x max represents the individual with the highest fitness value, rand represents a random number between 0 and 1, and x 3 、x 4 、x 5 、x 6 Indicates different individuals.
[0117] In the present invention, new gene combinations can be introduced into the population through mutation operations. The introduction of such randomness helps to avoid the individuals in the population being too similar, thereby preventing the algorithm from falling into a local optimal solution. New individuals are generated based on the difference combination of the individuals with the highest fitness and other randomly selected individuals, which ensures that there is sufficient difference between the new individuals and the existing individuals.
[0118] Optionally, the mutation probability P m Specifically:
[0119]
[0120] Among them, P m represents the mutation probability, P m,max represents the maximum mutation probability, P m,min represents the minimum mutation probability, δ represents the fitness value of the current individual, δ avg represents the average fitness value of the population, δ max Represents the maximum fitness value among individuals.
[0121] In the present invention, by adjusting the mutation probability with the change of individual fitness, the algorithm can adaptively select the appropriate mutation operation probability. Individuals with higher fitness can adopt a lower mutation probability to retain the current high-quality gene combination; while individuals with lower fitness adopt a higher mutation probability to increase the chance of generating a better solution.
[0122] S508: Determine whether the maximum number of iterations has been reached. If so, output the repair configuration information of the individual representative with the largest fitness value. Otherwise, return to continue iteration.
[0123] In the present invention, the genetic algorithm, with its evolution-based search mechanism, can effectively find the approximate optimal solution in a complex, nonlinear and multi-peak search space. In the fault repair scenario of the memory chip, the configuration space of the circuit may be very large and complex. Traditional optimization methods may find it difficult to quickly find the optimal solution. The genetic algorithm can efficiently search the entire space by simulating natural evolution.
[0124] In a possible implementation manner, before S1, the process further includes S101 to S103:
[0125] S101: Construct a simulated failure.
[0126] Optionally, the simulated faults include: faults based on a destruction mechanism, faults based on a mutation mechanism, and faults based on a wiring mechanism.
[0127] Optionally, the failure of the destruction mechanism is specifically: a failure caused by modifying a signal time characteristic or a signal value in a memory.
[0128] Optionally, the failure of the mutation mechanism is specifically: a failure caused by replacing one storage component with another storage component.
[0129] Optionally, the fault of the wiring mechanism is specifically: a fault caused by incorrectly connecting one storage component to another storage component.
[0130] In the present invention, by simulating multiple types of faults, the system can be tested under different fault scenarios. This diverse testing ensures that the system can cope with faults from different sources, such as signal distortion, hardware component failure, or connection errors, thereby improving the robustness and reliability of the system.
[0131] S102: Injecting simulated faults into the memory chip.
[0132] In a possible implementation manner, S102 specifically includes:
[0133] Based on VHDL, a very large scale integrated circuit hardware description language, simulated faults are injected into memory chips.
[0134] Among them, VHSIC Hardware Description Language (VHDL) is a hardware description language used to describe and simulate the behavior and structure of digital circuits. VHSIC Hardware Description Language VHDL is a mature prior art and will not be described in detail in the present invention.
[0135] S103: According to the detection result of the simulated fault, the convolutional neural network for self-detection is trained by the gradient descent method.
[0136] Among them, the gradient descent method is an iterative algorithm for optimizing functions, which is widely used in machine learning and deep learning, especially in training neural networks. The core idea of the gradient descent method is to optimize the performance of the model by continuously adjusting parameters so that the objective function (usually the loss function) gradually approaches the minimum value.
[0137] Specifically, the loss function when constructing a convolutional neural network for fault detection is:
[0138]
[0139] Among them, L represents the loss function, y i represents the classification result of the i-th simulated fault sample, represents the classification label of the i-th simulated fault sample, and n represents the total number of simulated fault samples.
[0140] The specific method for updating the parameters of the convolutional neural network is:
[0141]
[0142] Among them, θ t+1 represents the convolutional neural network parameters at the t+1th iteration, θ t represents the convolutional neural network parameters at the tth iteration, θ t-1 represents the convolutional neural network parameters at the t-1th iteration, η t represents the adaptive learning rate at the tth iteration, Represents the loss function L on the convolutional neural network parameters θ t is the gradient of , and β represents the momentum coefficient.
[0143] In the present invention, by introducing the momentum term, the update direction not only depends on the current gradient, but also takes into account the direction of the previous update. This method helps to accelerate the convergence speed when the gradient direction is consistent, so that the parameters can be approached faster when approaching the optimal solution.
[0144] Optionally, the adaptive learning rate is calculated as:
[0145] or t = or minutes + 1 2 ( or max - or minutes )[1+ cos ( t T max π)]
[0146] Among them, η min Represents the minimum learning rate, η max Represents the maximum learning rate, T max represents the maximum number of iterations, and cos represents the cosine function.
[0147] In the present invention, as the training progresses, the learning rate gradually decreases from the maximum value to the minimum value. This means that in the early stages of training, the learning rate is large, which helps to quickly explore the solution space and avoid falling into the local optimum; in the later stages of training, the learning rate gradually decreases, which helps to refine the parameter adjustment and reduce the risk of overfitting.
[0148] In the present invention, by constructing and injecting simulated faults, the system can reproduce the types of faults that may occur in the actual environment. This enables the convolutional neural network used for self-detection to be trained on data close to the real scene, thereby improving the accuracy and practicality of fault detection.
[0149] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0150] (1) In the present invention, CNN is used for self-detection, which can intelligently identify and classify different types of faults. With the continuous training and updating of CNN, its fault recognition ability can be continuously enhanced to adapt to new fault modes. It has good recognition ability when facing complex or atypical faults, thereby improving the accuracy of self-detection.
[0151] (2) In the present invention, repair strategies from simple to complex are tried in sequence, which not only ensures the rapid repair of simple faults, but also provides a deep repair path for complex faults. In particular, in the case of complex faults, genetic algorithms can be used to dynamically generate repair configuration information, which can adaptively respond to different fault scenarios and improve the chip repair capability.
[0152] Reference Manual Attached Figure 2 , showing a structural schematic diagram of a storage chip fault self-detection system provided by the present invention.
[0153] The present invention further provides a memory chip fault self-checking system 20, which is applied to the above-mentioned memory chip fault self-checking method, comprising:
[0154] Processor 201.
[0155] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the memory chip fault self-detection method of the method embodiment is implemented.
[0156] The memory chip fault self-checking system 20 provided by the present invention can execute the above-mentioned memory chip fault self-checking method and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate on them.
[0157] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0158] (1) In the present invention, CNN is used for self-detection, which can intelligently identify and classify different types of faults. With the continuous training and updating of CNN, its fault recognition ability can be continuously enhanced to adapt to new fault modes. It has good recognition ability when facing complex or atypical faults, thereby improving the accuracy of self-detection.
[0159] (2) In the present invention, repair strategies from simple to complex are tried in sequence, which not only ensures the rapid repair of simple faults, but also provides a deep repair path for complex faults. In particular, in the case of complex faults, genetic algorithms can be used to dynamically generate repair configuration information, which can adaptively respond to different fault scenarios and improve the chip repair capability.
[0160] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0161] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0162] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0163] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0164] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0165] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0166] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0167] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0168] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0169] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0170] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0171] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0172] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the storage chip fault self-detection method as described in the method embodiment is implemented.
[0173] A computer-readable storage medium provided by the present invention can implement the steps and effects of the storage chip fault self-detection method of the above method embodiment. To avoid repetition, the present invention will not go into details.
[0174] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0175] (1) In the present invention, CNN is used for self-detection, which can intelligently identify and classify different types of faults. With the continuous training and updating of CNN, its fault recognition ability can be continuously enhanced to adapt to new fault modes. It has good recognition ability when facing complex or atypical faults, thereby improving the accuracy of self-detection.
[0176] (2) In the present invention, repair strategies from simple to complex are tried in sequence, which not only ensures the rapid repair of simple faults, but also provides a deep repair path for complex faults. In particular, in the case of complex faults, genetic algorithms can be used to dynamically generate repair configuration information, which can adaptively respond to different fault scenarios and improve the chip repair capability.
[0177] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0178] There are a few points to note:
[0179] (1) The drawings of the embodiments of the present invention only involve structures related to the embodiments of the present invention. Other structures may refer to conventional designs.
[0180] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.
[0181] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.
[0182] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A memory chip fault self-check method, characterized in that: Applied to memory, including: S1: Obtain the number of row cells and column cells with defects, and count the defect locations; S2: Enable the spare row and spare column to quickly repair the fault; determine whether the repair is complete; if so, end the self-check; otherwise, enter S3; S3: Use convolutional neural network to perform self-detection to determine whether it is a typical fault; if so, enter S4; otherwise, enter S5; S4: directly perform repair according to the repair method corresponding to the typical fault stored in the fault repair library; determine whether the repair is completed; if so, end the self-check; otherwise, enter S5; S5: Generate repair configuration information using a genetic algorithm, download the repair configuration information, and repair the fault; determine whether the repair is complete; if so, end the self-check; otherwise, scrap the memory chip; The use of a convolutional neural network in S3 to determine whether it is a typical fault specifically includes: S301: constructing a storage unit state matrix according to the layout structure of the storage unit; S302: setting the values of the storage cells with faults in the storage cell state matrix to 1, and setting the values of the storage cells without faults in the storage cell state matrix to 0, to construct a binary state diagram; S304: performing a two-dimensional discrete Fourier transform on the binary state graph to obtain a frequency domain graph; S303: using a 5×5 detection window, counting the number of storage units with faults in each detection window in the binary state diagram as the fault density at the center position, and constructing a fault density map; S304: In the input layer of the convolutional neural network, a binary state diagram, a frequency domain diagram, and a fault density diagram are input through three channels; S305: In the three channels of the convolutional neural network, respectively, a state feature map, a frequency domain feature map, and a fault density feature map of the binary state map are independently extracted through a convolution layer and a pooling layer; S306: In the fully connected layer of the convolutional neural network, weighted fusion is performed on the state feature map, the frequency domain feature map, and the fault density feature map to obtain a fused feature map: ; Among them, H represents the fusion feature map, h1 represents the state feature map, ω1 represents the weight coefficient of the state feature map, h2 represents the frequency domain feature map, ω2 represents the weight coefficient of the frequency domain feature map, h3 represents the fault density feature map, and ω3 represents the weight coefficient of the fault density feature map; S307: In the classification layer of the convolutional neural network, the probability that the current fault belongs to each fault type is calculated according to the fused feature graph: ; ; Among them, P represents the probability that the current fault belongs to each fault type, p i represents the probability that the current fault belongs to the i-th fault type, Softmax represents the Softmax activation function, W represents the weight matrix, H represents the fusion feature map, and b represents the bias term; S308: Determine whether the probability that the current fault belongs to a certain fault type is greater than a preset probability value. If so, determine that the current fault belongs to a typical fault, and use the fault type with the largest probability value as the fault type of the current fault.
2. The memory chip fault self-checking method according to claim 1, characterized in that: Prior to S1, this also included: S101: build a simulated fault; S102: injecting the simulated fault into the memory chip; S103: According to the detection result of the simulated fault, a convolutional neural network for self-detection is trained by a gradient descent method.
3. The memory chip fault self-checking method according to claim 2, characterized in that: The simulated faults include: faults based on a destruction mechanism, faults based on a mutation mechanism, and faults based on a wiring mechanism; The failure of the destruction mechanism is specifically: a failure caused by modifying the signal time characteristics or signal value in the memory; The fault of the mutation mechanism is specifically: a fault caused by replacing one storage component with another storage component; The wiring mechanism failure is specifically a failure caused by incorrectly connecting one storage component to another storage component.
4. The memory chip fault self-checking method according to claim 2, characterized in that: The S102 is specifically: Based on the very large scale integrated circuit hardware description language VHDL, the simulated fault is injected into the memory chip.
5. The memory chip fault self-checking method according to claim 1, characterized in that: The enabling of the spare row unit and the spare column unit in S2 to quickly repair the fault specifically includes: S201: Determine whether the number of defective row units is greater than a defect threshold; if so, prioritize starting the spare row to repair the fault; otherwise, prioritize allocating the spare column to repair the fault; S202: When the spare column is allocated to repair the fault, determine whether there are remaining defective row units and column units; if so, activate the spare row to repair the fault.
6. The memory chip fault self-checking method according to claim 1, characterized in that: After S305 and before S306, the following steps are also included: S309: Introduce a channel attention mechanism to add attention weights to each channel, highlight the target area that contributes to fault detection, and suppress irrelevant areas that do not contribute to fault detection.
7. The memory chip fault self-checking method according to claim 6, characterized in that: The S309 specifically includes: S3091: compress the feature map: ; Wherein, Y represents the feature map after compression processing, S represents the compression function, X represents the input feature map, including the state feature map, the frequency domain feature map and the fault density feature map, (i, j) represents the coordinates of the similarity point, H represents the height of the feature map, and W represents the width of the feature map; S3092: Introduce the channel attention mechanism, add attention weights to each channel, and obtain the attention weight function: ; Among them, E represents the attention weight function, Y represents the feature map after compression processing, and W s represents the reconstruction of the fully connected layer, W r represents a compressed fully connected layer, Sigmoid represents the Sigmoid activation function, and ReLU represents the ReLU activation function; S3093: Use the attention weight function to multiply the feature map to highlight the target area that contributes to fault detection and suppress irrelevant areas that do not contribute to fault detection.
8. The memory chip fault self-checking method according to claim 1, characterized in that: The generation of repair configuration information by using a genetic algorithm in S5 specifically includes: S501: virtualizing the circuit structure of the memory chip using virtual reconfigurable circuit technology; S502: Encode the memory chip circuit structure using Cartesian genetic programming technology; S503: constructing a fitness function of a genetic algorithm with the goal of reducing the number of defective memory cells and the actual execution time of the circuit; S504: constructing an initial population using the repair configuration information stored in the fault repair library, wherein the population includes a plurality of individuals, each of which represents a type of repair configuration information; S505: Adopt the elite retention strategy to select individuals in the initial population; S506: performing crossover processing on the population after the selection processing; S507: Perform mutation processing on the population after the crossover processing; S508: Determine whether the maximum number of iterations has been reached; if so, output the repair configuration information of the individual representative with the largest fitness value; otherwise, return to continue iteration.
9. A memory chip fault self-checking system, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the memory chip fault self-detection method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Chip defect detection method and device based on active learning
CN114155213A
Chip appearance defect automatic detection method, electronic equipment and storage medium
CN117011260A