Convolutional neural network fault tolerance method and electronic equipment

By calculating and connecting checksums in the fault tolerance layer of the convolution neural network, determining the error position in the convolution result and correcting errors, the problem that is only suitable for low error rate environments in the prior art is solved, and efficient fault tolerance in high error rate environments is achieved.

CN120104397AActive Publication Date: 2025-06-06INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510594779.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing convolutional neural network fault tolerance technology is only suitable for low error rate environments and cannot be compatible with high error rate environments.

Method used

By calculating the checksum in the fault-tolerant layer of the convolution neural network and connecting it to the end of the feature map matrix and weight matrix, the row and column checksum difference in the convolution result is calculated, and the error location is determined and error correction is performed.

Benefits of technology

It improves the fault tolerance effect of convolutional neural networks at high error rates, and improves the full-scene adaptability and fault tolerance performance of fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104397A_ABST
    Figure CN120104397A_ABST
Patent Text Reader

Abstract

The invention provides a convolutional neural network fault tolerance method and an electronic device, and relates to the technical field of computers, and the method comprises the steps: calculating a first checksum according to an input first feature map matrix, connecting to the end of the first feature map matrix to obtain a second feature map matrix, calculating a second checksum according to a first weight matrix, and obtaining a second checksum; connecting to the tail of the first weight matrix to obtain a second weight matrix, calculating a first convolution result of the second feature map matrix and the second weight matrix, determining a first row checksum, a second row checksum, a first column checksum and a second column checksum, calculating a row checksum pair difference value and a column checksum pair difference value, and calculating a second convolution result of the second feature map matrix and the second weight matrix; and determining an error position according to the row checksum pair difference value and the column checksum pair difference value, and if the in-block row index and the in-block column index detect a plurality of errors on the secondary block row and the secondary block column, setting the in-block elements which may have errors to be zero for error correction, and obtaining a second convolution result after error correction. According to the invention, the fault tolerance effect of the convolutional neural network under a high error rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a convolutional neural network fault tolerance method and electronic equipment. Background Art

[0002] Algorithm-Based Fault Tolerance (ABFT) is a method of achieving fault tolerance by introducing redundancy and error detection mechanisms at the computational algorithm level, so that the system can continue to provide correct results even if errors or failures occur during execution. The advantage of this method is its efficiency and flexibility, and it can be widely used in computing environments that require high reliability and high performance.

[0003] In the field of software testing, it is necessary to distinguish the concepts of three terms: Fault (or defect), Error, and Failure. The following definition comes from the ISO 26262 automotive functional safety standard: 1) Fault: an abnormal condition that can cause an element or an item to fail, which may include permanent, intermittent and transient faults (especially soft errors).

[0004] 2) Error: discrepancy between a computed, observed or measured value or condition, and the true, specified or theoretically correct value or condition, that is, the difference between a computed, observed or measured value and the true, specified or theoretically correct value.

[0005] 3) Failure: ermination of an intended behaviour of an element or anitem due to a fault manifestation, i.e. the termination of the intended behaviour of an element or anitem due to a fault manifestation.

[0006] The relationship between the three is: failure may lead to errors, and errors may lead to failure. In the present invention, specifically, failure refers to the change of some bits of the values ​​of some memory elements from 0 to 1 or from 1 to 0 during the calculation process of some convolutional layers of the neural network (which can be used to simulate single-particle flip errors in space environments), errors refer to changes in convolution results in the neural network caused by changes in the values ​​of some elements, and failure refers to the inconsistency between the classification results output by the neural network and the expected results.

[0007] In the research of neural network fault tolerance technology, existing methods usually design a fault tolerance test framework or a fault tolerance enhancement model, in which the fault tolerance test framework can load multiple models and evaluate the robustness and recovery ability of the model under different fault conditions. Among them, fault injection technology is widely used to simulate various possible system failures, such as data loss, model parameter damage, and computing resource anomalies. When optimizing models and designing fault tolerance mechanisms, developers use fault injection to compare the performance differences between the original model before modification and the model after the introduction of the fault tolerance mechanism under various fault conditions. By quantitatively analyzing key indicators (such as accuracy, response time, and resource consumption), developers can verify the effectiveness of the fault tolerance effect. This process not only helps developers identify potential vulnerabilities, but also ensures the reliability and stability of the final model in real application scenarios, thereby improving the model's adaptability to abnormal conditions and overall performance.

[0008] The existing convolutional neural network fault-tolerance technology is only applicable to low error rate situations, and there is an urgent need to provide a convolutional neural network fault-tolerance method that is compatible with high error rate environments. Summary of the invention

[0009] The present invention provides a convolutional neural network fault-tolerance method and an electronic device, which are used to solve the defect that the fault-tolerance method in the prior art is only applicable to low error rate environments, and to provide a convolutional neural network fault-tolerance method compatible with high error rate environments.

[0010] The present invention provides a convolutional neural network fault tolerance method, comprising the following steps: according to the input first feature map matrix Calculate the first checksum , the first checksum Connected to the first feature map matrix At the end of the second feature map matrix ;in, , The fault-tolerant layer i feature map, represents the number of feature maps input to the fault-tolerant layer, , ; According to the first weight matrix Calculate the second checksum , the second checksum Connected to the first weight matrix The second weight matrix is ​​obtained at the end of ;in, ), The number of weight matrices representing the fault-tolerant layer; ; The fault-tolerant layer j A weight matrix, ; Calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result ;in, = , Represents the convolution operation; determines the first row checksum , the second line checksum , the first column checksum and the second column checksum ;in, ; Checksum of the second line The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of ; Checksum of the second column The The elements are represented as , Represents the first convolution result Bank of China No. , column number is Secondary blocks; calculate row checksum and difference Checksum and difference of columns ;in, , ; Checksum the difference according to the row Determine the first error location , based on the column checksum and the difference Determine the second error location ; Wherein, the first error position middle, Respectively represent the secondary block column index, the row index within the block, and the column index within the block; the second error position middle, Respectively represent the secondary block row index, the row index within the block, and the column index within the block; according to the first error position And the second error position Determine the row index and column index corresponding to the block In response to learning the row index and column index in the block according to the error detection situation If multiple errors are detected on both the secondary block row and the secondary block column, the position of the element in the block where the error may occur is obtained, and the error is corrected by setting the corresponding element in the block to 0; the second convolution result after error correction is obtained; wherein the second convolution result is the first convolution result after error correction Remove the remaining part of the last row of sub-blocks and the last column of sub-blocks.

[0011] According to a convolutional neural network fault tolerance method provided by the present invention, in the first error position And the second error position Determine the row index within the block and the column index within the block After the error detection situation, the method further includes: in response to obtaining the row index and the column index in the block according to the error detection situation If only one error is detected on the secondary block row, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the column checksum to obtain the difference. In response to obtaining the row index and column index in the block according to the error detection condition, If only one error is detected on the secondary block column, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the row checksum to obtain the difference. Add together for error correction.

[0012] According to a convolutional neural network fault tolerance method provided by the present invention, the first error position And the second error position Determine the row index and column index corresponding to the block The error detection condition comprises: according to the first error position Building the first dictionary ; wherein the first dictionary The key is the row index within the block and the column index within the block , the first dictionary The key value is the secondary block column index According to the second error position Create a second dictionary ; wherein the second dictionary The key is the row index within the block and the column index within the block , the second dictionary The key value is the secondary block row index ; In response to looping through the first dictionary and the second dictionary Obtain the row index and column index in the block The corresponding secondary block row index unique, then determine the row index and column index in the block Only one error is detected on the secondary block line; in response to traversing the first dictionary by looping and the second dictionary Obtain the row index and column index in the block The corresponding secondary block column index unique, then determine the row index and column index in the block Only one error is detected on the secondary block column; in response to traversing the first dictionary by looping and the second dictionary Obtain the row index and column index in the block The corresponding secondary block row index and secondary block column index If both are not unique, the row index and column index in the block are determined. Multiple errors are detected on both sub-block rows and sub-block columns.

[0013] According to a convolutional neural network fault tolerance method provided by the present invention, in the first feature map matrix according to the input Calculate the first checksum Previously, the method further includes: for the batch normalized fault-tolerant layer input, obtaining the first feature map matrix The feature values ​​that are not within a preset error range of the mean of the batch normalization result are identified, and the feature values ​​are truncated to preset values ​​within the preset error range.

[0014] According to a convolutional neural network fault tolerance method provided by the present invention, the batch normalization result approximately satisfies the normal distribution, and the preset error range is , , Respectively represent the mean and standard deviation of the batch normalization results, represents a preset integer; when the eigenvalue is less than the mean of the batch normalization result, the preset value is , when the feature value is greater than the mean of the batch normalization result, the preset value is ; Alternatively, the preset value is uniformly set to the mean of the batch normalization results.

[0015] According to a convolutional neural network fault tolerance method provided by the present invention, in the step of checking and verifying the difference between the rows Determine the first error location Before, the method further includes: checking the row and comparing the difference In response to the existence of at least one element whose value is not in range, then perform the checksum based on the row and the difference Determine the first error location action; among them, , Respectively represent the pre-obtained row checksum and difference The mean and standard deviation of the elements in , Represents a preset integer; in the difference value according to the column checksum Determine the second error location Before, the method further includes: checking the column and comparing the difference In response to the existence of at least one element whose value is not in If the column checksum is within the range, the difference is checked. Determine the second error location action; among them, , Represents the pre-obtained column checksum and difference value respectively The mean and standard deviation of the elements in , Represents a preset integer.

[0016] According to a convolutional neural network fault tolerance method provided by the present invention, in the checking of the rows and the difference Before traversing the elements in , the method further includes: replacing all convolutional layers of the convolutional neural network with the fault-tolerant layer, and setting a fault-tolerant control switch for each of the fault-tolerant layers to control whether to enable the fault-tolerant function; under error-free conditions, using the replaced convolutional neural network to perform reasoning on the test set data, and calculating the row checksum difference The mean of the elements in and standard deviation , and calculate the column checksum difference The mean of the elements in and standard deviation .

[0017] According to a convolutional neural network fault-tolerance method provided by the present invention, before the convolutional neural network receives input data for inference, the method also includes: controlling the turning on or off of the fault-tolerant control switches of each of the fault-tolerant layers according to the number of fault-tolerant layers with fault-tolerant functions turned on and the planning results of the distribution of the fault-tolerant layers with fault-tolerant functions turned on; wherein the planning results of the distribution of the fault-tolerant layers include giving priority to turning on the fault-tolerant control switches of the fault-tolerant layers with larger convolution kernel sizes; for fault-tolerant layers with the same priority, the fault-tolerant control switches are evenly turned on according to the network depth.

[0018] According to a convolutional neural network fault-tolerance method provided by the present invention, after obtaining the second convolution result after error correction, the method also includes: in response to reaching the cycle time for execution of the test phase, using the test set to perform inference and calculate the accuracy, and calculating the bit error rate based on the accuracy; in response to the bit error rate being greater than a first threshold, turning on the fault-tolerance control switches of more fault-tolerance layers; in response to the bit error rate being less than a second threshold, closing the fault-tolerance control switches of more fault-tolerance layers; wherein the second threshold is less than the first threshold; in response to the bit error rate being less than or equal to the first threshold and greater than or equal to the second threshold, keeping the setting of the fault-tolerance control switch of the fault-tolerance layer unchanged.

[0019] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a convolutional neural network fault tolerance method as described above is implemented.

[0020] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the convolutional neural network fault-tolerance methods described above.

[0021] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the convolutional neural network fault tolerance method described above is implemented.

[0022] The convolutional neural network fault tolerance method and electronic device provided by the present invention are based on the input first feature map matrix Calculate the first checksum , the first checksum Connect to the first feature map matrix At the end of the second feature map matrix , according to the first weight matrix Calculate the second checksum , the second checksum Connect to the first weight matrix The second weight matrix is ​​obtained at the end of , calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result , determine the first line checksum , the second line checksum , the first column checksum and the second column checksum , calculate the row checksum difference Checksum and difference of columns , based on the row checksum and difference Determine the first error location , based on the column checksum and the difference Determine the second error location , according to the first error position and the second error position Determine the row index and column index corresponding to the block In response to obtaining the row index and column index of the block according to the error detection situation If multiple errors are detected on both the secondary block rows and secondary block columns, the positions of the elements in the blocks where the errors may occur are obtained, and the errors are corrected by setting the corresponding elements in the blocks to 0, and the second convolution result after error correction is obtained, which improves the fault tolerance effect of the convolutional neural network under high error rates, thereby improving the full-scene adaptability and fault tolerance performance of the convolutional neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 This is one of the flow charts of the convolutional neural network fault tolerance method provided by the present invention.

[0025] Figure 2 This is one of the principle schematic diagrams of the convolutional neural network fault tolerance method provided by the present invention.

[0026] Figure 3 This is the second principle schematic diagram of the convolutional neural network fault tolerance method provided by the present invention.

[0027] Figure 4 This is the second flow chart of the convolutional neural network fault-tolerance method provided by the present invention.

[0028] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] Figure 1 It is one of the flow diagrams of the convolutional neural network fault-tolerant method provided by the present invention. The convolutional neural network receives input data for inference. During forward propagation, the fault-tolerant method is executed once each time a fault-tolerant layer with a fault-tolerant function turned on is passed. The fault-tolerant layer is a convolutional layer protected by the fault-tolerant method. If the fault-tolerant function is turned on, the fault-tolerant method can be executed and play a protective role. Figure 1 As shown, the fault tolerance method includes: Step S1: According to the input first feature map matrix Calculate the first checksum , the first checksum Connected to the first feature map matrix At the end of the second feature map matrix ;in, , The fault-tolerant layer i feature map, represents the number of feature maps input to the fault-tolerant layer, , .

[0031] The present invention utilizes the distributive law of convolution operation to generate the checksum. Specifically, there is the following relationship: , (1) The same ones are: , (2) Figure 2 This is one of the principle schematic diagrams of the convolutional neural network fault tolerance method provided by the present invention. Figure 2 As shown, the secondary block , whose size is recorded as R×R, R can be calculated by inputting feature map size H, convolution kernel size F, padding P, and stride S, and the formula is:

[0032] In the fault-tolerant method of the present invention, a first checksum is added , which is the bracket part in the right half of formula (1). And the first checksum Connect to the first feature map matrix At the end of the second feature map matrix .

[0033] Step S2: According to the first weight matrix Calculate the second checksum , the second checksum Connected to the first weight matrix The second weight matrix is ​​obtained at the end of ;in, ), The number of weight matrices representing the fault-tolerant layer; ; The fault-tolerant layer j A weight matrix, .

[0034] In the fault-tolerant method of the present invention, a second checksum is added , which is the bracket part in the right half of formula (2). And the second checksum Connect to the first weight matrix The second weight matrix is ​​obtained at the end of .

[0035] The execution order of step S1 and step S2 can be interchanged, or they can be executed simultaneously.

[0036] Step S3: Calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result ;in, = , Represents a convolution operation.

[0037] Step S4: Determine the checksum of the first row , the second line checksum , the first column checksum and the second column checksum ;in, ; Checksum of the second line The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of ; Checksum of the second column The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block.

[0038] From the first convolution result Get the first line checksum , which is the result of the first convolution Remove the last row of sub-blocks from the last column of sub-blocks. Or, based on Calculate the checksum of the first row .according to Calculated Each element in , thus obtaining the second line checksum The second line checksum Each element of is the result of the first convolution After removing the last column of sub-blocks and the last row of sub-blocks, the sub-blocks are summed up in rows.

[0039] From the first convolution result Get the checksum of the first column , that is Remove the last column of sub-blocks from the last row of sub-blocks. Or based on Calculate the checksum of the first column .according to Calculated Each element in , thereby obtaining the checksum of the second column The second column is the checksum Each element of is the result of the first convolution After removing the last column of sub-blocks and the last row of sub-blocks, the sub-blocks are summed up in columns.

[0040] Step S5: Calculate the row checksum difference Checksum and difference of columns ;in, , .

[0041] First line checksum and the second line checksum Construct a row checksum pair, the first column checksum and the second column checksum Construct a column checksum pair. If no error occurs, there should be as well as Established, so by calculating the row checksum difference Checksum and difference of columns ,Will , and , This comparison allows soft errors to be detected, located, and, under certain conditions, corrected.

[0042] Step S6: Checksum the difference according to the row Determine the first error location , based on the column checksum and the difference Determine the second error location ; Wherein, the first error position middle, Respectively represent the secondary block column index, the row index within the block, and the column index within the block; the second error position middle, They represent the secondary block row index, the row index within the block, and the column index within the block respectively.

[0043] In a high error rate environment, multiple errors are likely to occur in a single convolution layer, which may cause the first convolution result There are multiple single elements in different rows and columns or multiple elements in entire rows and columns that are damaged.

[0044] Record the first convolution result The index of an element in is ( ),in and It is a secondary block The index of and Is the element in the secondary block The index in the row. Then when the row checksum is and When there is a difference, you can get the wrong secondary block Column index of and the index of the element within the sub-block and , similarly when the column checksum is and When there is a difference, you can get the wrong secondary block The row index of and the index of the element within the sub-block and .

[0045] Step S7: according to the first error position And the second error position Determine the row index and column index corresponding to the block error detection situation.

[0046] You can use the index corresponding to the element in the same sub-block and The first error position Middle secondary block column index The number of elements in the same sub-block, and the index of the element and The second error position Middle secondary block row index The number of blocks corresponding to the row index and column index in the block error detection situation.

[0047] Step S8: In response to obtaining the row index and column index in the block according to the error detection condition, If multiple errors are detected on both the secondary block rows and the secondary block columns, the positions of the elements in the blocks where the errors may occur are obtained, and the errors are corrected by setting the corresponding elements in the blocks to 0.

[0048] Elements with different indices in a secondary block are independent of each other in terms of checksums, so multiple soft errors at different locations in a secondary block can be detected and corrected. When correcting errors, just add the difference in the checksum back. When errors at the same location occur in multiple secondary blocks, there are three situations: 1) Errors with the same position in multiple blocks occur in secondary blocks in the same row but different columns; 2) Errors with the same position in multiple blocks occur in secondary blocks in different rows and the same column; 3) Errors with the same position in multiple blocks occur in secondary blocks at different rows and columns.

[0049] In the first two cases, the error will accumulate in the same position in the row checksum or column checksum, but still in different positions in the column checksum or row checksum, so it can still be detected and corrected normally. But in the third case, the error accumulates in the same position in both the row checksum and the column checksum, and two problems will arise: First, it is impossible to distinguish the actual wrong secondary block index Figure 3 This is the second schematic diagram of the principle of the convolutional neural network fault tolerance method provided by the present invention. , when the row checksum detects two (or more) error positions and , and the column checksum pair also detects two (or more) error locations as and ,Right now Figure 3 In the first line of the checksum and the first column checksum In the same block at index (3,3), the first row checksum and the first column checksum Two errors are detected. Then the first convolution result is There are 4 possible error positions in the corresponding Figure 3For the part marked with 1, 2, 3, and 4, taking the case where 2 errors are detected in both rows and columns as an example, there are 7 error situations: ① Block 1 and 4 errors; ② Block 2 and 3 errors; ③ Block 1, 2, 4 errors; ④ Block 1, 3, 4 errors; ⑤ Block 1, 2, 3 errors; ⑥ Block 2, 3, 4 errors; ⑦ Block 1, 2, 3, 4 errors.

[0050] This situation is more likely to occur when the batch size is larger, the feature map size is smaller, and the number of channels is larger. In common convolutional neural networks, the later convolutional layers need to capture higher-level and more abstract features, so the number of channels is increased to encode more information. At the same time, due to the pooling operation or convolution with a stride greater than 1, the size of the feature map continues to decrease, so this situation must be considered.

[0051] If an error causes errors in at most one row or column of elements, these errors will be accumulated in the first row checksum or the first column checksum at different positions, and the error can be simply added back to the original value to achieve error correction. When multiple errors are accumulated in the first row checksum and the first column checksum at the same time, obtaining the correct value actually becomes a linear algebra problem: extract these errors to form a matrix separately, assuming that the column checksum detects N and the row checksum detects M, that is, the sum of the elements of each row and column of the matrix with N rows and M columns is known, and the value of each element in the matrix is ​​required. A necessary but not sufficient condition for the problem to have a solution is N=M, which is not necessarily satisfied. Even if it is satisfied, there may be multiple solutions, so the original value may not be accurately calculated. In addition, when the error rate is high, the error value will expand to inf (infinity) after a small number of layers, and the original value cannot be estimated at this time.

[0052] Since it is found through experiments that there is no need to accurately restore the original value, the present invention adopts the method of setting zero to correct the error, that is, after obtaining the row index and column index in the block according to the error detection situation When multiple errors are detected on both the secondary block rows and the secondary block columns, the positions of the elements in the blocks where the errors may occur are obtained, and the errors are corrected by setting the corresponding elements in the blocks to 0. The network inference accuracy can also be well restored by suppressing the propagation and accumulation of error values.

[0053] Step S9, obtaining a second convolution result after error correction; wherein the second convolution result is the first convolution result after error correction Remove the remaining part of the last row of sub-blocks and the last column of sub-blocks.

[0054] After error correction is completed, the first convolution result The remaining part after removing the last row of secondary blocks and the last column of secondary blocks is the final second convolution result after error correction.

[0055] The convolutional neural network fault tolerance method provided by the present invention is based on the first feature map matrix input Calculate the first checksum , the first checksum Connect to the first feature map matrix At the end of the second feature map matrix , according to the first weight matrix Calculate the second checksum , the second checksum Connect to the first weight matrix The second weight matrix is ​​obtained at the end of , calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result , determine the first line checksum , the second line checksum , the first column checksum and the second column checksum , calculate the row checksum difference Checksum and difference of columns , based on the row checksum and difference Determine the first error location , based on the column checksum and the difference Determine the second error location , according to the first error position and the second error position Determine the row index and column index corresponding to the block In response to obtaining the row index and column index of the block according to the error detection situation If multiple errors are detected on both the secondary block rows and secondary block columns, the positions of the elements in the blocks where the errors may occur are obtained, and the errors are corrected by setting the corresponding elements in the blocks to 0, and the second convolution result after error correction is obtained, which improves the fault tolerance effect of the convolutional neural network under high error rates, thereby improving the full-scene adaptability and fault tolerance performance of the convolutional neural network.

[0056] According to a convolutional neural network fault tolerance method provided by the present invention, in the first error position And the second error position Determine the row index within the block and the column index within the block After the error detection situation, the method further includes: in response to obtaining the row index and the column index in the block according to the error detection situation If only one error is detected on the secondary block row, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the column checksum to obtain the difference. In response to obtaining the row index and column index in the block according to the error detection condition, If only one error is detected on the secondary block column, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the row checksum to obtain the difference. Add together for error correction.

[0057] If the row index and column index in the block are known according to the error detection situation Only one error is detected on the secondary block row, and one or more errors can be detected on the secondary block column, which can be checked by column checksum and difference Specifically, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the column checksum to obtain the difference value. Perform error correction.

[0058] If the row index and column index in the block are known according to the error detection situation Only one error is detected on the secondary block column, and one or more errors can be detected on the secondary block row, which can be checked by row checksum and difference. Specifically, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the row checksum to obtain the difference. Add together for error correction.

[0059] The convolutional neural network fault tolerance method provided by the present invention obtains the row index and column index of the block in response to the error detection situation. If only one error is detected on the secondary block row, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the column checksum. Perform error correction, in response to obtaining the row index and column index in the block according to the error detection situation If only one error is detected on the secondary block column, the value of the secondary block element where the error is detected is obtained, and the difference between the secondary block element value and the row checksum is calculated. Error correction is performed by adding the numbers together, which improves the simplicity and efficiency of error correction under low error rates.

[0060] According to a convolutional neural network fault tolerance method provided by the present invention, the first error position And the second error position Determine the row index and column index corresponding to the block The error detection condition comprises: according to the first error position Building the first dictionary ; wherein the first dictionary The key is the row index within the block and the column index within the block , the first dictionary The key value is the secondary block column index According to the second error position Create a second dictionary ; wherein the second dictionary The key is the row index within the block and the column index within the block , the second dictionary The key value is the secondary block row index ; In response to looping through the first dictionary and the second dictionary Obtain the row index and column index in the block The corresponding secondary block row index unique, then determine the row index and column index in the block Only one error is detected on the secondary block line; in response to traversing the first dictionary by looping and the second dictionary Obtain the row index and column index in the block The corresponding secondary block column index unique, then determine the row index and column index in the block Only one error is detected on the secondary block column; in response to traversing the first dictionary by looping and the second dictionary Obtain the row index and column index in the block The corresponding secondary block row index and secondary block column index If both are not unique, the row index and column index in the block are determined. Multiple errors are detected on both sub-block rows and sub-block columns.

[0061] According to the first error position and the second error position Determine the row index within the block and the column index within the block When the error detection situation occurs, according to the first error location Building the first dictionary , the first dictionary The key is the row index within the block and the column index within the block , the first dictionary The key value is the secondary block column index According to the second error location Create a second dictionary , the second dictionary The key is the row index within the block and the column index within the block , the second dictionary The key value is the secondary block row index .

[0062] If you loop through the first dictionary and the second dictionary Get the row index and column index within the block The corresponding secondary block row index unique, then determine the row index and column index within the block If only one error is detected on a secondary block row, the difference can be added back directly for error correction.

[0063] If you loop through the first dictionary and the second dictionary Get the row index and column index within the block The corresponding secondary block column index unique, then determine the row index and column index within the block If only one error is detected on the secondary block column, the difference can be directly added back for error correction.

[0064] If you loop through the first dictionary and the second dictionary Get the row index and column index within the block The corresponding secondary block row index and secondary block column index If both are not unique, determine the row index and column index within the block. Multiple errors were detected on both subblock rows and subblock columns. and , traverse all its combinations, all possible error positions , , ,……, The corresponding elements in the block are set to 0.

[0065] The convolutional neural network fault tolerance method provided by the present invention establishes a first dictionary and the second dictionary , by looping through the first dictionary and the second dictionary Determine the row index and column index corresponding to the block Improved error detection for row index within block and column index within block The acquisition efficiency of the error detection situation.

[0066] According to a convolutional neural network fault tolerance method provided by the present invention, in the first feature map matrix according to the input Calculate the first checksum Previously, the method further includes: for the batch normalized fault-tolerant layer input, obtaining the first feature map matrix The feature values ​​that are not within a preset error range of the mean of the batch normalization result are identified, and the feature values ​​are truncated to preset values ​​within the preset error range.

[0067] The present invention utilizes the characteristics of resnet and other networks to perform batch normalization before the convolution layer, and first numerically truncates the convolution layer input according to its parameters to suppress the propagation and accumulation of errors in the calculation process. The batch processing process can include calculating the mean and variance of the features in each mini-batch of data, and using these statistics to standardize the input to have zero mean and unit variance.

[0068] The preset error range can be set according to the mean and standard deviation of the batch normalization results. When performing numerical truncation, for the batch normalized fault-tolerant layer input, the first feature map matrix is ​​obtained The feature values ​​that are not within the preset error range of the mean of the batch normalization result are truncated to the preset value within the preset error range.

[0069] The present invention can also be applied to other networks that use batch normalization after the convolution layer, such as Inception-v2 / v3, DenseNet, MobileNet, ShuffleNet, EfficientNet, etc.

[0070] The convolutional neural network fault tolerance method provided by the present invention obtains the first feature map matrix by inputting the batch normalized fault tolerance layer The eigenvalues ​​that are not within the preset error range of the mean of the batch normalization result are truncated to the preset values ​​within the preset error range, which reduces the impact of error propagation with large absolute values ​​and further improves the system fault tolerance performance.

[0071] According to a convolutional neural network fault tolerance method provided by the present invention, the batch normalization result approximately satisfies the normal distribution, and the preset error range is , , Respectively represent the mean and standard deviation of the batch normalization results, Represents a preset integer; When the eigenvalue is less than the mean of the batch normalization result, the default value is , when the eigenvalue is greater than the mean of the batch normalization result, the default value is ; Alternatively, the preset value is uniformly set to the mean of the batch normalization results.

[0072] In this embodiment, the batch normalization result approximately satisfies the normal distribution. If the batch normalization result approximately satisfies the standard normal distribution, the mean of the batch normalization result is =0, standard deviation of batch normalization results =1. The preset error range is .

[0073] When the feature value is less than the mean of the batch normalization result, the preset value can be , or it can be [ ]; when the feature value is greater than the mean of the batch normalization result, the default value is , or it can also be [0, ] other values ​​between ].

[0074] To avoid the propagation of erroneous values, the eigenvalues ​​may be directly set to the mean of the batch normalization results when the eigenvalues ​​are less than the mean of the batch normalization results or when the eigenvalues ​​are greater than the mean of the batch normalization results, such as directly setting the eigenvalues ​​to 0.

[0075] in, is a preset integer. After experiments, When the integer is less than 16, better results can be achieved.

[0076] The convolutional neural network fault-tolerance method provided by the present invention further improves the fault-tolerance performance by reasonably setting a preset value for truncating the eigenvalue when the eigenvalue is less than the mean of the batch normalization result and the eigenvalue is greater than the mean of the batch normalization result.

[0077] According to a convolutional neural network fault tolerance method provided by the present invention, in the step of checking and verifying the difference between the rows Determine the first error location Before, the method further includes: checking the row and comparing the difference In response to the existence of at least one element whose value is not in range, then perform the checksum based on the row and the difference Determine the first error location action; among them, , Respectively represent the pre-obtained row checksum and difference The mean and standard deviation of the elements in , Represents a preset integer; in the difference value according to the column checksum Determine the second error location Before, the method further includes: checking the column and comparing the difference In response to the existence of at least one element whose value is not in If the column checksum is within the range, the difference is checked. Determine the second error location action; among them, , Respectively represent the pre-obtained column checksum and difference The mean and standard deviation of the elements in , Represents a preset integer.

[0078] Checksum based on row Determine the first error location Before that, you need to determine the difference based on the row checksum The error position is detected. Check the difference between the rows and the The method to detect the error position is to check the line and the difference Traverse the elements in the table to determine the row checksum and the difference Are the elements in Within the range; if the row checksum is The elements in If the range is within the range, it means that the difference is checked according to the row No error position was detected; if the row checksum is There is at least one element whose value is not in If the range is within the range, it means that the difference is checked according to the row If the error position is detected, the row checksum is performed. Determine the first error location action. , Represents the pre-obtained row checksum and difference value respectively The mean and standard deviation of the elements in , Represents a preset integer.

[0079] Checksum based on column Determine the second error location Before, you need to determine the difference based on the column checksum The error position is detected. Check the difference between the columns and the checksum The method to detect the error position is: column check and difference Traverse the elements in the table to determine the column checksum and the difference Are the elements in Within the range; if the column checksum is the difference The elements in If the column is within the range, it means that the difference is checked based on the column No error position was detected; if the column checksum is different There is at least one element whose value is not in range, it means that the difference is checked based on the column If the error position is detected, the column checksum is performed. Determine the second error location Among them, , Represents the pre-obtained column checksum and difference value respectively The mean and standard deviation of the elements in , Represents a preset integer.

[0080] The convolutional neural network fault tolerance method provided by the present invention can verify the rows and the difference between them. Iterate over the elements in to determine whether they are within the preset range determined by the mean and variance, and check the columns and the difference. It traverses the elements in the table to determine whether they are within the preset range determined by the mean and variance, improving the row check and difference Determine the first error location And check the difference based on the column Determine the second error location reliability.

[0081] According to a convolutional neural network fault tolerance method provided by the present invention, in the checking of the rows and the difference Before traversing the elements in , the method further includes: replacing all convolutional layers of the convolutional neural network with the fault-tolerant layer, and setting a fault-tolerant control switch for each of the fault-tolerant layers to control whether to enable the fault-tolerant function; under error-free conditions, using the replaced convolutional neural network to perform reasoning on the test set data, and calculating the row checksum difference The mean of the elements in and standard deviation , and calculate the column checksum difference The mean of the elements in and standard deviation .

[0082] Row checksum difference The mean of the elements in and standard deviation , and calculate column checksums and differences The mean of the elements in and standard deviation , which can be determined through a test set before the model is applied.

[0083] Checksum and difference in rows Before traversing the elements in , all convolutional layers of the convolutional neural network are replaced with fault-tolerant layers, and a fault-tolerant control switch is set for each fault-tolerant layer to control whether the fault-tolerant function is enabled.

[0084] Under error-free conditions, use the replaced convolutional neural network to perform reasoning on the test set data and calculate the row checksum difference The mean of the elements in and standard deviation , and calculate column checksums and differences The mean of the elements in and standard deviation Since it is executed under error-free conditions, in order to reduce time overhead, the fault-tolerant control switches of each fault-tolerant layer can be turned off.

[0085] The convolutional neural network fault tolerance method provided by the present invention replaces all convolutional layers of the convolutional neural network with fault-tolerant layers, and sets a fault-tolerant control switch for each fault-tolerant layer to control whether to enable the fault-tolerant function. Under error-free conditions, the replaced convolutional neural network is used to perform reasoning on the test set data to calculate the row checksum and the difference. The mean of the elements in and standard deviation , and calculate column checksums and differences The mean of the elements in and standard deviation , providing a basis for determining the error location.

[0086] According to a convolutional neural network fault-tolerance method provided by the present invention, before the convolutional neural network receives input data for inference, the method also includes: controlling the turning on or off of the fault-tolerant control switches of each of the fault-tolerant layers according to the number of fault-tolerant layers with fault-tolerant functions turned on and the planning results of the distribution of the fault-tolerant layers with fault-tolerant functions turned on; wherein the planning results of the distribution of the fault-tolerant layers include giving priority to turning on the fault-tolerant control switches of the fault-tolerant layers with larger convolution kernel sizes; for fault-tolerant layers with the same priority, the fault-tolerant control switches are evenly turned on according to the network depth.

[0087] It is possible to pre-plan how many fault-tolerant layers to enable the fault-tolerant function. The number of fault-tolerant layers with the fault-tolerant function enabled is called the number of fault-tolerant layers. It is also possible to pre-plan which fault-tolerant layers to enable the fault-tolerant function. Therefore, before the convolutional neural network receives input data for reasoning, the fault-tolerant control switches of each fault-tolerant layer are controlled to be turned on or off according to the planning results of the number of fault-tolerant layers with the fault-tolerant function enabled and the distribution of the fault-tolerant layers with the fault-tolerant function enabled.

[0088] It is preferred to turn on the fault-tolerant switch of fault-tolerant layers with larger convolution kernel size, because these layers bear more computing tasks; for fault-tolerant layers of the same priority, they are evenly selected according to the network depth. Because the number of errors is exponentially related to the number of propagation layers, averaging the error propagation path length can minimize the number of errors.

[0089] The convolutional neural network fault-tolerance method provided by the present invention further improves the fault-tolerance performance by reasonably planning the number of fault-tolerance layers enabled according to the fault-tolerance function and the distribution of the fault-tolerance layers enabled according to the fault-tolerance function, and controlling the opening or closing of the fault-tolerance control switches of each fault-tolerance layer.

[0090] According to a convolutional neural network fault-tolerance method provided by the present invention, after obtaining the second convolution result after error correction, the method also includes: in response to reaching the cycle time for execution of the test phase, using the test set to perform inference and calculate the accuracy, and calculating the bit error rate based on the accuracy; in response to the bit error rate being greater than a first threshold, turning on the fault-tolerance control switches of more fault-tolerance layers; in response to the bit error rate being less than a second threshold, closing the fault-tolerance control switches of more fault-tolerance layers; wherein the second threshold is less than the first threshold; in response to the bit error rate being less than or equal to the first threshold and greater than or equal to the second threshold, keeping the setting of the fault-tolerance control switch of the fault-tolerance layer unchanged.

[0091] Most of the existing ABFT methods for neural network fault tolerance only propose a verification and construction scheme and design corresponding experiments for effectiveness analysis. In fact, various fault-tolerant methods need to be within a certain error rate range to be effective and feasible. In practical applications, the fault-tolerant effect will be affected by many factors, including the environment, hardware, and the fault-tolerant ability of the model itself. For example, wafer-level chips have much higher manufacturing defects and bit error rates than traditional chips due to their complex manufacturing process and large area; different hardware (such as FPGA, ASIC, etc.) also have different error rates; at the same time, different model structures and different layers of the model have different fault-tolerant capabilities. When the error rate is very high, most fault-tolerant methods will not be able to obtain correct results in time due to excessive overhead even if they can correct errors normally. At this time, it may be reasonable to sacrifice some error coverage to reduce overhead, but there is currently a lack of methods that can flexibly balance the two.

[0092] The present invention designs a workflow at the network level that can dynamically adjust the number of fault-tolerant layers according to the detection error rate. After the convolutional neural network completes one or more batch inference tasks, it determines whether to enter the test phase. If the set cycle time has been reached, it enters the test phase, otherwise it continues to execute the inference task. The test phase process includes: a) Perform reasoning on the test set and observe the accuracy. The correspondence between the accuracy and bit error rate of the test set has been obtained through fault injection experiments in an error-free environment. The bit error rate can be estimated based on the network output accuracy on the test set in an error-free environment (it may not be equal to the actual bit error rate, but the estimated bit error rate can be used to decide whether to strengthen the fault tolerance strength). In terms of fault injection tools, the MLIR fault injection tool can be used, or other fault injection tools based on pytorch, such as OPFI and pytorchFI, can also be used.

[0093] b) If the estimated bit error rate is too high, the fault tolerance control switches of more fault tolerance layers are set to true (i.e., turned on); if the estimated bit error rate is too low, the fault tolerance control switches of more fault tolerance layers are set to false (i.e., turned off). If the estimated bit error rate is within an acceptable range, the settings of the fault tolerance control switches of the fault tolerance layers remain unchanged.

[0094] c) If the error rate is extremely high, try truncating the convolution input value to ±3 and changing it to zero. This can further reduce the inference time in some high error rate cases.

[0095] The convolutional neural network fault-tolerance method provided by the present invention performs reasoning and calculates the accuracy using a test set in response to the cycle time of reaching the test phase execution, calculates the bit error rate according to the accuracy, and in response to the bit error rate being greater than a first threshold, turns on the fault-tolerance control switches of more fault-tolerance layers; in response to the bit error rate being less than a second threshold, turns off the fault-tolerance control switches of more fault-tolerance layers; in response to the bit error rate being less than or equal to the first threshold and greater than or equal to the second threshold, keeps the setting of the fault-tolerance control switch of the fault-tolerance layer unchanged, can adjust the setting of turning on or off the fault-tolerance function of the fault-tolerance layer according to the real-time fault-tolerance effect, introduces flexible fault-tolerance levels, restores the network output accuracy and reduces the overhead as much as possible in a high error rate environment, and realizes a dynamic trade-off between classification accuracy and overhead.

[0096] Figure 4 This is the second flow chart of the convolutional neural network fault tolerance method provided by the present invention. The present invention first performs initialization in an error-free environment, and then periodically runs the fault tolerance phase and the test phase in an error-prone environment. The overall flow chart is as follows: Figure 4 As shown, the method includes: 1. Initialization phase: The system accepts a trained convolutional neural network model and a partial test set of the dataset used for its training. It does not need to use the entire test set, only about 100 images are needed.

[0097] Step 1: The system replaces all convolutional layers in the network model with convolutional layers protected by the ABFT fault-tolerant method, which are called fault-tolerant layers. The fault-tolerant layer adds checksum calculation and checksum checking before and after calculating the convolution in the forward propagation. In addition, compared with the ordinary convolutional layer, it has four more attributes: the mean and standard deviation of the difference between the row checksum pairs and the column checksum pairs, and a switch controls whether to enable the fault-tolerant function. If not enabled, the convolution result is directly returned like the ordinary convolutional layer.

[0098] Step 2: Use the replaced convolutional neural network to perform inference on a small amount of test set data, with the goal of calculating the error-free checksum and difference The mean and standard deviation , used to set the error threshold , and calculate column checksums and differences The mean and standard deviation , used to set the error threshold The default parameter k=4, each convolutional layer has an independent , , , , and .

[0099] 2. Fault-tolerant stage: After initialization, the fault-tolerant enhanced convolutional neural network receives data for reasoning in a faulty environment (referring to the actual application environment of the model). During forward propagation, the following fault-tolerant process is executed once each time a fault-tolerant layer with fault-tolerant function enabled is passed: a) For the convolutional layer input after batch normalization (generally the default parameters are set to mean 0 and variance 1), all values ​​not in the range of [-3,3] are truncated to ±3; b) For the input feature map matrix ,calculate and connected to the end of the matrix to get , for the weight matrix calculate and connected to the end of the matrix to get ; c) Calculation and The convolution result is recorded as ; d) From Get row checksum , that is Calculate another row checksum , No. The elements are represented as , that is After removing the last column of sub-blocks and the last row of sub-blocks, sum the sub-blocks in rows; from Get the column checksum , that is Remove the last column of all sub-blocks from the last row of sub-blocks. Calculate another column checksum , No. The elements are represented as , that is After removing the last column of sub-blocks and the last row of sub-blocks, the sub-blocks are summed in columns; e) Calculate the difference between the row checksum pairs Difference from column checksum pair ; f) Yes Traverse the elements in the table to determine whether their values ​​are within the threshold If it is not within the range, it is considered an error and the error position is recorded. Traverse the elements in the table to determine whether their values ​​are within the threshold If it is not within the range, it is considered an error and the error position is recorded. The error location The meaning is (sub-block column index, row index within block, column index within block), The error location Meaning: (sub-block row index, row index within block, column index within block); g) According to and The error position is represented by the row index within the block and the column index within the block, that is, Create two dictionaries for the keys and , traverse to determine whether multiple errors are detected on the secondary block row / column at a certain intra-block index; h) Use double loop to key Traversal and ,for and :like or , indicating that the index in this block only detects one error on the row or column sub-block, and the difference can be directly added back for error correction; like and , indicating that the index in this block detected multiple errors on both the row and column sub-blocks. and , traverse all its combinations, for all possible error positions , , ,…… Set to zero to suppress error output; i) Return the convolution result after error correction, that is The remainder after removing the last row and column of secondary blocks.

[0100] 3. Testing phase After the convolutional neural network completes one or more batches of inference tasks, it determines whether to enter the test phase. If the set cycle time has been reached, it enters the test phase, otherwise it continues to execute the inference task. The test phase process is as follows: a) Perform inference on the test set and observe the accuracy. The correspondence between the accuracy and bit error rate of the test set has been obtained through fault injection experiments in an error-free environment. The bit error rate can be estimated based on the network output accuracy on the test set in an error-free environment (it may not be equal to the actual bit error rate, but the estimated bit error rate can be used to decide whether to strengthen the fault tolerance strength).

[0101] b) If the estimated error rate is too high, the fault-tolerance switches of more fault-tolerant layers are set to true; if the estimated error rate is too low, the fault-tolerance switches of more fault-tolerant layers are set to false. The selection of fault-tolerant layers has different priorities. It is preferred to turn on the fault-tolerance switches of fault-tolerant layers with larger kernel sizes because these layers bear more computing tasks; for fault-tolerant layers of the same priority, they are evenly selected according to the network depth. Because the number of errors is exponentially related to the number of propagation layers, averaging the error propagation path length can minimize the number of errors.

[0102] c) If the error rate is extremely high, try truncating the convolution input value to ±3 and changing it to zero. This can further reduce the inference time in some high error rate cases.

[0103] The following further illustrates the fault tolerance performance of the convolutional neural network fault tolerance method provided by the present invention through experimental data.

[0104] Table 1 is a summary table of experimental data. For the convenience of illustration, only some average value data are taken for different key points.

[0105] The experimental data is based on resnet50, the datasets are CIFAR10 and CIFAR100, and the test set contains all 10,000 images. The fault injection tool is MRFI, the fault injection type is activation, the fault injection selector is RandomPositionByRate, and the fault model is a FloatRandomBitFlip-float32 type value.

[0106] Table 1

[0107] Table 1 - Continuation

[0108] Among them, "3" means that if the eigenvalue is in the range of [0,3], it is set to 3; if the eigenvalue is in the range of [-3,0], it is set to -3.

[0109] Key point 1: The optimal number of fault-tolerant layers is different for different bit error rates.

[0110] The example table is shown in Table 2.

[0111] Table 2

[0112] In a low error rate environment (bit error rate = 1e-6), only 4 layers of fault tolerance are needed to restore the network classification accuracy to more than 99%. Using more fault tolerance layers will increase the inference time.

[0113] In a higher error rate environment, 4-layer fault tolerance has an absolute advantage in terms of overhead, and the inference time is almost the same as that of no fault tolerance in all three error rate levels, even when the bit error rate is 10 -4 It can also restore the network classification accuracy equivalent to random guessing to more than 70%. However, if a higher network classification accuracy is required, the number of fault-tolerant layers must be increased instead of considering using three 4-layer fault-tolerant networks for voting: Taking the 10-classification task CIFAR10 as an example, the network classification accuracy is 75% with 4-layer fault tolerance. Use three 4-layer fault-tolerant networks to vote, and choose the one with more votes as the final result. If the classification results of the three networks are different, one of them is randomly selected. After a simple calculation, the final classification accuracy is about 85.8%, which is still lower than 16-layer fault tolerance, and the gap will be even greater for the 100-classification task. Therefore, if a high accuracy is required in a high error rate environment, it is still necessary to choose to increase the number of fault-tolerant layers. Excessive overhead can be greatly reduced by truncating the error value of the convolutional layer input, see the next key point.

[0114] Beneficial technical effect: The appropriate number of fault-tolerant layers can be selected based on the estimated bit error rate, overhead, and classification accuracy requirements.

[0115] Key point 2: Error value truncation can reduce error correction time and improve accuracy.

[0116] The example table is shown in Table 3.

[0117] Table 3

[0118] It can be seen that the more fault-tolerant layers there are, the more the error value truncation of the convolution input can reduce the error correction time and improve the accuracy. The in-place error rate is as high as 10 -4 When , 53-layer fault tolerance + error value truncation can achieve the best recovery effect. Although the time cost is still more than 10 times that of 4-layer fault tolerance, it can achieve a higher classification accuracy than using 10 4-layer fault tolerance network voting (the latter has a final accuracy of about 92.8%).

[0119] Beneficial technical effect: Error value truncation can be used when the error rate is high, reducing time overhead and improving accuracy.

[0120] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the convolutional neural network fault tolerance method, the method comprising: according to the input first feature map matrix Calculate the first checksum , the first checksum Connected to the first feature map matrix At the end of the second feature map matrix ;in, , The fault-tolerant layer i feature map, represents the number of feature maps input to the fault-tolerant layer, , ; According to the first weight matrix Calculate the second checksum , the second checksum Connected to the first weight matrix The second weight matrix is ​​obtained at the end of ;in, ), The number of weight matrices representing the fault-tolerant layer; ; The fault-tolerant layer j A weight matrix, ; Calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result ;in, = , Represents the convolution operation; determines the first row checksum , the second line checksum , the first column checksum and the second column checksum ;in, ; Checksum on the second line The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of ; Checksum of the second column The The elements are represented as , Represents the first convolution result Bank of China No. , column number is Secondary blocks; calculate row checksum and difference Checksum and difference of columns ;in, , ; Checksum the difference according to the row Determine the first error location , based on the column checksum and the difference Determine the second error location ; Wherein, the first error position middle, Respectively represent the secondary block column index, the row index within the block, and the column index within the block; the second error position middle, Respectively represent the secondary block row index, the row index within the block, and the column index within the block; according to the first error position And the second error position Determine the row index and column index corresponding to the block In response to learning the row index and column index in the block according to the error detection situation If multiple errors are detected on both the secondary block row and the secondary block column, the position of the element in the block where the error may occur is obtained, and the error is corrected by setting the corresponding element in the block to 0; the second convolution result after error correction is obtained; wherein the second convolution result is the first convolution result after error correction Remove the remaining part of the last row of sub-blocks and the last column of sub-blocks.

[0121] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0122] On the other hand, the present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the convolutional neural network fault tolerance method provided by the above methods, the method comprising: according to the input first feature map matrix Calculate the first checksum , the first checksum Connected to the first feature map matrix At the end of the second feature map matrix ;in, , The fault-tolerant layer i feature map, represents the number of feature maps input to the fault-tolerant layer, , ; According to the first weight matrix Calculate the second checksum , the second checksum Connected to the first weight matrix The second weight matrix is ​​obtained at the end of ;in, ), The number of weight matrices representing the fault-tolerant layer; ; The fault-tolerant layer j A weight matrix, ; Calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result ;in, = , Represents the convolution operation; determines the first row checksum , the second line checksum , the first column checksum and the second column checksum ;in, ; Checksum of the second line The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of ; Checksum of the second column The The elements are represented as , Represents the first convolution result Bank of China No. , column number is Secondary blocks; calculate row checksum and difference Checksum and difference of columns ;in, , ; Checksum the difference according to the row Determine the first error location , based on the column checksum and the difference Determine the second error location ; Wherein, the first error position middle, Respectively represent the secondary block column index, the row index within the block, and the column index within the block; the second error position middle, Respectively represent the secondary block row index, the row index within the block, and the column index within the block; according to the first error position And the second error position Determine the row index and column index corresponding to the block In response to learning the row index and column index in the block according to the error detection situation If multiple errors are detected on both the secondary block row and the secondary block column, the position of the element in the block where the error may occur is obtained, and the error is corrected by setting the corresponding element in the block to 0; the second convolution result after error correction is obtained; wherein the second convolution result is the first convolution result after error correction Remove the remaining part of the last row of sub-blocks and the last column of sub-blocks.

[0123] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the convolutional neural network fault tolerance method provided by the above methods, the method comprising: according to the input first feature map matrix Calculate the first checksum , the first checksum Connected to the first feature map matrix At the end of the second feature map matrix ;in, , The fault-tolerant layer i feature map, represents the number of feature maps input to the fault-tolerant layer, , ; According to the first weight matrix Calculate the second checksum , the second checksum Connected to the first weight matrix The second weight matrix is ​​obtained at the end of ;in, ), The number of weight matrices representing the fault-tolerant layer; ; The fault-tolerant layer j A weight matrix, ; Calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result ;in, = , Represents the convolution operation; determines the first row checksum , the second line checksum , the first column checksum and the second column checksum ;in, ; Checksum on the second line The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of ; Checksum of the second column The The elements are represented as , Represents the first convolution result Bank of China No. , column number is Secondary blocks; calculate row checksum and difference Checksum and difference of columns ;in, , ; Checksum the difference according to the row Determine the first error location , based on the column checksum and the difference Determine the second error location ; Wherein, the first error position middle, Respectively represent the secondary block column index, the row index within the block, and the column index within the block; the second error position middle, Respectively represent the secondary block row index, the row index within the block, and the column index within the block; according to the first error position And the second error position Determine the row index and column index corresponding to the block In response to learning the row index in the block and the column index in the block according to the error detection situation If multiple errors are detected on both the secondary block row and the secondary block column, the position of the element in the block where the error may occur is obtained, and the error is corrected by setting the corresponding element in the block to 0; the second convolution result after error correction is obtained; wherein the second convolution result is the first convolution result after error correction Remove the remaining part of the last row of sub-blocks and the last column of sub-blocks.

[0124] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0125] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A convolutional neural network fault-tolerance method, wherein the convolutional neural network receives input data for inference, and executes the fault-tolerance method once each time a fault-tolerance layer with a fault-tolerance function enabled is passed through the convolutional layer protected by the fault-tolerance method during forward propagation, wherein: The fault-tolerant method comprises: According to the first feature map matrix of the input Calculate the first checksum , the first checksum Connected to the first feature map matrix At the end of the second feature map matrix ;in, , The fault-tolerant layer i feature map, represents the number of feature maps input to the fault-tolerant layer, , ; According to the first weight matrix Calculate the second checksum , the second checksum Connected to the first weight matrix The second weight matrix is ​​obtained at the end of ;in, ), Represents the number of weight matrices of the fault-tolerant layer; ; The fault-tolerant layer j A weight matrix, ; Calculate the second feature map matrix and the second weight matrix Convolution of , get the first convolution result ;in, = , Represents the convolution operation; Determine the first line checksum , the second line checksum , the first column checksum and the second column checksum ;in, ; Checksum on the second line The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of ; Checksum of the second column The The elements are represented as , Represents the first convolution result Bank of China No. , column number is The secondary block of Calculate row checksum difference Checksum and difference of columns ;in, , ; Checksum difference according to the row Determine the first error location , based on the column checksum and the difference Determine the second error location ; Wherein, the first error position middle, Respectively represent the secondary block column index, the row index within the block, and the column index within the block; the second error position middle, Respectively represent the secondary block row index, the row index within the block, and the column index within the block; According to the first error location And the second error position Determine the row index and column index corresponding to the block Error detection situations; In response to obtaining the row index and the column index in the block according to the error detection condition If multiple errors are detected on both the secondary block row and the secondary block column, the position of the element in the block where the error may occur is obtained, and the error is corrected by setting the corresponding element in the block to 0; Obtain a second convolution result after error correction; wherein the second convolution result is the first convolution result after error correction Remove the remaining part of the last row of sub-blocks and the last column of sub-blocks.

2. The convolutional neural network fault tolerance method according to claim 1, characterized in that: In the first error position And the second error position Determine the row index within the block and the column index within the block After the error detection situation, the method further comprises: In response to obtaining the row index and the column index in the block according to the error detection condition If only one error is detected on the secondary block row, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the column checksum to obtain the difference. Add together for error correction; In response to obtaining the row index and the column index in the block according to the error detection condition If only one error is detected on the secondary block column, the secondary block element value where the error is detected is obtained, and the secondary block element value is compared with the row checksum to obtain the difference. Add together for error correction.

3. The convolutional neural network fault tolerance method according to claim 2, characterized in that: The first error position And the second error position Determine the row index and column index corresponding to the block Error detection conditions include: According to the first error location Building the first dictionary ; wherein the first dictionary The key is the row index within the block and the column index within the block , the first dictionary The key value is the secondary block column index ; According to the second error position Create a second dictionary ; wherein the second dictionary The key is the row index within the block and the column index within the block , the second dictionary The key value is the secondary block row index ; In response to looping through the first dictionary and the second dictionary Obtain the row index and column index in the block The corresponding secondary block row index unique, then determine the row index and column index in the block Only one error was detected on the secondary block line; In response to looping through the first dictionary and the second dictionary Obtain the row index and column index in the block The corresponding secondary block column index unique, then determine the row index and column index in the block Only one error was detected on the secondary block column; In response to looping through the first dictionary and the second dictionary Obtain the row index and column index in the block The corresponding secondary block row index and secondary block column index If both are not unique, the row index and column index in the block are determined. Multiple errors are detected on both sub-block rows and sub-block columns.

4. The convolutional neural network fault tolerance method according to claim 1, characterized in that: In the first feature map matrix according to the input Calculate the first checksum Previously, the method also included: For the batch normalized fault-tolerant layer input, get the first feature map matrix The feature values ​​that are not within a preset error range of the mean of the batch normalization result are identified, and the feature values ​​are truncated to preset values ​​within the preset error range.

5. The convolutional neural network fault tolerance method according to claim 4, characterized in that: The batch normalization result approximately satisfies the normal distribution, and the preset error range is , , Respectively represent the mean and standard deviation of the batch normalization results, Represents a preset integer; When the feature value is less than the mean of the batch normalization result, the preset value is , when the feature value is greater than the mean of the batch normalization result, the preset value is ; Alternatively, the preset value is uniformly set to the mean of the batch normalization results.

6. The convolutional neural network fault tolerance method according to claim 1, characterized in that: In the difference value according to the row checksum Determine the first error location Previously, the method also included: Checksum the difference of the row In response to the existence of at least one element whose value is not in range, then perform the checksum based on the row and the difference Determine the first error location action; among them, , Respectively represent the pre-obtained row checksum and difference The mean and standard deviation of the elements in , Represents a preset integer; In the difference value according to the column checksum Determine the second error location Previously, the method also included: Checksum the difference of the column In response to the existence of at least one element whose value is not in If the column checksum is within the range, the difference is checked. Determine the second error location action; among them, , Represents the pre-obtained column checksum and difference value respectively The mean and standard deviation of the elements in , Represents a preset integer.

7. The convolutional neural network fault tolerance method according to claim 6, characterized in that: In the pair of row checksums, the difference Before traversing the elements in , the method further includes: Replacing all convolutional layers of the convolutional neural network with the fault-tolerant layers, and setting a fault-tolerant control switch for each of the fault-tolerant layers to control whether to enable the fault-tolerant function; Under error-free conditions, the replaced convolutional neural network is used to perform reasoning on the test set data to calculate the row checksum difference The mean of the elements in and standard deviation , and calculate the column checksum difference The mean of the elements in and standard deviation .

8. The convolutional neural network fault tolerance method according to claim 7, characterized in that: Before the convolutional neural network receives input data for inference, the method further includes: Controlling the turning on or off of the fault-tolerant control switches of each of the fault-tolerant layers according to the number of fault-tolerant layers with the fault-tolerant function turned on and the planning result of the distribution of the fault-tolerant layers with the fault-tolerant function turned on; Among them, the planning result of the fault-tolerant layer distribution includes preferentially turning on the fault-tolerant control switch of the fault-tolerant layer with a larger convolution kernel size; for fault-tolerant layers of the same priority, the fault-tolerant control switch is evenly turned on according to the network depth.

9. The convolutional neural network fault tolerance method according to claim 7, characterized in that: After obtaining the second convolution result after error correction, the method further includes: In response to reaching the cycle time of the test phase execution, using the test set to perform reasoning and calculate the correctness, and calculating the bit error rate according to the correctness; In response to the bit error rate being greater than a first threshold, turning on the fault tolerance control switches of more fault tolerance layers; In response to the bit error rate being less than a second threshold, closing the fault tolerance control switches of more fault tolerance layers; wherein the second threshold is less than the first threshold; In response to the bit error rate being less than or equal to the first threshold and greater than or equal to the second threshold, the setting of the fault-tolerant control switch of the fault-tolerant layer is kept unchanged.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the convolutional neural network fault tolerance method as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • FPGA-based neural network accelerator

    CN109948788A

  • Method and device for checking AI calculation

    CN115705487A

  • Fault detectable and tolerant neural network

    US20200074287A1

  • Error detection in convolutional operations

    US20240020419A1