Approximate fault detection and fault correction method and system

By optimizing the threshold for fault detection and positioning, combined with Bayesian algorithm to screen and process fault deviation values, the problems of high overhead and limited fault tolerance in large-scale computing systems are solved, and efficient fault detection and recovery are achieved.

CN120386319APending Publication Date: 2025-07-29HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510484334.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing fault detection and recovery technologies are highly complex in large-scale computing systems, have high resource consumption, and have limited fault tolerance, making it difficult to adapt to complex fault types.

Method used

By optimizing the threshold for fault detection and positioning, the Bayesian algorithm is used to select the target performance design requirements, filter out the fault deviation values higher than the optimization selection, perform row and column verification deviation vector processing, and correct the fault points if necessary.

Benefits of technology

On the premise of ensuring fault tolerance, the overhead of fault detection and recovery is reduced, and the computing efficiency and resource utilization are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386319A_ABST
    Figure CN120386319A_ABST
Patent Text Reader

Abstract

The invention discloses an approximate fault detection and fault correction method and system, and the method comprises the steps: setting a fault detection and fault positioning threshold value based on the target performance provided by a user, and maintaining a fault deviation value higher than the threshold value after optimization selection in the subsequent fault detection and fault positioning process, therefore, the expensive fault positioning overhead is reduced, and the unnecessary overhead is reduced as much as possible on the premise of ensuring the fault-tolerant capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computers, and particularly relates to a method and system for approximate fault detection and fault correction. Background Art

[0002] In existing fault detection and fault recovery technologies, it usually relies on exact comparison based on checksums. This method calculates the checksum of data and compares the calculated value with the expected value to determine whether the data has a fault. These methods have been proven to have a certain fault tolerance. In simple small-scale systems, they can effectively identify errors such as bit flips during data transmission, ensuring data integrity and the normal operation of the system.

[0003] However, this method is often accompanied by high computational complexity and resource consumption. In large-scale computing systems, the amount of data is extremely huge, and each detection requires a large amount of computational resources and time to calculate the checksum and make comparisons. With the continuous increase in the scale and complexity of computing systems, the overhead of traditional fault detection and fault recovery becomes increasingly unbearable. From the perspective of hardware resources, more processor capabilities, memory and other resources are required to support frequent checksum calculations. From the aspect of time overhead, this may lead to system response delays and affect the overall performance. Moreover, in the face of some complex and hidden fault types, this method also has limitations, and its fault tolerance faces more and more challenges. There is an urgent need for the emergence of new fault detection and recovery technologies to adapt to the increasingly complex computing environment.

[0004] The invention patent with the publication number CN117009125A discloses a fault detection method, device, electronic device and storage medium, which relates to the field of computer technology. The method includes: obtaining a corresponding connection relationship topology diagram according to the connection relationships of each electronic device, and converting the connection relationship topology diagram into a corresponding matrix; obtaining the alarm data of the target computer cluster, and obtaining the alarm topology diagram of each electronic device according to the association relationships between the alarm data, and converting the alarm topology diagram into a corresponding matrix; obtaining a device fault matrix according to the device adjacency matrix and the alarm topology diagram adjacency matrix, and inputting the device fault adjacency matrix into a fault location analysis model to output a fault detection result. By obtaining the combination of the connection relationships of each electronic device and the association relationships between the alarm data and inputting them into the fault location analysis model, it is judged whether the electronic device has a fault according to the output result. It can more accurately analyze the electronic device that actually has a fault.

[0005] The invention patent application with the publication number CN119669093A discloses a testing method, system, device, electronic device and medium for a fault-tolerant system, which relates to the field of computer technology, especially the fields of artificial intelligence, software testing and model training technology. The specific implementation solution is as follows: receiving a test case, where the test case includes a target fault type and expected behavior data. Then, based on the preset correspondence between various fault types and preset faults, determining the target fault corresponding to the target fault type. Next, injecting the target fault into the training device of the model to be tested. And obtaining the actual behavior data from the fault-tolerant system, where the actual behavior data is the data generated by the coping behavior executed by the fault-tolerant system to cope with the target fault. After that, comparing the actual behavior data and the expected behavior data to obtain the test result of the fault-tolerant system.

[0006] The above two patent applications have a certain improvement in fault tolerance, but the overhead is still relatively large. Therefore, there is an urgent need to propose an approximate fault detection and fault recovery system and method, aiming to reduce the fault tolerance overhead on the premise of ensuring the fault tolerance ability. Summary of the Invention

[0007] The present invention provides a method for approximate fault detection and fault correction, which relaxes the indicators of fault detection and fault recovery and reduces the fault tolerance overhead on the premise of ensuring the fault tolerance ability.

[0008] The method for approximate fault detection and fault correction provided by the specific embodiments of the present invention includes:

[0009] Optimally selecting the thresholds for fault detection and fault location of the target model according to the target performance design requirements provided by the user;

[0010] Calculating the output matrix fault deviation value for the operations of the target model, and screening out the operations corresponding to the fault deviation values higher than the fault detection threshold obtained by the optimal selection;

[0011] Calculating the row and column check deviation vectors for the screened operations, and deleting the fault deviation values smaller than the fault location threshold obtained by the optimal selection in the row and column check deviation vectors to obtain the screened row and column fault deviation value vectors;

[0012] When there are single row and column fault deviation values in the screened row and column fault deviation value vectors respectively, the position of the element corresponding to the intersection of the single row and column fault deviation values is the screened fault point, and adding the single row and column fault deviation values to the fault value corresponding to the screened fault point to obtain the corrected fault-free value, thus completing the fault correction.

[0013] Preferably, when there are multiple fault deviation values in the rows and columns of the filtered row and column fault deviation value vectors respectively, the positions of the elements corresponding to the intersections of the multiple row and column fault deviation values are the selected fault points, and the selected fault points are set to zero.

[0014] Preferably, the Bayesian algorithm is used to optimize the selection of the thresholds for fault detection and fault location of the target model, so that the overhead of fault detection and fault correction meets the set requirements.

[0015] Preferably, using the Bayesian algorithm to optimize the selection of the thresholds for fault detection and fault location of the target model includes:

[0016] First, create a Bayesian optimizer and initialize the sample set, and generate a set of given observations in the search space;

[0017] Then, use Bayesian to obtain the posterior distribution, use the improved acquisition function to guide the search direction, and under the guidance of the posterior distribution and the acquisition function, achieve the directional exploration of the parameter space in t rounds of iteration;

[0018] Finally, return the searched threshold configuration for fault detection and fault location.

[0019] Preferably, the operations of the target model include convolution operations or GEMM operations.

[0020] Preferably, calculating the output matrix fault deviation value for the operations of the target model includes:

[0021] Assume that matrix multiplication is C = A·B, where the matrix dimension is assumed to be N, and α is an N-dimensional all-1 vector [1, 1,..., 1];

[0022] First, calculate the column sum vector A of matrix A checksum = αA and the row sum vector B of matrix B checksum = Bα T , and generate the predicted matrix checksum C through dot product operation L = A checksum ·

[0023] B checksum , and at the same time, accumulate all elements of the output matrix to obtain the actual matrix checksum

[0024]

[0025] Then calculate the deviation MSD = |C L - C R |.

[0026] Preferably, calculating the fault deviation values of the rows and columns of the output matrix of the screened operations to obtain the row and column check deviation vectors includes:

[0027] Summing each row of the output matrix of the screened operations to obtain the row sum vector, and summing each column of the output matrix to obtain the column sum vector;

[0028] First, calculate the checksum product of A L_Row and matrix B to generate the predicted output matrix row check vector C checksum = A checksum · B. At the same time, calculate the product of matrix A and B L_Col to obtain the predicted output matrix column check vector C checksum = A · B

[0029] Then, sum the output matrix C column by column and row by row to obtain the actual row check vector C R_Row = αC and the column check vector C R_Col = Cα T ;

[0030] By calculating the deviation between the predicted vector and the actual vector, the row and column check deviation vector R / CSD = [|C L_Row - C R_Row |, |C L_Col - C R_Col |] is obtained.

[0031] The present invention also provides an approximate fault detection and fault correction system, including a multi-threshold collaborative optimization module, a fault detection threshold setting module, a fault detection module, a fault location threshold setting module, a fault location module, and a fault correction module;

[0032] The multi-threshold collaborative optimization module is used to optimize and select the thresholds for fault detection and fault location of the target model according to the target performance design requirements provided by the user;

[0033] The fault detection threshold setting module and the fault location threshold setting module are used to respectively send the optimized and selected fault detection threshold and fault location threshold to the fault detection module and the fault location module;

[0034] The fault detection module is used to calculate the fault deviation value of the operations of the target model, and screen out the operations corresponding to the fault deviation values higher than the optimized and selected fault detection threshold;

[0035] The fault location module is used to calculate the fault deviation values of the rows and columns of the output matrix of the screened operations to obtain the row and column check deviation vector, and delete the fault deviation values smaller than the optimized and selected fault location threshold in the row and column check deviation vector to obtain the screened row and column fault deviation value vector;

[0036] The fault correction module is used to, when there are single row and column fault deviation values in the fault deviation value vectors of the filtered rows and columns respectively, take the position of the element corresponding to the intersection of the single row and column fault deviation values as the filtered fault point, and add the fault deviation values of the single row and column to the fault value corresponding to the filtered fault point to obtain the corrected fault-free value, thereby completing the fault correction.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] The present invention sets the thresholds for fault detection and fault location based on the target performance provided by the user. During the subsequent fault detection and fault location processes, the fault deviation values higher than the optimized selected thresholds are retained, thereby reducing the expensive fault location overhead, so as to minimize unnecessary overhead on the premise of ensuring the fault tolerance ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flowchart of a method for approximate fault detection and fault correction provided by a specific embodiment of the present invention;

[0040] Figure 2 It is a block diagram of an approximate fault detection and fault correction system provided by a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further details an approximate fault detection and fault recovery system and method provided in the embodiments of the present invention with reference to the accompanying drawings.

[0042] A specific embodiment of the present invention provides a method for approximate fault detection and fault correction, Figure 1 as shown, including:

[0043] S1. Optimally select the thresholds for fault detection and fault location of the target model according to the target performance design requirements provided by the user: Since additional overhead will be generated during the process of fault detection and fault correction of the target model, in order to achieve the target accuracy provided by the user and minimize the fault tolerance overhead as much as possible, a specific embodiment of the present invention first optimally selects the thresholds for fault detection and fault location of the target model. The fault detection threshold and fault location threshold after the optimal selection can locate the fault position, so as to meet the target accuracy provided by the user and lower fault tolerance overhead.

[0044] In a specific embodiment, the present embodiment uses the Bayesian algorithm to optimally select the thresholds for fault detection and fault location of the target model, so that the overhead of fault detection and fault correction meets the set requirements.

[0045] Specifically, the method for optimizing and selecting the thresholds for fault detection and fault location of the target model using the Bayesian algorithm provided in this embodiment includes:

[0046] First, create a Bayesian optimizer and initialize the sample set to generate a set of given observations in the search space;

[0047] Then, use Bayesian to obtain the posterior distribution, and use the improved acquisition function to guide the search direction. Under the guidance of the posterior distribution and the acquisition function, perform directional exploration of the parameter space in t rounds of iteration;

[0048] Finally, return the threshold configuration for fault detection and fault location found by the search.

[0049] S2. The fault detection stage provided in this embodiment includes: calculating the output matrix fault deviation value for the operations of the target model, and screening out the operations corresponding to the fault deviation values higher than the optimized and selected fault detection threshold.

[0050] Specifically, this embodiment detects the convolution or GEMM operations of the target model to obtain the fault deviation values of the convolution or GEMM operations, compares the fault deviation values with the optimized and selected fault detection threshold obtained in step S1, and calls the subsequent convolution or GEMM operations higher than the fault detection threshold, and ignores the convolution or GEMM operations smaller than the fault detection threshold to reduce the fault tolerance overhead.

[0051] In a specific embodiment, calculating the output matrix fault deviation value for the operations of the target model in this embodiment includes:

[0052] Assume that matrix multiplication is C = A·B, where the matrix dimension is assumed to be N, and α is an all-1 vector [1, 1,..., 1] of dimension N;

[0053] First, calculate the column sum vector A of matrix A checksum = αA and the row sum vector B of matrix B checksum = Bα T , generate the predicted matrix checksum C through dot product operation L = A checksum ·

[0054] B checksum , and at the same time, accumulate all elements of the output matrix to obtain the actual matrix checksum

[0055]

[0056] Then calculate the deviation MSD = |C L - C R |.

[0057] S3. The fault location stage provided in this embodiment includes: calculating the fault deviation values of the rows and columns of the output matrix of the selected operations to obtain the row and column check deviation vectors, converting multiple non-correctable fault situations in the rows and columns into correctable situations of single faults in the rows and columns, and then deleting the fault deviation values smaller than the optimized selected fault location threshold in the row and column check deviation vectors, reducing the fault tolerance overhead while ensuring the target accuracy, and obtaining the filtered fault deviation value vectors of the rows and columns.

[0058] In a specific embodiment, calculating the fault deviation values of the rows and columns of the output matrix of the selected operations to obtain the row and column check deviation vectors provided in this embodiment includes:

[0059] Summing each row of the output matrix of the selected operations to obtain the row sum vector, and summing each specific column to obtain the column sum vector;

[0060] First, calculate the checksum product of A L_Row and matrix B to generate the predicted output matrix row check vector C checksum = A checksum · B. At the same time, calculate the product of matrix A and B L_Col to obtain the predicted output matrix column check vector C checksum = A · B

[0061] Then sum each column and each row of the output matrix C to obtain the actual row check vector C R_Row = αC and the column check vector C R_Col = Cα T ;

[0062] By calculating the deviation between the predicted vector and the actual vector, the row and column check deviation vector R / CSD = [|C L_Row - C R_Row |, |C L_Col - C R_Col |] is obtained.

[0063] S4. The fault correction stage provided in this embodiment includes: when there are single row and column fault deviation values in the filtered fault deviation value vectors of the rows and columns respectively, the position of the element corresponding to the cross of the single row and column fault deviation values is the selected fault point, and adding the single row and column fault deviation values to the fault value corresponding to the selected fault point to obtain the corrected fault-free value, directly correcting the fault situation and efficiently completing the task of fault correction.

[0064] In a specific embodiment, the fault correction stage provided in this embodiment further includes a multi-fault situation. When there are multiple row and column fault deviation values respectively in the filtered row and column fault deviation value vectors, the positions of the elements corresponding to the intersection of the multiple row and column fault deviation values are the filtered fault points. The filtered fault points are set to zero to avoid the influence of the fault values on subsequent calculations.

[0065] The embodiment of the present invention also provides an approximate fault detection and fault correction system, as Figure 2 shown, including a multi-threshold collaborative optimization module 11, a fault detection threshold setting module 12, a fault detection module 13, a fault location threshold setting module 14, a fault location module 15, and a fault correction module 17.

[0066] The multi-threshold collaborative optimization module 11 provided in this embodiment is used to optimize and select the thresholds for fault detection and fault location of the target model according to the target performance design requirements provided by the user;

[0067] The fault detection threshold setting module 12 and the fault location threshold setting module 14 provided in this embodiment are used to send the optimized and selected fault detection threshold and fault location threshold to the fault detection module 13 and the fault location module 15 respectively;

[0068] The fault detection module 13 provided in this embodiment is used to calculate the fault deviation value of the operation of the target model and filter out the operations corresponding to the fault deviation values higher than the optimized and selected fault detection threshold;

[0069] The fault location module 15 provided in this embodiment is used to calculate the row and column fault deviation values of the output matrix of the filtered operations to obtain the row and column check deviation vectors, and delete the fault deviation values smaller than the optimized and selected fault location threshold in the row and column check deviation vectors to obtain the filtered row and column fault deviation value vectors;

[0070] The fault correction module 17 provided in this embodiment is used to, when there are single row and column fault deviation values respectively in the filtered row and column fault deviation value vectors, the positions of the elements corresponding to the intersection of the single row and column fault deviation values are the filtered fault points, and add the single row and column fault deviation values to the fault values corresponding to the filtered fault points to obtain the corrected fault-free values, thereby completing the fault correction.

[0071] The approximate fault detection and fault correction system provided in this embodiment further includes a storage module 10 for storing input, output, and intermediate data, including but not limited to the training and forward data required for the fault detection, fault location, and fault correction processes.

[0072] Although the present invention has been described by way of preferred embodiments, the present invention is not limited to the embodiments described herein, and various changes and modifications made without departing from the scope of the present invention are also included.

Claims

1. A method for approximate fault detection and fault correction, characterized in that, Including: Optimally select the thresholds for fault detection and fault location of the target model according to the target performance design requirements provided by the user; Calculate the output matrix fault deviation value for the operations of the target model, and filter out the operations corresponding to when the fault deviation value is higher than the optimally selected fault detection threshold; Calculate the row and column check deviation vectors for the filtered operations, and delete the fault deviation values smaller than the optimally selected fault location threshold in the row and column check deviation vectors to obtain the filtered row and column fault deviation value vectors; When there are single row and column fault deviation values in the filtered row and column fault deviation value vectors respectively, the position of the element corresponding to the intersection of the single row and column fault deviation values is the filtered fault point, and add the single row and column fault deviation values to the fault value corresponding to the filtered fault point to obtain the corrected fault-free value, thus completing fault correction.

2. The method for approximate fault detection and fault correction according to claim 1, wherein When there are multiple row and column fault deviation values in the filtered row and column fault deviation value vectors respectively, the position of the element corresponding to the intersection of the multiple row and column fault deviation values is the filtered fault point, and perform a zeroing operation on the filtered fault point.

3. The method for approximate fault detection and fault correction according to claim 1, characterized in that Use the Bayesian algorithm to optimally select the thresholds for fault detection and fault location of the target model, so that the overhead of fault detection and fault correction meets the set requirements.

4. The method for approximate fault detection and fault correction according to claim 3, characterized in that, Using the Bayesian algorithm to optimally select the thresholds for fault detection and fault location of the target model, including: First, create a Bayesian optimizer and initialize the sample set, and generate a set of given observation results in the search space; Then use Bayesian to obtain the posterior distribution, use the improved acquisition function to guide the search direction, and under the guidance of the posterior distribution and the acquisition function, achieve the directional exploration of the parameter space in t rounds of iteration; Finally, return the searched threshold configuration for fault detection and fault location.

5. The method for approximate fault detection and fault correction according to claim 1, characterized in that The operations of the target model include convolution operations or GEMM operations.

6. The method for approximate fault detection and fault correction according to claim 1, characterized in that Calculating the output matrix fault deviation value for the operations of the target model, including: Assume matrix multiplication is C = A·B, where the assumed matrix dimension is N, and α is an N-dimensional all-1 vector [1, 1,..., 1]; First, calculate the column sum vector A of matrix A checksum = αA and the row sum vector B of matrix B checksum = Bα T , generate the prediction matrix checksum C through dot product operation L = A checksum ·B checksum , at the same time, accumulate all elements of the output matrix to obtain the actual matrix checksum Then calculate the deviation between the predicted value and the actual value MSD = |C L - C R |.

7. The method for approximate fault detection and fault correction according to claim 6, wherein Calculating the row and column fault deviation values of the output matrix of the filtered operations to obtain the row and column check deviation vectors, including: Sum each row of the output matrix of the filtered operations to obtain the row sum vector, and sum each column of the output matrix to obtain the column sum vector; First, calculate A checksum The product with matrix B to generate the row check vector C of the predicted output matrix L_Row = A checksum · B. At the same time, calculate the product of matrix A and B checksum to obtain the column check vector C of the predicted output matrix L_Col = A · B checksum ; Then, perform column-wise and row-wise summations on the output matrix C to obtain the actual row check vector C R_Row = αC and the column check vector C R_Col = Cα T ; The row and column check deviation vector R / CSD = [|C L_Row - C R_Row |, |C L_Col - C R_Col |] is obtained by calculating the deviation between the predicted vector and the actual vector.

8. An approximate fault detection and fault correction system, characterized in that, Including a multi-threshold collaborative optimization module, a fault detection threshold setting module, a fault detection module, a fault location threshold setting module, a fault location module, and a fault correction module; The multi-threshold collaborative optimization module is used to optimally select the thresholds for fault detection and fault location of the target model according to the target performance design requirements provided by the user; The fault detection threshold setting module and the fault location threshold setting module are used to send the optimally selected fault detection threshold and fault location threshold to the fault detection module and the fault location module respectively; The fault detection module is used to calculate the fault deviation value for the operations of the target model, and filter out the operations corresponding to when the fault deviation value is higher than the optimally selected fault detection threshold; The fault location module is used to calculate the fault deviation values of the rows and columns of the output matrix of the selected operations to obtain the row and column check deviation vectors, and delete the fault deviation values smaller than the fault location threshold obtained by optimal selection in the row and column check deviation vectors to obtain the fault deviation value vectors of the filtered rows and columns; The fault correction module is used to, when there are single row and column fault deviation values in the fault deviation value vectors of the filtered rows and columns respectively, the position of the element corresponding to the intersection of the single row and column fault deviation values is the selected fault point, and add the single row and column fault deviation values to the fault value corresponding to the selected fault point to obtain the corrected fault-free value, thereby completing the fault correction.

Citation Information

Patent Citations

  • Fault detection method and device, electronic equipment and storage medium

    CN117009125A

  • Testing method, system and device for fault-tolerant system, electronic equipment and medium

    CN119669093A