Completeness check
By using an approximation circuit to compare approximate and exact results with a deviation threshold, the method efficiently detects errors in processing circuits, addressing inefficiencies in existing detection methods.
Patent Information
- Application Number
- JP2024574541
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-28
- Filing Date
- 2023-06-01
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods for detecting intermittent failures in processing circuits, such as wayward cores, are either costly in terms of execution overhead or require expensive replicated calculation logic, making them inefficient and power-intensive.
Incorporating an approximation circuit within the processing circuit to calculate an approximate result of a computation, which is then compared with the exact result to detect errors using a deviation threshold, thereby reducing the need for redundant execution or replicated cores.
This approach provides a quick and efficient method to detect errors in processing circuits by comparing approximate and exact results, reducing power and area requirements while maintaining performance.
Smart Images

Figure 2025520568000001_ABST
Abstract
Description
Technical Field
[0001] This technique relates to the field of data processing. More specifically, this technology relates to integrity checking.
Background Art
[0002] In a processor, a fault that affects the operation of the processor may occur. Such a fault may be caused by one or more of several possible factors, such as manufacturing problems, system aging, or bugs in hardware or software design. If not addressed, these faults can lead to errors in the calculations performed by the processor. In some cases, such faults may occur only in a non-deterministic manner even in the same calculation executed on the same processor or core, resulting in an incorrect result in one instance but not in another. Some faults in the processing circuit lead to easily identifiable errors, but some faults result in valid but incorrect results generated from the calculation, making the detection of these faults more difficult. A core that exhibits such behavior may be referred to as a "wayward core".
[0003] “Cores that not count”, Peter H. Hochschild et al., 2021, Proc. 18th Workshop on Hot Topics in Operating Systems (HotOS 2021) describes some known cases of such wayward cores that lead to intermittent faults.
Summary of the Invention
[0004] In one exemplary configuration, an apparatus is provided, the apparatus comprising a processing circuit that executes instructions, the processing circuit including a computing circuit that responds to one or more instructions that require a computation to be performed to calculate a result of the computation, an approximation circuit that, in response to the one or more instructions and independently of the computing circuit, calculates an approximate result of the computation, and a completeness check circuit that compares the result of the computation performed by the computing circuit with the approximate result of the computation performed by the approximation circuit and, in response to determining that a difference between the result of the computation and the approximate result of the computation is greater than a deviation threshold, detects an error in the processing circuit, thereby performing a completeness check.
[0005] In another exemplary configuration, a method is provided that includes executing instructions by a processing circuit, calculating a result of a computation in response to one or more instructions that require the computation to be performed, calculating an approximate result of the computation independently of the computation in response to the one or more instructions, comparing the result of the computation with the approximate result of the computation, and detecting an error in the processing circuit in response to determining that a difference between the result of the computation and the approximate result of the computation is greater than a deviation threshold, thereby performing a completeness check.
[0006] In a further exemplary configuration, a computer program is provided for controlling a host data processing device to provide an instruction execution environment, the computer program including processing program logic for executing instructions, the processing program logic including calculation program logic that responds to one or more instructions that require a calculation to be performed to calculate a result of the calculation, approximation program logic that calculates an approximate result of the calculation, independently of the calculation program logic, in response to the one or more instructions, and integrity check program logic that performs an integrity check by comparing a result of a calculation performed by the calculation program logic with an approximate result of the calculation performed by the approximation program logic, and detecting an error in the processing program logic in response to determining that a difference between the result of the calculation and the approximate result of the calculation is greater than a deviation threshold value.
Brief Description of the Drawings
[0007] Further aspects, features, and advantages of the present technique will become apparent upon reading the following description of examples in conjunction with the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0008] Before considering the examples with reference to the accompanying drawings, the following description of the examples is provided.
[0009] The intermittent failures of the whimsical core as described above can lead to errors in the calculations performed by the processing circuit in several possible ways. For example, the intermittent failures can cause the following. ● Data encrypted on one whimsical core can only be decrypted by that core; ● Data is corrupted during copying; ● Violation of lock semantics; and ● Calculation errors, for example, int(1.1^53) returns 0, while int(1.1^52) correctly returns 142.
[0010] The present technology provides an apparatus for detecting intermittent failures, particularly an apparatus for detecting intermittent failures that lead to calculation errors. The detection of such failures may also be referred to as a integrity check.
[0011] One possible approach for detecting intermittent failures in a processing circuit involves redundant execution of an application, i.e., executing the code more than once and comparing the resulting outputs. However, such an approach involves a large overhead for execution and is particularly costly since each calculation has to be repeated.
[0012] Another possible approach for detecting failures is to provide replicated calculation logic, for example, by providing two cores (referred to as dual-core lockstep) that operate in parallel and execute the same workload using the outputs of the cores that are compared in each cycle. However, this approach is expensive in terms of the required area and power since it requires provision of an additional core along with the circuitry necessary to compare the results of the two cores.
[0013] The inventors have recognized that in many cases, an exact value as a comparison target for the result obtained by the calculation circuit that executed the calculation is not required to check the completeness of a specific calculation. Rather, it may be sufficient to calculate an approximate result of the calculation and compare the approximate result with the result obtained by the calculation circuit. If a fault has caused an error in a calculation that has resulted in a sufficient deviation in the result, the approximate result will be different enough from the obtained result to be able to determine that an error has occurred or is likely to have occurred. On the other hand, even if the approximate result is different from the result obtained by the calculation circuit, if the difference is a relatively small amount, this may indicate that the calculation is likely to have produced the correct result.
[0014] Therefore, there is no need to provide replicated calculation logic such as separate cores. Instead, the device has an approximation circuit that can obtain an approximate result using less power and / or area than is required to independently calculate another exact result. The approximate result may also be calculated in parallel with the result obtained by the calculation circuit, thereby providing a more efficient process for completeness checking than can be achieved when the same circuit is used to repeat the calculation one or more times.
[0015] According to the technology described in this specification, an apparatus is provided that includes a processing circuit for executing instructions. The processing circuit may include a central processing unit (CPU) or a graphics processing unit (GPU), or a part thereof. In particular, the processing circuit may include an arithmetic logic unit (ALU) or a floating point unit (FPU) such as those found in a CPU or GPU. The processing circuit has a calculation circuit that responds to one or more instructions that require calculations to be performed. For example, the calculation circuit may be a logic unit of an ALU or FPU configured to perform a specific type of calculation. When the processing circuit executes an instruction that requires that type of calculation to be performed, the calculation circuit performs the calculation and calculates a calculation result.
[0016] However, a fault in the calculation circuit can lead to the generation of incorrect calculation results. In some cases, the fault may not be obvious to the extent that it can be said that the calculation circuit has completely failed to produce a result. Alternatively, the fault may instead lead to the calculation circuit providing an inaccurate but otherwise valid calculation result. This fault can be intermittent in the sense that it causes an error in the result of a calculation in some situations or at some times, but not in other situations or at other times, even when the conditions under which the calculation is performed are the same.
[0017] Therefore, an approximation circuit and a integrity check circuit are provided to check the integrity of the calculation circuit itself and to determine whether the calculation result is likely to be accurate. The approximation circuit is configured to independently calculate an approximate result of a calculation against which the result obtained by the calculation circuit can be compared. This is performed simultaneously with the calculation by the calculation circuit, which can reduce the impact on performance of performing this check.
[0018] When the calculation result from the calculation circuit and the approximation result from the approximation circuit become available, the integrity check circuit performs an integrity check by comparing the calculation result and the approximation result. Based on the difference between the approximation result and the calculation result, the integrity check circuit can determine whether there is an error in the processing circuit. This determination regarding the existence of an error utilizes a deviation threshold that indicates the expected or acceptable difference between the approximation result and the result from the calculation circuit. Since the approximation result is not calculated precisely, some deviation is expected between the approximation result and the result from the calculation circuit even when the processing circuit is functioning correctly. Thus, the deviation threshold provides a way to determine whether this deviation is sufficient to determine that an error has been detected. Therefore, a small deviation between the calculation result from the calculation circuit and the approximation result may be due to the approximation result not being precise, but if the calculation circuit produces a result that is significantly different from the expected result determined by the approximation, this may indicate a malfunction in the calculation circuit (or actually the approximation circuit).
[0019] The integrity check circuit can be configured to take several possible actions in response to the detection of an error. For example, the integrity check circuit can indicate that an error has been detected by raising an exception or saving a specific value to a register. This can be used to indicate to the software that an error has been detected and thus that the calculation result may be incorrect, or, for example, to help detect erratic cores that are prone to intermittent malfunctions and enable these cores to be replaced. In addition or alternatively, the integrity check circuit may repeat the calculation.
[0020] In some examples, the integrity check can be initiated based on an integrity check start instruction, which enables the integrity check to be controlled using software. That is, a programmer or compiler may include an integrity check start instruction in the program code such that execution of the integrity check start instruction by the processing circuit causes the integrity check to be performed.
[0021] This can be implemented by using the integrity check start instruction in combination with an integrity check end instruction. If one or more instructions to cause the calculation circuit to perform a calculation follow the integrity check start instruction, execution of the integrity check start instruction causes the approximation circuit to calculate an approximation result of those calculations. Next, in response to execution of the integrity check end instruction, the processing circuit signals the end of a sequence of instructions for which an integrity check should be performed such that the integrity check circuit is called to perform the integrity check and compare the results obtained by the calculation circuit and the approximation circuit.
[0022] In the case of the calculation circuit, the location of the operands on which the calculation is based can be specified by the instruction being executed. For example, a particular form of instruction may indicate the register in which an operand is stored, or a register that contains the address in memory from which an operand is retrieved. It will be understood that the operands may be indicated in other ways, such as using an offset, or encoded in the instruction itself.
[0023] The approximate circuit can determine the location of the operand on which it is based when it operates from the same instructions as described above. However, in some examples, the integrity check start instruction specifies one or more registers that include the operand on which the calculation represented by the subsequent instructions is based when it operates. In this case, the approximate circuit is configured to obtain one or more operands from the one or more specified registers. This approach enables the operand to be placed in any of several registers that may be provided within the device, thereby providing flexibility regarding where the operand can be stored before invoking the integrity check mechanism.
[0024] However, in some examples, to avoid the need to specify registers in the integrity check start instruction, one or more operands on which the approximate circuit is based when it operates can be determined in advance, for example, by being defined as architectural behavior. Thus, in response to the integrity check start instruction, the approximate circuit can obtain the operand from a predetermined register.
[0025] In some cases, the approximate circuit may only be suitable for approximating the results of a particular type of calculation supported by the calculation circuit. For other types of calculations supported by the calculation circuit, due to the nature of the calculation and / or the configuration of the approximate circuit, the approximate circuit may not be able to determine the approximate result, or may not be able to determine the approximate result with sufficient accuracy to enable the integrity check to be performed using the approximate circuit.
[0026] In this case, the approximate circuit may further include an approximation suitability check circuit for determining whether a particular calculation is suitable for approximation by the approximate circuit. Based on the determination of whether a particular calculation is suitable for approximation, the approximation suitability check circuit is configured to permit or prevent, respectively, the execution of the integrity check for that particular calculation.
[0027] The approximate fitness check circuit may be configured to determine whether a particular calculation is suitable for approximation based on several possible factors. In some examples, the approximate fitness check circuit utilizes two or more different mechanisms to determine whether a calculation is suitable for approximation.
[0028] One such approach involves determining that a particular calculation is suitable for approximation based on the determination that the calculation involves evaluating a function that is mathematically smooth at the point where the function is to be evaluated. The fact that the function is mathematically smooth at that point means that the function can be differentiated and the derivative can be evaluated at that point. For example, if the calculation involves raising one operand to the power of another operand, the function being evaluated is smooth and the function can be differentiated.
[0029] In determining whether a particular calculation is suitable for approximation, the approximate fitness check circuit may also take into account the behavior of the derivative of the function being evaluated, in which case it is found that only calculations involving evaluating a function with a derivative that does not change too rapidly are suitable for approximation.
[0030] In some cases, the function being evaluated will directly correspond to the instruction being executed (e.g., if the function involves addition, an addition instruction may be given). However, in some cases, the function will correspond to several instructions (e.g., if the instruction set architecture does not provide a multiplication instruction, the function may correspond to one or more addition instructions and branch instructions).
[0031] This condition regarding fitness for approximation may be used, for example, when an approximation circuit depends on the derivative of a function to determine an approximate result.
[0032] In some examples, the approximate fitness check circuit is configured to use at least one of a neural network, a random forest, and a decision tree to determine whether a particular calculation is suitable for approximation by the approximate circuit. These structures may be trained to recognize functions for which the approximate circuit can generate reliable approximation results. These structures may be trained based on the observed results of the computing circuit and / or the approximate circuit for various functions evaluated using the device, or these structures may be pre-trained before being deployed to the device, in which case a pre-trained model used by the neural network, random forest, or decision tree is provided to the device.
[0033] Another technique for identifying whether a calculation is suitable for approximation by the approximate circuit involves identifying the calculation as suitable for approximation based on determining that the calculation corresponds to an operation from a predetermined list of operations. For example, the approximate fitness check circuit may be provided with an indication of a particular operation (which may correspond to a particular instruction or sequence of instructions) known to be suitable for approximation using the approximate circuit. If the calculation corresponds to such an operation, thereby, the approximate fitness check circuit may determine that the calculation is suitable for approximation.
[0034] The approximate circuit itself can utilize various techniques for calculating the approximate result. For example, a neural network can be used to determine the approximate result. However, in some examples, the approximate circuit maintains and utilizes calculation result history information based on previous calculations performed by the computing circuit. In this way, the approximate circuit can utilize the results from the computing circuit and base the calculation of the approximate result on these previous results. In this case, the approximate circuit refers to the calculation result history information to calculate the approximate result of the calculation.
[0035] The calculation result history information can include an indication of previous results calculated by a calculation circuit and gradient information indicating how an operation being evaluated to perform a calculation changes in response to inputs to the operation. The operation may be a mathematical function or may include two or more sub-operations such that the operation requires one or more functions that are evaluated as part of the calculation. Thus, the gradient information can reflect the dependency of the overall operation, including the functions forming the sub-operations of the operation, on the inputs to the operation.
[0036] To calculate an approximate result of a particular calculation, the approximation circuit can be configured to estimate the result of the calculation using the corresponding previous result and gradient information and the inputs to be used. For example, the approximation can select a previous result obtained using inputs similar to the inputs for the calculation to be approximated and then identify a gradient indicating how the result of the calculation varies as the input varies. This can then be used to derive an approximate result of the calculation.
[0037] This provides a quick and efficient way to determine an approximation to the result of a calculation that can be compared to the result provided by the calculation circuit. If the result from the calculation circuit and the approximate result differ by more than a deviation threshold, this can be regarded as an indication that an error has occurred, which may be the result of a malfunction in the processing circuit.
[0038] The calculation result history information may be supplemented by the approximation circuit based on the observed results from the calculation circuit. Specifically, the approximation circuit can update the calculation result history in response to the calculation circuit calculating the result of a given calculation and based on that result of the given calculation calculated by the calculation circuit. For example, the approximation circuit can store the result of the given calculation and the input that resulted in that result. Based on the previously obtained results and other results for that calculation, the approximation circuit can also calculate and store new gradient information to be used when calculating the approximate result.
[0039] In some examples, the calculation result history information is updated each time a calculation is performed using the calculation circuit. However, in some examples, this update may be performed only for calculations for which a completeness check is executed.
[0040] The deviation threshold can be obtained in several possible ways. For example, the deviation threshold can be set in a system register, in which case the completeness check circuit can obtain the deviation threshold from the system register. Thus, the deviation threshold used can be modified by changing the value stored in the register.
[0041] In some examples, the approximation circuit not only determines an approximation result but also determines a level of confidence associated with that approximation. If a calculation result history is used to perform the approximation, this confidence level can be determined, for example, based on the difference between one or more inputs for which there were entries in the calculation result history and one or more inputs on which the calculation was based. Regardless of how the confidence level is established, this confidence level can be used to determine the deviation threshold to be used such that the deviation threshold is smaller when the confidence level is higher.
[0042] When a completeness check start instruction and a completeness check end instruction are used to control the completeness check process, at least one of the completeness check start instruction and the completeness check end instruction can cause the processing circuit to specify a particular deviation threshold in response to the execution of this instruction and cause the completeness check circuit to use that particular deviation threshold.
[0043] Here, a specific example will be described with reference to the figures.
[0044] FIG. 1 schematically illustrates an example of a data processing system 2 to which the techniques described herein may be applied. The apparatus 2 has a processing circuit 4 in the form of a processing pipeline 4 that includes a plurality of pipeline stages. In this example, the pipeline stages include a fetch stage 6 for fetching instructions from an instruction cache 8, a decode stage 10 for decoding fetched program instructions to generate micro-operations to be processed by the remaining stages of the pipeline, an issue stage 12 for checking whether the operands required for a micro-operation are available in a register file 14 and issuing a micro-operation for execution when the operands required for a given micro-operation become available, an execution stage 16 for performing a data processing operation corresponding to the micro-operation by processing the operands read from the register file 14 to generate a result value, and a write-back stage 18 for writing the result of the processing back to the register file 14. This is merely an example of a possible pipeline architecture, and it will be understood that other systems may have additional stages or stages of different configurations. For example, in an out-of-order processor, an additional register renaming stage may be included for mapping architectural registers specified by program instructions or micro-operations to physical register specifiers that identify physical registers within the register file 14.
[0045] The execution stage 16 includes several processing units for executing processing operations of different classes. For example, the execution units may include an arithmetic / logic unit (ALU) 20 for performing arithmetic or logical operations, a floating point unit (FPU) 22 for performing operations on floating point values, a branch unit 24 for evaluating the result of a branch operation and adjusting the program counter representing the current execution point accordingly, and a load / store unit 28 for performing load / store operations to access data in the memory systems 8, 30, 32, 34. In this example, the memory system includes a level 1 data cache 30, a level 1 instruction cache 8, a shared level 2 cache 32, and a main system memory 34. This is only an example of a possible memory hierarchy, and it will be understood that other arrangements of caches can be provided. The specific types of processing units 20-28 shown in the execution stage 16 are merely examples, and other implementations may have different sets of processing units, or may include multiple instances of the same type of processing unit so as to be able to process multiple micro-operations of the same type in parallel. FIG. 1 is only a simplified representation of some components of a possible processor pipeline architecture, and it will be understood that the processor can include many other elements not shown for the sake of simplicity, such as a branch prediction mechanism, or an address translation or memory management mechanism.
[0046] Device 2 may have, for example, one or more faults resulting from problems in the manufacture or aging of the system. In particular, the processor may have one or more intermittent faults that occur only occasionally and lead to errors in the calculations performed by device 2.
[0047] FIG. 2 schematically shows a processing circuit 4 provided with a integrity check circuit 50 for checking the calculation results executed by a calculation circuit 40. The calculation circuit 40 can correspond to, for example, the ALU 20 or FPU 22 shown in FIG. 1, or elements thereof. The calculation circuit 40 responds to an instruction executed by the processing circuit 4 that requests a calculation to be executed to calculate one or more results of the calculation. A failure of the processing circuit 4 may lead to the calculation circuit 40 providing an inaccurate result as a result of the calculation. This type of error may otherwise appear to have a valid result at first glance, and since the indication indicating the error is the only thing indicating that the result is incorrect, it may be difficult to detect.
[0048] Therefore, the integrity check circuit 50 is provided to check the results generated by the calculation circuit 40. Instead of causing the calculation circuit 40 to repeat the calculation to check the correctness of the just-executed calculation or providing a replication circuit that can execute a parallel calculation of the results to compare the results, the integrity check circuit 50 is provided with an approximation circuit 60 configured to calculate an approximate result of the calculation, and the result calculated by the calculation circuit 40 can be compared with the approximate result as shown in comparison 64. Since the approximation circuit 60 is not expected to generate an accurate result for the calculation, a deviation threshold is used when comparing (64) the approximate result with the result from the calculation circuit 40, and when the difference between the result from the calculation circuit 40 and the approximation exceeds the deviation threshold, an error that may indicate a detected failure in the processing circuit 4 is shown.
[0049] The above difference may be considered to exceed the deviation threshold when the difference is greater than a certain value, or may be considered to exceed the deviation threshold when the difference is equal to or greater than a certain value.
[0050] As shown in FIG. 2, various mechanisms can be provided to calculate the approximation result. However, the approximation circuit 60 maintains calculation result history information 62 based on the previous results calculated by the calculation circuit 40. Next, the approximation circuit 60 can calculate the approximation result using this calculation result history information 62. For example, the approximation circuit 60 may determine the approximation result based on the fact that the approximation result will be similar to the previous result based on the previous results of calculations including similar operations with similar operands. The calculation result history information 62 may include or be used to calculate gradient information indicating how the operation evaluated to execute the calculation changes according to the input to the operation. Based on this gradient information, the approximation circuit 60 can take into account the difference between the operand used in the previous calculation and the operand for which the approximation result is to be executed.
[0051] The approximation circuit 60 is configured to update the calculation result history information 62 based on the result of the calculation circuit 40. In some cases, the integrity check circuit 50 is not necessarily called to perform an integrity check for all calculations performed by the calculation circuit 40. In this case, the approximation circuit 60 may update the calculation result history information 62 when the calculation is performed even if the integrity check is not performed. However, in some examples, the approximation circuit 60 is configured to update the calculation result history information 62 only when the integrity check is called in order to avoid supplementing the calculation result history information 62 with result information for which the integrity check was not performed.
[0052] Since the calculation result history information 62 is supplemented from the results of the calculation circuit 40 in the example of FIG. 2, the approximation circuit 60 is trained during the use of the device 2, and a "warm-up" period may be experienced in which the calculation result history information 62 is supplemented and the accuracy of the approximation performed by the approximation circuit 60 increases.
[0053] In some cases, the approximation circuit 60 may not be suitable for approximating all the calculations that can be performed by the calculation circuit 40. Therefore, an approximation suitability check circuit 52 is provided to determine for which calculations the approximation circuit 60 can provide a reliable approximation of the result.
[0054] When the integrity check circuit 50 is called for a calculation for which the approximation suitability check circuit 52 has determined that it is not suitable for approximation, the approximation suitability check circuit 52 may suppress the integrity check operation for that calculation. On the other hand, when the approximation suitability check circuit 52 determines that a calculation can be approximated by the approximation circuit 60, the approximation suitability check circuit 52 enables the integrity check operation to proceed.
[0055] The determination of whether a particular calculation can be accurately approximated can be performed in several ways. For example, the approximation suitability check circuit 52 can operate based on smoothness detection 54 and can be configured to determine whether the calculation includes evaluating a function that is mathematically smooth at the points where the function is to be evaluated. Calculating the gradient, such as the dependence of the calculation result on variations in the input when the function is smooth, may become possible. Therefore, when the calculation involves evaluating a smooth function, the approximation circuit 60 may be able to determine the approximation result using the calculation result history information 62 and neighboring results from the gradient information.
[0056] However, in some examples, the approximate fitness check circuit 52 utilizes a list of appropriate operations 56 such that if the calculation consists of appropriate operations for approximation, the approximate fitness check circuit 52 enables the result to be approximated by the approximate circuit 60. In some examples, the approximate fitness check circuit 52 utilizes a neural network, a random forest, or a decision tree 58. These structures may be trained in situ using data observed from the operation of the processing circuit 4, or the neural network / random forest / decision tree may be provided pre-trained such that it can already classify calculations that are to be performed as approximate or otherwise suitable.
[0057] Figure 3 shows an exemplary entry 70 of calculation result history information. The entry 70 represents a single previous calculation executed by the calculation circuit 40 and can be used by the approximate circuit 60 when calculating an approximate result of a further calculation. The entry 70 includes an operation identifier (ID) 72 for identifying a specific operation format of the operation for which the entry 70 represents a result. The entry 70 also includes an indication showing one or more inputs 74 used as operands in the operation represented by the entry 70. Next, the entry 76 shows one or more results 76 generated by that calculation and gradient information 78 indicating how the result of the operation changes in response to the inputs. This gradient information 78 can be calculated, for example, based on the differences between the inputs and results for two previous calculations performed using that operation.
[0058] Figure 4 shows an example of program code adapted to call a completeness check. Figure 4 shows a function for evaluating a power. On the left side of Figure 4, the function signature is shown, and on the right side of Figure 4 are the instructions executed by the processing circuit to execute the function.
[0059] As shown in FIG. 4, this function begins with a completeness check start instruction, i.e., the START_IC instruction. The START_IC instruction indicates to the processing circuit 4 that a completeness check operation is to be performed regarding the calculations executed by subsequent instructions. The START_IC instruction can also indicate the identifier (ID) of a specific element of the completeness check circuit 50 configured to perform a completeness check on the elements of the calculation circuit 40 that are processing the calculation. The START_IC instruction also specifies the number of arguments of the calculation for which the completeness check is to be performed. In the case of this figure, there are two arguments x and y.
[0060] The completeness check circuit 40 may be configured to obtain arguments from a predetermined register (e.g., from the first two general-purpose registers within the register file 14), or in some examples, the START_IC instruction may indicate the location of the arguments.
[0061] When the START_IC instruction is used to indicate the start of a calculation for which a completeness check is to be performed, subsequent instructions (not shown) are used to cause the processing circuit 4, specifically the calculation circuit 40, to execute that calculation. A completeness check end instruction, i.e., the END_ID instruction, is provided to the completeness check circuit 50 to indicate that the calculation for which the completeness check operation is to be performed has been completed. The completeness check circuit 50 is configured to perform a comparison 64 between the result from the calculation circuit 40 and an approximate result in response to the completeness check end instruction.
[0062] Thus, a programmer or compiler can indicate that a completeness check is to be performed for a particular calculation and the range of instructions representing that calculation.
[0063] Figure 5 shows another example of program code adapted to call a completeness check. In the example of Figure 5, the calculation of dividing a by b has two results, a quotient and a remainder, as shown on the left side of Figure 5. Thus, in this example, the instructions for calculating the quotient undergo a separate completeness check operation from the instructions used to calculate the remainder. This can be seen on the right side of Figure 5, where there is a first completeness check start instruction and a first completeness check end instruction with a first identifier id1 placed around the instructions for calculating the quotient, followed by a second completeness check start instruction and a second completeness check end instruction with a second identifier id2 placed around the instructions for calculating the remainder. In this way, the calculation of both results undergoes a completeness check operation and is able to detect any obstacles that would lead to an error during the calculation of either result.
[0064] Figure 6 is a flowchart showing the operation of the completeness check. As shown in Figure 6, at step 602, it is determined whether a completeness check should be executed. This determination can be made based on whether a completeness check start instruction indicating that the completeness check should be executed has been received. Additionally, at step 602, a determination may be made as to whether the calculation to be executed is suitable for approximation. Thus, if the calculation is not suitable for approximation, the completeness check is not executed for that calculation.
[0065] If the completeness check is executed, the calculation circuit 40 proceeds to calculate the result at step 606. The approximation circuit 60 also calculates an approximate result at step 604. These results are then compared at step 608, where it is determined whether the difference between the result and the approximate result exceeds a deviation threshold. If the difference does not exceed the threshold, the calculation is considered to have passed the completeness check. At this time, the calculation result history information 62 may be supplemented using the result of the calculation.
[0066] On the one hand, when the difference between the result and the approximate result exceeds the deviation threshold, in step 610, the integrity check circuit 50 determines that an error has occurred.
[0067] The error can be indicated, for example, by writing a value to a register in the register file 14 or by generating an exception. The error may indicate a fault in the processing circuit 4 and can thus be used as a basis for deciding whether to replace the processing circuit 4 or compensate for the fault in another way. However, since the error determination is based on approximation, a single error determination may not be sufficient to conclusively determine the presence of a fault, and the presence of a fault can only be determined instead when repeated errors are detected by the integrity check circuit 50.
[0068] FIG. 7 illustrates a simulator implementation that may be used. The above-described embodiments implement the present invention in terms of an apparatus and method for operating specific processing hardware that supports the technique, but it is also possible to provide an instruction execution environment according to the embodiments described herein implemented by the use of a computer program. Such a computer program is often referred to as a simulator as long as it provides a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators including dynamic binary translators. Typically, the simulator implementation may be executed on a host processor 714, optionally execute a host operating system 712, and support a simulator program 704. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that execute at a reasonable speed, but such an approach may be justified in certain situations, such as when it is desired to execute native code on another processor for reasons of compatibility or reuse. For example, the simulator implementation may provide an instruction execution environment having additional functionality not supported by the host processor hardware, or may provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is described in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, 1990 Winter USENIX Conference, pages 53-63.
[0069] Previously, embodiments have been described with reference to specific hardware components or features, but in simulated embodiments, equivalent functionality can be provided by suitable software components or features. For example, a specific circuit may be implemented as computer program logic in a simulated embodiment. Similarly, memory hardware such as registers or caches may be implemented as software data structures in simulated embodiments. In arrangements where one or more of the hardware elements referred to in the foregoing embodiments are present in host hardware (e.g., host processor 714), some simulated embodiments may use the host hardware where appropriate.
[0070] The simulator program 704 may be stored on a computer-readable storage medium (which may be a non-transitory medium), but provides a program interface (instruction execution environment) to target code 702 (which may include applications, operating systems, and hypervisors) that is the same as the interface of the hardware architecture modeled by the simulator program 704. Specifically, the simulator program 704 includes simulator code that provides processing program logic 706, calculation program logic 708, and approximation program logic 710 corresponding to the processing circuit 4, the calculation circuit 40, and the approximation circuit 60, respectively. Thus, the program instructions of the target code 702, including the completeness check start instruction and the completeness check end instruction described above, can be executed from within the instruction execution environment using the simulator program 704, and as a result, a host computer 714 that does not actually have the hardware functions of the device 2 described above can emulate these functions.
[0071] A mechanism for quickly and efficiently checking the results of calculations executed by the processing circuit 4, which avoids the need to duplicate the entire calculation logic or execute the calculation more than once, has been described. The integrity check operation can also be selectively used, together with an instruction provided to indicate to the processing circuit 4 whether the integrity check operation should be executed.
Claims
1. An apparatus comprising a processing circuit that executes commands, wherein the processing circuit a calculation circuit that calculates a result of the calculation in response to one or more commands requesting execution of the calculation; an approximation circuit that calculates an approximate result of the calculation independently of the calculation circuit in response to the one or more commands; a completeness check circuit, comparing the result of the calculation executed by the calculation circuit with the approximate result of the calculation executed by the approximation circuit; detecting an error in the processing circuit in response to determining that a difference between the result of the calculation and the approximate result of the calculation is greater than a deviation threshold, and a completeness check circuit configured to perform a completeness check; An apparatus comprising a processing circuit that executes commands.
2. In response to execution of a completeness check start command, the processing circuit causes the approximation circuit to start calculating the approximate result for one or more commands following the completeness check start command, The apparatus according to claim 1, wherein the processing circuit causes the completeness check circuit to perform the completeness check in response to execution of a completeness check end command.
3. The calculation is performed based on one or more operands, The completeness check start command specifies one or more registers, The apparatus according to claim 2, wherein the approximation circuit is configured to obtain the one or more operands from the one or more specified registers.
4. The calculation is performed based on one or more operands, The apparatus according to claim 1 or 2, wherein the approximation circuit is configured to obtain the one or more operands from one or more predetermined registers.
5. The approximation circuit determines whether a particular calculation is suitable for approximation by the approximation circuit, and in response to determining that the particular calculation is suitable for approximation by the approximation circuit, is configured to enable only the completeness check to be performed for the particular calculation. The apparatus according to any one of claims 1 to 4, comprising an approximation suitability check circuit.
6. The apparatus according to claim 5, wherein the approximation suitability check circuit is configured to determine that the particular calculation is suitable for approximation based on determining that the calculation includes evaluating a function that is mathematically smooth at a point at which the function is to be evaluated.
7. The approximate fitness check circuit is configured to determine whether a particular calculation is suitable for approximation by the approximate circuit using at least one of a neural network, a random forest, and a decision tree, for the apparatus according to claim 5 or 6.
8. The approximate fitness check circuit is configured to identify the calculation as being suitable for approximation by the approximate circuit in response to determining that the calculation corresponds to an operation from a predetermined list of operations, for the apparatus according to any one of claims 5 to 7.
9. The approximate circuit maintains calculation result history information based on a result of a previous calculation calculated by the calculation circuit, and is configured to calculate the approximate result of the calculation with reference to the calculation result history information, for the apparatus according to any one of claims 1 to 8.
10. The calculation result history information includes an indication indicating the previous result of the calculation calculated by the calculation circuit, and gradient information indicating how an operation evaluated to execute the calculation changes depending on the input to the operation, for the apparatus according to claim 9.
11. The approximate circuit is configured to calculate the approximate result of the calculation based on the corresponding previous result and the gradient information among the previous results from the calculation result history information, for the apparatus according to claim 10.
12. The approximate circuit updates the calculation result history information based on the result of the given calculation in response to the calculation circuit calculating the result of the given calculation, for the apparatus according to any one of claims 9 to 11.
13. Updating the calculation result history includes storing the result of the given calculation, calculating new gradient information based on the result of the given calculation and the previous result, and storing the new gradient information, for the apparatus according to claim 12.
14. The calculation circuit is an arithmetic logic unit (ALU) or a part of an ALU, for the apparatus according to any one of claims 1 to 13.
15. The calculation circuit is a floating point unit (FPU) or a part of an FPU, for the apparatus according to any one of claims 1 to 13.
16. The apparatus according to any one of claims 1 to 15, wherein the integrity check circuit is configured to obtain the deviation threshold value from a system register.
17. The approximation circuit is configured to determine a level of reliability associated with the approximation, The apparatus according to any one of claims 1 to 15, wherein the deviation threshold value is based on the level of reliability.
18. The apparatus according to claim 2, or, when dependent on claim 2, the apparatus according to any one of claims 3 to 15, wherein the processing circuit causes the integrity check circuit to use the specific deviation threshold value as the deviation threshold value in response to execution of at least one of the integrity check start command and the integrity check end command that specify the specific deviation threshold value.
19. A method comprising: executing instructions by a processing circuit; calculating a result of the calculation in response to one or more instructions that require the calculation to be executed; calculating an approximate result of the calculation independently of the calculating, in response to the one or more instructions; performing an integrity check; comparing the result of the calculation and the approximate result of the calculation; detecting an error in the processing circuit in response to determining that a difference between the result of the calculation and the approximate result of the calculation is greater than a deviation threshold value; and performing by. A method including.
20. A computer program for controlling a host data processing device to provide an instruction execution environment, the computer program including: processing program logic for executing instructions, wherein the processing program logic includes: calculation program logic for calculating a result of the calculation in response to one or more instructions that require the calculation; approximate program logic for calculating an approximate result of the calculation independently of the calculation program logic in response to the one or more instructions; integrity check program logic for performing an integrity check by: comparing the result of the calculation executed by the calculation program logic and the approximate result of the calculation executed by the approximate program logic; detecting an error in the processing program logic in response to determining that a difference between the result of the calculation and the approximate result of the calculation is greater than a deviation threshold value; and integrity check program logic configured to perform by. A computer program including