Validation of integrity confirmation information

Asynchronous extension processing circuitry in data processing systems validates integrity confirmation information to counter electromagnetic fault injection, enhancing detection and prevention of attacks while improving system security and efficiency.

WO2025210333A1PCT designated stage Publication Date: 2025-10-09ARM LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/050544
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-03-17
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing data processing systems are vulnerable to electromagnetic fault injection attacks, particularly dual probe electromagnetic fault injection, which is difficult to detect and correct, and existing methods to mitigate such attacks are inefficient and predictable.

Method used

Implementing extension processing circuitry that operates asynchronously to the data processing pipeline, performing computational tasks independently and validating integrity confirmation information to compare against further instances, making it difficult for attackers to synchronize fault injection.

Benefits of technology

Enhances the ability to detect and prevent electromagnetic fault injection by introducing asynchronous computational tasks and integrity validation, reducing the likelihood of successful attacks and improving system security and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025050544_09102025_PF_FP_ABST
    Figure GB2025050544_09102025_PF_FP_ABST
Patent Text Reader

Abstract

There is provided an apparatus, a method, and a computer program. The apparatus is provided with a data processing pipeline configured to perform data processing, and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline. The extension processing circuitry is configured to perform the computational task asynchronously to the data processing operations. The data processing pipeline comprises decoding circuitry responsive to one or more integrity confirmation instructions to trigger the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task. The integrity confirmation information is suitable for comparison against confirmation data obtained from one or more further instances of the computational task.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] VALIDATION OF INTEGRITY CONFIRMATION INFORMATION

[0002] This invention relates to data processing. Furthermore, this invention relates to an apparatus, a method, and a computer program.

[0003] The validation that a computational task has been performed correctly provides improved confidence in the apparatus performing the computational task and allows the apparatus to identify inconsistencies in processing that may arise, for example, due to malicious activity aiming to disrupt processing operations of the apparatus or obtaining secure data held by the apparatus.

[0004] According to some configurations of the present techniques there is provided an apparatus comprising: a data processing pipeline configured to perform data processing; and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline, the extension processing circuitry configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline, wherein the data processing pipeline comprises decoding circuitry responsive to one or more integrity confirmation instructions to trigger the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task, the integrity confirmation information suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

[0005] According to some configurations of the present techniques there is provided a method of operating an apparatus comprising a data processing pipeline configured to perform data processing, and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline, the extension processing circuitry configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline, the method comprising: in response to one or more integrity confirmation instructions, triggering the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task, the integrity confirmation information suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

[0006] According to some configurations of the present techniques there is provided a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: data processing pipeline program logic configured to perform data processing; and extension processing program logic associated with the data processing pipeline program logic and configured to perform a computational task delegated by the data processing pipeline program logic in response to a delegation signal received from the data processing pipeline program logic, the extension processing circuitry configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline program logic, wherein: the data processing pipeline program logic comprises decoding program logic responsive to one or more integrity confirmation instructions to trigger the extension processing program logic to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task, the integrity confirmation information suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

[0007] In some configurations the computer program is stored on a computer readable storage medium.

[0008] In some configurations the computer readable storage medium is a non-transitory computer readable storage medium. The present techniques will be described further, by way of example only, with reference to configurations thereof as illustrated in the accompanying drawings, in which:

[0009] Figure 1 schematically illustrates an apparatus according to some configurations of the present techniques;

[0010] Figure 2 schematically illustrates an apparatus according to some configurations of the present techniques;

[0011] Figure 3 schematically illustrates an apparatus according to some configurations of the present techniques;

[0012] Figure 4 schematically illustrates operation of an apparatus according to some configurations of the present techniques;

[0013] Figure 5 schematically illustrates an electromagnetic fault injection probe according to some configurations of the present techniques;

[0014] Figure 6 schematically illustrates an electromagnetic fault injection probe according to some configurations of the present techniques;

[0015] Figure 7 schematically illustrates an apparatus according to some configurations of the present techniques;

[0016] Figure 8 schematically illustrates an apparatus according to some configurations of the present techniques;

[0017] Figure 9 schematically illustrates an apparatus according to some configurations of the present techniques;

[0018] Figure 10 schematically illustrates signals between a data processing pipeline and extension processing circuitry according to some configurations of the present techniques;

[0019] Figure 11 schematically illustrates signals between a data processing pipeline and extension processing circuitry according to some configurations of the present techniques;

[0020] Figure 12 schematically illustrates signals between a data processing pipeline and extension processing circuitry according to some configurations of the present techniques;

[0021] Figure 13 schematically illustrates an apparatus according to some configurations of the present techniques;

[0022] Figure 14 schematically illustrates an apparatus according to some configurations of the present techniques;

[0023] Figure 15 schematically illustrates signals between a data processing pipeline and extension processing circuitry according to some configurations of the present techniques; Figure 16 schematically illustrates signals between a data processing pipeline and extension processing circuitry according to some configurations of the present techniques;

[0024] Figure 17 schematically illustrates signals between a data processing pipeline and extension processing circuitry according to some configurations of the present techniques;

[0025] Figure 18 schematically illustrates an apparatus according to some configurations of the present techniques;

[0026] Figure 19 schematically illustrates an apparatus according to some configurations of the present techniques;

[0027] Figure 20 schematically illustrates an apparatus according to some configurations of the present techniques;

[0028] Figure 21 schematically illustrates a sequence of steps carried out by an apparatus according to some configurations of the present techniques; and

[0029] Figure 22 schematically illustrates a simulator implementation of an apparatus according to some configurations of the present techniques.

[0030] Electromagnetic fault injection is a technique for injecting faults into devices. Strong magnetic fields may be applied in the vicinity of electronic circuits to introduce transient bit flips or to permanently set the state of registers and / or memories. Such attacks may be difficult to detect and / or correct for their effects. One approach that could be used in relation to such attacks is to run the same code twice on two cores and compare the information about the results and / or trace data generated by the two cores. Whilst such techniques could, in principle, work for a single electromagnetic fault, they would still be vulnerable to dual probe electromagnetic fault injection. Dual probe electromagnetic fault injection is a technique made possible by improved resolution of the electromagnetic probes in used in electromagnetic fault injection and relies on the application of a same fault into each of the two cores.

[0031] According to some configurations of the present techniques there is provided an apparatus comprising: a data processing pipeline configured to perform data processing, and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline. The extension processing circuitry is configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline. The data processing pipeline comprises decoding circuitry responsive to one or more integrity confirmation instructions to trigger the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task. The integrity confirmation information is suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

[0032] An apparatus comprising a data processing pipeline can be required to perform a limitless variety of data processing operations as defined by the sequence of instructions provided to it. In order to efficiently perform those data processing operations, the data processing pipeline may be configured with a variety of functional units, each with a given specialised type of data processing ability, such as arithmetic logic units (ALUs), floating point (FP) units, load / store units, and so on. Yet even with such specialised functional units being provided as part of the data processing pipeline, some apparatuses may be provided with extension processing circuitry that is configured to perform particular tasks or functions that are frequently executed. Examples of such functions or tasks include operations to transfer data in memory (e.g., memcpy), or to initialise an area of memory (e.g., memset), compression, encryption, and string processing, although the present techniques are not limited to these particular examples.

[0033] The extension processing circuitry is associated with the data processing pipeline and is configured to perform such a function (a delegated task) in response to a delegation signal received from the data processing pipeline. Such extension processing circuitry may also be referred to as a threadlet extension (TE) herein. The sequence of operations carried out by the extension processing circuitry to perform the defined function may also be referred to as a threadlet herein. The extension processing circuitry is closely associated (tightly coupled) with the data processing pipeline both in terms of its physical location and the in terms of resources shared by the extension processing circuitry and the processing pipeline. Nevertheless, the extension processing circuitry is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline. The data processing pipeline may also be referred to as the CPU herein. Threadlets are functions or collections of operations that can be executed asynchronously relative to other CPU activity once launched. The asynchronous operation of the extension processing circuitry with respect to the data processing pipeline is possible because, unlike some prior art techniques, the extension processing circuitry receives a directive or command from the thread currently executing on the CPU and performs the required operations independently, that is without requiring a stream of instructions from the CPU that directly control or influence its internal operation. The CPU is therefore free to continue executing other code and potentially reduce overall runtime by overlapping the execution of the instruction stream after the directive or command is sent to the extension processing circuitry with the operation of the extension processing circuitry. The directive or command sent to the extension processing circuitry to initiate the delegated task may be generated in response to an extension start instruction defined for this purpose in the instruction set of the data processing pipeline. Accordingly, the decoding circuitry may be responsive to the extension start instruction to issue the delegation signal to the extension processing circuitry to delegate the delegated task to the extension processing circuitry. Because of the tight integration of the extension processing circuitry with the data processing pipeline, the extension processing circuitry can be launched rapidly and its state can be checked in a short amount of time (e.g. of the order of a few ns) relative to some prior art techniques, which would require a great many CPU cycles for launching commands or performing synchronisation operations.

[0034] Because the extension processing circuitry is configured to perform the computational task asynchronously to the data processing operations of the data processing pipeline, the manner in which the computational task is performed is not constrained by the physical hardware that is provided as part of the data processing pipeline. Indeed, the operations carried out by the extension processing circuitry when performing the computational task may differ from the operations carried out by the processing pipeline when performing the same computational task.

[0035] The data processing pipeline is responsive to integrity confirmation instructions to control the extension processing circuitry to perform one or more operations to validate the integrity confirmation information obtained by the extension processing circuitry when performing the computational task. The one or more operations may comprise operations to obtain the integrity confirmation data and or to validate the integrity confirmation data once obtained. The integrity confirmation instructions are instructions that are part of an Instruction Set Architecture (ISA). The ISA defines the set of instructions that can be used by a programmer or compiler to control the processing circuitry. Instructions in the ISA are decoded by decoding circuitry which is provided to receive the instructions of the ISA and to generate control signals to control the processing circuitry. Instructions included in the ISA are carefully chosen and are limited by the encoding space provided for the instructions. Often the inclusion of a new instruction may result in an existing instruction being removed and, as such, the designer of the ISA may require significant motivation to incorporate new instructions into the ISA.

[0036] An apparatus implementing a given ISA has freedom to choose the precise circuitry arrangements that are provided for each instruction in the ISA and the manner in which instructions of the ISA are implemented may vary from system to system. In other words, the ISA defines the function that the apparatus should achieve in response to a given instruction, but provides freedom to the designer to implement the instruction in any way and according to particular design constraints associated with the apparatus. Provision of a carefully chosen instruction as part of an ISA provides advantages in terms of improved code density and or added functionality which can improve overall system performance.

[0037] The integrity confirmation instructions are provided as part of the ISA and allow a programmer or compiler to control the data processing pipeline to trigger the validation operations performed by the extension processing circuitry. The ISA may also be provided with one or more instructions to control the processing circuitry to trigger the computational task and to probe the extension processing circuitry for confirmation as to whether the computational task has been completed. Whilst the instructions in the ISA are performed by the processing circuitry according to a sequence of instructions (which may be performed in-order or out-of-order dependent on the configuration of the processing pipeline), the timing and operation of the computational task performed by the extension processing circuitry is asynchronous to these instructions and is not directly under the control of the programmer or compiler. The operations performed by the extension processing circuitry may operate in a manner that is different to the instructions set out in the ISA. The inventors have recognised that the dual core electromagnetic fault injection technique uses both the physical separation of the points at which the faults are injected and the similarities between the hardware into which the faults are injected in order to cause the same fault in multiple instances of a hardware task. The provision of extension processing circuitry that is asynchronous to the processing pipeline makes it difficult for an attacker to identify a timing for injecting a fault. This reduces a likelihood that dual probe electromagnetic fault injection will be successful.

[0038] The extension processing circuitry is configured to obtain integrity confirmation information whilst performing the computational task. This integrity confirmation information is suitable for comparison against confirmation data that is obtained from one or more further instances of the computational task. Where it is determined that the integrity confirmation information matches the confirmation data, it can be confirmed that the computational task and the one or more further instances of the computational task executed in a same manner and therefore the likelihood of a fault being injected is low. On the other hand, where it is determined that the integrity confirmation information does not match the confirmation data, it can be determined that at least one of the computational task or the one or more further instances of the computational task did not execute correctly. The integrity confirmation information may take a variety of forms and may represent result data, intermediate result data, or one or more parameters indicative of the computational task being performed (e.g., information relating to addresses written to or read from by the extension processing circuitry whilst performing the computational task). The provision of extension processing circuitry configured to operate in this manner enables the apparatus to identify instances of electromagnetic fault injection, or other erroneous computations (e.g., caused by particle strikes).

[0039] In some configurations the one or more further instances comprises an instance of the computational task performed by the processing pipeline. In such configurations, the computational task is performed on both the extension processing circuitry and the processing pipeline. The computational task performed on the extension processing circuitry may be triggered before, during, or after the computational task performed on the processing pipeline. Because the extension processing circuitry operates asynchronously to the processing pipeline, the performance of the computational task by the extension processing circuitry may be further delayed relative to the performance of the instance of the computational task performed by the processing pipeline. Using two separate hardware structures to perform the computational task introduces further differences between the manner in which the computational tasks are performed introducing additional difficulties for an attacker attempting to use an electromagnetic probe for electromagnetic fault injection due to differences in the physical layout of the data processing pipeline and the extension processing circuitry.

[0040] In some configurations one or more further instances of processing operations comprises a previous instance of the computational task performed by the extension processing circuitry. Using the extension processing circuitry to perform repeat instances of the computational task, both asynchronously to the processing operations performed by the data processing pipeline, introduces further temporal ambiguity in the processing performed to complete the computational task. Because both instances of the computational task are performed asynchronously, it will be increasingly difficult for an attacker to identify when to trigger a fault using electromagnetic fault injection. Using the extension processing circuitry to perform multiple instances of the computational task allows the data processing pipeline to perform one or more different processing operations in parallel to the performance of the multiple instances of the computational task performed by the extension processing circuitry which can improve the throughput of some use cases.

[0041] The control of the extension processing circuitry by the data processing pipeline may be achieved in a variety of ways. In some configurations the one or more integrity confirmation instructions comprise an extension start integrity confirmation instruction specifying the computational task; and the decoding circuitry is responsive to the extension start integrity confirmation instruction, as one of the one or more operations, to clear the integrity confirmation information, and to generate the delegation signal to trigger the extension processing circuitry to perform the computational task. The extension start integrity confirmation instruction may be provided as an explicit integrity confirmation instruction that is distinct from a delegation start instruction (e.g., that may be provided to start a delegated task on the extension processing circuitry for which integrity confirmation is not required). Alternatively, the extension start integrity confirmation instruction may be combined with the delegation start instruction with a specific parameter, e.g., an immediate value, specified in the delegation start instruction to indicate to the decoding circuitry that the delegation start instruction is the extension start integrity confirmation instruction. The specific parameter may alternatively be provided in a control register or as a control flag that can be set in advance, for example, by privileged software executing instructions on the data processing pipeline.

[0042] In some configurations the decoding circuitry is responsive to the extension start integrity confirmation instruction to trigger the data processing pipeline to perform a further instance of the computational task. The extension start integrity confirmation instruction can therefore be used as a single instruction to trigger the extension processing pipeline to perform the computational task asynchronously to the data processing pipeline and to cause the data processing pipeline to perform the further instance of the computational task. The instance and the further instance of the computational task may be performed sequentially with one or more further tasks performed in between the instance and the further instance. Alternatively, the instance and the further instance of the computational task may be performed consecutively without any intervening tasks being performed by the extension processing circuitry.

[0043] In some configurations the extension start integrity confirmation instruction triggers a branch to instructions comprised in the further instance. The data processing pipeline is responsive to the extension start integrity confirmation instruction to trigger the extension processing circuitry and then to branch to a branch target address, for example, identified in the extension start integrity confirmation instruction, at which the instructions that comprise the further instance are located. The extension start integrity confirmation instruction may be an unconditional branch instruction that causes the data processing pipeline to unconditionally branch to the instructions comprised in the further instance. Alternatively, the extension start integrity confirmation instruction may be encoded as a conditional branch instruction where the branch is triggered or not triggered based on a specific flag indicating whether or not the extension start integrity confirmation instruction is enabled or disabled.

[0044] In some configurations the extension processing circuitry is responsive to the delegation signal to perform plural instances of the computational task. The delegation signal, triggered in response to the extension start integrity confirmation instruction therefore provides a single instruction which, when received by the decoder circuitry of the processing pipeline causes the processing pipeline to trigger multiple instances of the computational task. In some configurations the extension processing circuitry is configured to store the integrity confirmation information for each of the plural instances. For example, the integrity confirmation information for each of the plural instances can be stored and, subsequent to the extension processing circuitry performing the plural instances of the computational task, compared to one another. Alternatively, the integrity confirmation information may be stored for the two most recent instances of the plural instances and, where a further instance of the computational task is being performed, prior to the further instance, the integrity confirmation information for the two most recent instances can be compared with an error being indicated if the two most recent instances result in different integrity confirmation information.

[0045] In some configurations the extension processing circuitry is responsive to the delegation signal to implement a variable length delay between each of the plural instances. The inclusion of a variable length delay adds an additional level of unpredictability to the timing at which the extension processing circuitry performs the computational task. The length of the variable length delay may be based on any random or pseudo-random metadata stored by the data processing pipeline or the extension processing circuitry. For example, the length of the variable length delay may be based on a clock value, a hash of a program counter value currently being used by the data processing pipeline, and / or a hash of one or more data values stored in the extension processing circuitry.

[0046] Whilst the number of the plural instances may be defined in a variety of different ways, in some configurations the extension processing circuitry is configured to perform the plural instances until the data processing pipeline has completed the further instance computational task. Typically, the processing of the computational task by the extension processing circuitry will take less time that the processing of the computational task by the data processing pipeline. This is because the extension processing circuitry is typically purpose built to perform one or more functions efficiently. As a result, the extension processing circuitry may be able to perform several instances of the computational task whilst the data processing pipeline is performing the further instance. In some alternative configurations, the number of the plural instances may be defined in a register, or specified as an immediate value in the extension start integrity confirmation instruction. In some configurations, the number of the plural instances may be determined based on both the number of instances that can be performed before the data processing pipeline has performed the further instance, and a number specified as an intermediate value in the extension start integrity confirmation instruction. For example, the number of the plural instances may, in some configurations be set to the minimum of the number of instances that can be performed before the data processing pipeline has performed the further instance and the number specified as an intermediate value. Alternatively, the number of the plural instances may, in some configurations be set to the maximum of the number of instances that can be performed before the data processing pipeline has performed the further instance and the number specified as an intermediate value.

[0047] In some configurations the apparatus comprises a plurality of extension processing circuits, the plurality of extension processing circuits comprising the extension processing circuitry; and the decoding circuitry is responsive to the extension start integrity confirmation instruction, to clear the integrity confirmation information for two or more extension processing circuits of the plurality of extension processing circuits, and to generate the delegation signal to trigger the two or more extension processing circuits to each perform the computational task. The extension processing circuits may each be associated with, e.g., closely coupled with, a same data processing pipeline. Alternatively, at least one of the two or more extension processing circuits may be associated with a different data processing pipeline. Each of the plurality of extension processing circuits operates asynchronously to the data processing pipeline and may operate asynchronously to one another. The two or more extension processing circuits may be specified, for example, using an intermediate parameter in the extension start integrity confirmation instruction. Alternatively, the two or more extension processing circuits may be identified in a control register associated with the data processing pipeline.

[0048] Whilst the delegation signal may be provided to each of the extension processing circuits at the same time or substantially the same time, in some configurations the decoding circuitry is configured to implement a delay between triggering each of the two or more extension processing circuits. The delay may for instance be a variable length delay or a fixed length delay. The delay may be based on a counter, a hash of a program counter value, and / or a hash of one or more data values stored in registers of the data processing pipeline. The delay implemented by the decoding circuitry may be in addition to a further delay implanted by the extension processing circuitry which maybe either a fixed length delay or a variable length delay.

[0049] In some configurations the decoding circuitry is responsive to a compare extension processing instruction to generate control signals to cause the data processing pipeline to: trigger each extension processing circuit of the two or more extension processing circuits to return the integrity confirmation information maintained by that extension processing circuit to the data processing pipeline; and compare the integrity confirmation information returned by each extension processing circuit, and in response to a determination that the integrity confirmation information returned by at least one of the two or more extension processing circuits differs from another of the two or more extension processing circuits, to signal a fault. The compare extension processing instruction is an instruction of the instruction set architecture. The two or more extension processing circuits may be specified, for example, using an intermediate parameter in the compare extension processing instruction. Alternatively, the two or more extension processing circuits may be identified in a control register associated with the data processing pipeline. In some configurations alternative, the comparison of extension processing circuitry may be configured to automatically return (e.g., broadcast) the validation information subsequent to performing the computational task, i.e., as a response to the extension start integrity confirmation instruction. This information may then be stored in temporary storage circuitry associated with the data processing pipeline or one of the extension processing circuits to be subsequently compared in response to the compare extension processing instruction.

[0050] In some configurations the decoding circuitry is responsive to a compatibility check instruction to generate compatibility check control signals; and the data processing pipeline is responsive to the compatibility check control signals to trigger at least one of the plurality of extension processing circuits to indicate whether the at least one of the plurality of extension processing circuits is capable of maintaining the integrity confirmation information. Some apparatuses may be provided with at least one instance of extension processing circuitry that is not capable of maintaining integrity confirmation instruction, for example, because the at least one instance of the extension processing circuitry is of an older design and / or arranged to minimise circuit area or power consumption. The provision of the compatibility check instruction allows the programmer or compiler to determine, for a given instance of the plurality of extension processing circuits, which of those extension processing circuits are capable of maintaining the integrity confirmation information.

[0051] In some configurations the data processing pipeline is configured to store selection information indicative of whether an integrity confirmation state is in an enabled state or a disabled state for each of the plurality of extension processing circuits; the two or more extension processing circuits comprises each of the plurality of extension processing circuits indicated as having integrity confirmation enabled; and the decoding circuitry is responsive to an integrity confirmation control instruction specifying at least one of the plurality of extension processing circuits and a new integrity confirmation state, to update the integrity confirmation state for that one of the plurality of extension processing circuits to the new integrity confirmation state. The selection information therefore provides an indication of which of the plurality of extension processing circuits is currently in an enabled state, i.e., is able to store integrity confirmation information whilst performing a delegated task. The selection information may be stored, for example, in a control register having a plurality of bits, each associated with one of the plurality of extension processing circuits. Bits within the control register may be set to a first value (e.g., one of a logical zero and a logical one) to indicate that the extension processing circuit associated with that bit is able to generate integrity confirmation information, and may be set to a second value (e.g., the other of a logical zero and a logical one) to indicate that the extension processing circuitry associated with that bit is not able to generate integrity confirmation information. The integrity confirmation control instruction allows a programmer or compiler to change the selection information associated with each of the plurality of extension processing circuits. In some configurations, where one of the plurality of extension processing circuits is unable to generate integrity confirmation information, the selection information associated with that extension processing circuit may be hard wired to indicate that the integrity confirmation state is disabled for that extension processing circuit.

[0052] Whilst the integrity confirmation information may be returned by the extension processing circuitry subsequent to completion of the computational task, in some configurations the one or more integrity confirmation instructions comprise a return integrity information instruction; and the decoding circuitry is responsive to the return integrity information instruction to trigger the extension processing circuitry, as one of the one or more operations, to return a current value of the integrity confirmation information relating to a previously delegated task to the data processing pipeline. The extension processing circuitry may store the integrity confirmation information for each instance of the computational task performed on that extension processing circuitry. Alternatively, where multiple instances of the computational task are performed by the extension processing circuitry, the extension processing circuitry may automatically, e.g., in response to completion of each instance of the computational task, perform a comparison between instances of the integrity confirmation information. In such configurations, the extension processing circuitry may, where all instances of the integrity confirmation information are the same, store a single instance of the integrity confirmation information and, where one or more instances of the integrity confirmation information differs from the others, the extension processing circuitry may store an error code. The data processing pipeline may perform one or more actions in response to the error code being returned in response to the mismatched integrity information instruction. For example, the data processing pipeline may trigger an exception and / or branch to an error handling code.

[0053] In some configurations the extension processing circuitry is responsive to returning the current value, to disable storing of the integrity confirmation information. The extension processing circuitry may subsequently operate in the disabled state until the storing of integrity confirmation information is subsequently enabled. In some configurations, the return integrity information instruction may encode information indicating that the extension processing circuitry should disable storing of the integrity confirmation information subsequent to returning the current value. The information may be provided as an immediate parameter in the return integrity information instruction or through one or more different encodings of the return integrity information instruction.

[0054] In some configurations the data processing pipeline is responsive to a determination that the integrity confirmation information differs from the confirmation data, to generate a fault signal.

[0055] The fault signal may take a variety of forms and, in some configurations the fault signal comprises at least one of raising an exception to interrupt a current process; triggering the current process to branch to a fault handling code sequence; and / or triggering an indication of the fault signal to be stored. The use of a fault handling sequence, code, or mechanism, and / or storing the fault signal may allow the fault to be dealt with at a later stage of processing or may be used to signal to higher privileged software that a fault has occurred. In some configurations, the fault handler may be used to trigger a further one or more instances of the computational task to determine if the fault is a reoccurring fault.

[0056] In some configurations the extension processing circuitry comprises an accumulation register to store the integrity confirmation information, and the extension processing circuitry is configured to accumulate information indicative of result data generated during the computational task into the accumulation register. The integrity confirmation data is therefore accumulated from the result data and is dependent on the results produced by the extension processing circuitry rather than how those results are produced. The integrity confirmation data is therefore independent of precise sequence of operations performed in order to achieve the computational task.

[0057] In some configurations the information indicative of result data comprises a hash of the result data. The information indicative of the result data may be hashed along with existing data stored in the accumulation register to generate a hashed value that is dependent on the current result data and all previous result data generated during the computational task.

[0058] Whilst the extension processing circuitry may take any form, in some configurations the extension processing circuitry comprises a further data processing pipeline, the further data processing pipeline having a reduced processing capability relative to the data processing pipeline. Hence, whilst the extension processing circuitry may be provided as specialist processing circuitry configured to perform one or more specific processing tasks, in some configurations, the extension processing circuitry is capable of performing a range of processing tasks. The provision of a further data processing pipeline as the extension processing circuitry may provide additional degrees of reliability, flexibility and security.

[0059] In some configurations the extension processing circuitry and the data processing pipeline are each implemented using a different hardware configuration. For example, the physical layout of components required for each of the extension processing circuitry and the data processing pipeline may be different. In addition, the functional operation of the hardware used for the extension processing circuitry and for the data processing pipeline may be different with some functionality that is only available to the data processing pipeline and some functionality that is only available to the extension processing circuitry.

[0060] In some configurations the extension processing circuitry is configured to perform the computational task using a first sequence of processing operations different to a second sequence of processing operations required to perform the computational task using the data processing pipeline. The provision of extension processing circuitry that has a different hardware configuration and / or that performs a different sequence of processing operations to the data processing pipeline is advantageous in protecting against dual probe electromagnetic fault injection. Because the sequence of operations performed by the data processing pipeline and the extension processing circuitry are different, an electromagnetic field applied to each of these circuits will likely result in a different fault being triggered. For example, the electromagnetic probe generating an electromagnetic field in the vicinity of the data processing pipeline may trigger a plurality of bits in the data processing pipeline to be set resulting in a first fault being propagated throughout the data processing pipeline. However, a second electromagnetic probe generating the same electromagnetic field in the vicinity of the extension processing circuitry may cause a different set of bits to be set resulting in a different (second) fault to be propagated through the processing operations performed by the extension processing circuitry. Generally, each of these faults (the first fault and the second fault) may result in a different change in the integrity confirmation information that is output by the data processing apparatus and the extension processing circuitry. When the integrity confirmation information is subsequently verified, the difference resulting from the first and second fault can be identified and a fault can be indicated.

[0061] In some configurations the extension processing circuitry and the data processing pipeline are fabricated as a single integrated circuit with at least some physical components of the data processing pipeline provided in close physical proximity to one another. In contrast to CPUs (data processing pipelines), which are often provided in separate, non-overlapping, portions of a chip with each area corresponding to circuitry of a single CPU, threadlets (extension processing circuitry) are tightly coupled to components within the host CPU and may be provided with components that overlap and / or are shared with the CPU. The close physical proximity of the CPU and the extension processing circuitry increases the difficulty of arranging electromagnetic probes in such a way as to deterministically interfere with both the CPU and the extension processing circuitry. For example, the components of the CPU and the extension processing circuitry may be closer to one another than the resolution of the EM probes which may, for example, be less than 500 micrometres, e.g., less than 400 micrometres, 200 micrometres, or 100 micrometres. In some configurations the extension processing circuitry shares, or has direct access to, one or more functional components of the data processing pipeline. For example, the extension processing circuitry may have direct access to a load unit and / or a store buffer provided as part of the data processing pipeline. Alternatively, or in addition, the data extension processing circuitry may share at least part of a translation lookaside buffer or cache structure associated with the data processing pipeline. In some configurations, the extension processing circuitry is configured to share the data processing pipeline’s path to memory.

[0062] Particular configurations of the present techniques will now be described in relation to the accompanying figures.

[0063] Figures 1 to 4 schematically illustrate details of a data processing apparatus provided with extension processing circuitry according to some configurations of the present techniques.

[0064] Figure 1 schematically illustrates a data processing apparatus 10 according to some examples. The data processing apparatus 10 is schematically shown to have a pipelined configuration, which for the purposes of brevity and clarity is shown in a conceptual representation here. The illustrated pipeline stages comprise an instruction cache 11, a fetch stage 12, a decode stage 13, a micro-op cache 14, an issue stage 15, and a register access stage 16. A sequence of instructions is retrieved from memory (not shown) and cached in the instruction cache 11. The fetch stage 12 controls which instructions are retrieved as the sequence of instructions and these instructions are then decoded in the decode stage 13. This decoding essentially identifies the type of each instruction, as well as any further operands specified by the instruction, and generates control signals to control the remainder of the apparatus to perform the data processing operation(s) defined by the instruction. Decoding the instructions may comprise splitting an instruction into one or more micro-ops, and these micro-ops can be cached in the micro-op cache 14. The final stage of the pipeline before execution is the issue stage 15, where instructions (or micro-ops) are queued pending the availability of the register values they specify as operands and the corresponding functional unit of the data processing pipeline which will carry out the defined operation. Generally, the data processing operation(s) defined by the instructions are carried out by the functional units that form part of the data processing pipeline, namely the load / store unit 17, the execute unit 18, and the execute unit 19. These latter execute units may for example be arithmetic logic units (ALUs), floating point units (FPUs), and so on. The functional units that form part of the data processing pipeline perform their data processing operations on data values which are provided from a set of registers (conceptually represented by the register access stage 16 in the figure) and result values of those data processing operations are returned to the set of registers. The load / store unit 17 is provided for the purpose of storing values from the set of registers to the memory system, of which only a level 1 cache 21 and a level 2 cache 22 are shown in the figure. The LI cache 21 is private to the data processing apparatus 10 and the L2 cache 22 may be shared with another data processing apparatus, when part of a wider data processing system. The data processing apparatus 10 is also shown to comprise a branch unit 20, which monitors execution flow of the sequence of instructions and seeks to predict, based on previous execution history, whether a given branch will be taken or not. The predictions from the branch unit 20 inform the sequence of instructions caused to be fetched by the fetch stage 12.

[0065] The data processing apparatus 10 further comprises extension processing circuitry 23, which is provided to support efficient performance of one or more defined functions, which have been established to be impactful and ubiquitous for the data processing operations which this data processing apparatus 10 carries out. Example functions of this type have been found to include tasks or functions such as memcpy, memset, compression, encryption, and string processing, although the present techniques are not limited to these particular examples. The extension processing circuitry is closely associated with the data processing pipeline and is configured to perform the defined function (also referred to herein as a delegated task) in response to a delegation signal received from the data processing pipeline. The extension processing circuitry 23 is an example of a threadlet extension (TE) according to the present techniques. The sequence of operations it carries out to perform the defined function is referred to as a threadlet herein. The extension processing circuitry 23, although closely associated with the data processing pipeline, is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline. The data processing pipeline may also be referred to as the CPU herein. Threadlets are functions or collections of operations that can be executed asynchronously relative to other CPU activity once launched. The directive or command sent to the extension processing circuitry 23 to initiate the delegated task is generated in response to an extension start instruction defined for this purpose in the instruction set of the data processing pipeline. Thus, an extension start instruction progresses along the data processing pipeline in the manner that any other CPU instruction would, but when the decoding circuitry 13 identifies the extension start instruction it can signal directly to the extension processing circuitry 23. The close integration of the extension processing circuitry 23 with data processing pipeline is illustrated by the fact that the extension processing circuitry 23 has direct access to the load / store unit 17, and thus it shares the data processing pipeline’s path to memory. The extension processing circuitry 23 also has access to the set of registers 16, such that for example, the extension start instruction can specify one or more registers as operands, and the values from these registers are then passed directly to the extension processing circuitry 23 in association with the command sent to initiate the delegated task. Upon completion of the task, results of the delegated task can be returned to the register values via an extension synchronisation instruction.

[0066] Figure 2 schematically illustrates a data processing apparatus 30 according to some examples. It will be noted that the arrangement of components of the data processing apparatus 30 is similar to that of the components of the data processing apparatus 10 shown in Figure 1. One difference is that whilst the data processing apparatus 10 of Figure 1 is intended to represent an in-order processor, the data processing apparatus 30 is an out-of-order processor. As one consequence of this the data processing pipeline of the data processing apparatus 30 comprises a rename stage 35, allowing the data processing apparatus 30 to vary the order in which it executes instructions of the sequence of instructions, such that they can be executed in an order dictated by when their operands become available, and the availability of functional units, rather than the order in which they appear in the sequence. The illustrated pipeline stages comprise an instruction cache 31, a fetch stage 32, a decode stage 33, a micro-op cache 34, the rename stage 35, an issue stage 36, and a register access stage 37. A sequence of instructions is retrieved from memory (not shown) and cached in the instruction cache 31. Instructions pass through the data processing pipeline in the manner described above with reference to the data processing apparatus 10 of Figure 1, with the further register renaming that is performed by the rename stage 35. The functional units of the data processing pipeline in this example are the load unit 38, the store unit 39, the FPU

[0067] 41, the integer ALU 42, and the vector unit 43. The throughput of the FPU 41, the integer ALU

[0068] 42, and the vector unit 43 is sufficient that a result cache 44 is provided an intermediary before results of their data processing are returned to the registers 37. A branch prediction unit 45 is also provided and its predictions inform the operation of the fetch stage 32.

[0069] The data processing apparatus 30 further comprises extension processing circuitry (“threadlet extension”) 49, which is provided to support efficient performance of one or more defined functions, which have been established to be impactful and ubiquitous for the data processing operations which this data processing apparatus 30 carries out. The extension processing circuitry 49 is closely associated with the data processing pipeline and is configured to perform the defined function in response to a delegation signal received from the data processing pipeline. In the example of Figure 2, this delegation signal is shown emanating from the issue queue stage 36. Notably, this is after the rename stage 35, such that the extension processing circuitry 49 can operate with respect to the physical registers of the set of registers 37 according to the same mapping of architectural registers used for the rest of the apparatus. As in the example of Figure 1, the data processing pipeline (instruction cache 31 through to the register read stage 37, the load / store units 38 and 39, and the functional units 41-45) may also be referred to as the CPU. The threadlet extension 49 operates asynchronously relative to other CPU activity once launched. The directive or command sent to the extension processing circuitry 49 to initiate the delegated task is generated in response to an extension start instruction defined for this purpose in the instruction set of the data processing pipeline. The close integration of the extension processing circuitry 49 with data processing pipeline also apparent in this example by the fact that the extension processing circuitry 49 has direct access to the load unit 38 and the store buffer 40, and thus it shares the data processing pipeline’s path to memory. The extension processing circuitry 49 also has access to the set of registers 37, such that for example, the extension start instruction can specify one or more registers as operands, and the values from these registers are then passed directly to the extension processing circuitry 49 in association with the command sent to initiate the delegated task. Note that the output of the branch prediction unit 45 is also provided to the extension processing circuitry 49. Upon completion of the task, results of the delegated task can be returned to the register values via an extension synchronisation instruction.

[0070] Figure 3 schematically illustrates a data processing apparatus 50 according to some examples. This example provides a comparison to the examples of Figure 1 and Figure 2, in which examples the extension processing circuitry was closely embedded with the data processing pipeline, to the extent that those instances of extension processing circuitry may be considered to be within the CPU. In the example apparatus 50 of Figure 3, the CPU 51 and the extension processing circuitry (threadlet extension) 52 are not as closely integrated. For example this is illustrated by the fact that each has its own path to memory, with an LI cache 53 private to the CPU 51 and an LI cache 54 private to the threadlet extension 52. They share the L2 cache 55. Nevertheless, the threadlet extension 52 remains tightly coupled to the CPU 51, and can be launched quickly when an extension start instruction is encountered in the CPU pipeline specifying the function this threadlet extension 52 performs. The threadlet extension 52 can get data directly from CPU registers at the start of its execution. Upon completion, it can return values via an extension synchronisation instruction. Figure 3 also shows the threadlet extension 52 as having its own private TLB 56, in which it can cache currently used address translations. As a preparatory step before or associated with the delegation signal, content from the TLB 57 in the CPU 51 can be copied into the private TLB 56 in order to pre-warm this cache before the threadlet begins operation.

[0071] Figure 4 is a state diagram illustrating an example set of states between which extension processing circuitry (TE) transitions in some examples. Initially the TE is in an IDLE state 60. When an extension start (XSTART) instruction is encountered by the data processing pipeline, a delegation signal can cause the TE to switch to the SETUP state 61. This may also require a signal indicating that the XSTART instruction has been committed to be asserted. In the SETUP state 61, certain actions necessary for preparing the TE can be performed, for example, in examples in which the TE has a separate path to memory (as in the case of Figure 3), one setup task is the transfer of relevant entries currently in the CPU’s TLB to a private TLB within the TE. This enables the TE to perform translations independently at a faster rate than if it were to rely entirely on the existing translation mechanism within the CPU. If the TE has been in a clock-gated or power-gated condition when in the IDLE state 60, the SETUP state 61 may also comprise the task of exiting the TE from that clock-gated or power-gated condition. Once the SETUP state 61 is complete the TE can switch to the RUNNING state 62. If the TE encounters a memory fault during its processing, it asserts a signal which will raise an interrupt within the CPU, causing it to stop executing the main thread and switch to a handler. The TE switches to the INTERRUPTED state 63. An indication of the source of the fault is placed in a special syndrome system register, the address associated with the fault is stored in the fault address system register, and a bit in the Program Status Register (PSR) will be set enabling the handler to quickly determine the source of the fault. Setting a bit in the PSR makes communicating the resumption of the threadlet straightforward, because the handler can reset the relevant bit in the SPSR and when the CPSR is restored from the SPSR during exception return, the TE can detect the resetting of this bit and resume executing. The TE will also switch to the INTERRUPTED state 63 if the main thread gets switched out, e.g. during a context- switch initiated by the operating system. In the INTERRUPTED state 63, the TE may be clock-gated or power-gated, unless some other thread launches a new command directed at it or the associated thread returns resumes execution or the handler returns. The TE returns from the INTERRUPTED state 63 to the RUNNING state 62 via the RELOAD state 64 in which any context or state relevant to its execution, which was previously saved to memory, can be restored. This might be the case if another thread made use of a TE which was previously interrupted. Finally, when the extension reaches the end of the offloaded granule of computation (the delegated task) it moves to the IDLE state 60. The TE will advertise completion of the task, so that an extension synchronisation instruction (XSYNC) can pick up that “done” signal and, if required, provide a return value to a specified register. If the TE has any lingering data in its private caches it might also need to flush these entries upon completion.

[0072] An example of using threadlets is now set out. The programmer or compiler identifies functions whose execution in custom hardware (extension processing circuitry) satisfies the costbenefit thresholds in their use-case. An instruction (such as XSTART) is used to launches a command within the designated CPU extension. An example use written in pseudo-code (for such an identified function “funcX”) is as follows: funcA () {

[0073] XSTART {xO - x3}, #imm_op / / funcX(a, b, c, d);

[0074] II

[0075] 12

[0076] 13

[0077] 14

[0078] XSYNC xO, #imm_op }

[0079] Thus, within the function funcA, the XSTART instruction initializes the CPU extension and transfers to the extension processing circuitry the parameters (a, b, c, d) for funcX, which are in registers xO, xl, x2, x3 respectively. The XSTART instruction in this example also specifies the immediate value #imm_op, which defines the specific function to be carried out. For example, whilst there might only be one instance of extension processing circuitry, it may be capable of performing more than one function, or at least more than one variant of a function, and the immediate value #imm_op can select the desired variant and / or function. In other examples there may be more than one instance of extension processing circuitry and the immediate value #imm_op can select between them. Depending on the setup, the extension could also automatically get a copy of relevant entries in the TLB. The extension processing circuitry then carries out the task required (funcX) and during its execution, the CPU is free to carry on executing other instructions II, 12, 13, 14, etc. At some point in the future, the CPU executes an extension synchronisation instruction (XSYNC) which automatically checks whether the extension has completed or not. If it has not, for some variants of the extension synchronisation instruction, the CPU will wait for the delegated task to complete. Other variants of the extension synchronisation instruction (e.g. the XSYNCS variant) allow the CPU can carry on executing other code (if there are alternative routines available or stop executing and wait for completion of the extension (typically if there is nothing else to execute in the interim). There are a range of variations of XSTART and XSYNC proposed herein, and these are discussed in more detail with reference to the figures which follow.

[0080] Figure 5 schematically illustrates the technique of dual probe electromagnetic fault injection. Dual probe electromagnetic fault injection relies on the use of strong localised electromagnetic fields, generated at the tips of probes, to cause transient bit flips or to permanently set one or more bits stored in an electronic circuit implemented on a chip 70 (e.g., a silicon chip or a flexible integrated circuit printed on a flexible substrate). The chip 70 may implement a number of different electronic circuit designs which may operate either independently or in combination with one another. In the illustrated configuration, the chip 70 implements two separate circuitry instances, including a first circuit 76 and a second circuit 78. The first circuit 76 and the second circuit are separated from one another by a physical separation distance. The first circuit 76 and the second circuit 78 may communicate with one another through one or more communication channels or via one or more additional circuits (not illustrated) that are implemented on the chip 70.

[0081] An attacker seeking to use an electromagnetic probe to inject a fault could attempt to do so by positioning a first electromagnetic fault injection probe 72 with its tip at a location near to the first circuit 76. By carefully manipulating the probe placement and experimenting with timing of the attack, a skilled attacker may be able to manipulate the first electronic circuit either to operate in an unexpected way or to reveal secret information. In order to mitigate against use of a single electromagnetic fault injection probe, such as the first electromagnetic fault injection probe 72, one approach could be to use the second circuit 78 operating in a same way as the first circuit 76 to repeat one or more computational tasks performed by the first circuit 76 and comparing the results of these circuits or information relating to the performance of the operations carried out by these circuits. If the information stored by both of the first circuit 76 and the second circuit 78 differ from one another then it can be identified that a fault has occurred in at least one of the first circuit 76 and the second circuit 78.

[0082] Dual probe electromagnetic fault injection is a technique which, due to the increasing resolution of electromagnetic fault injection probes, could be used to circumvent this approach through the use of a second electromagnetic injection probe 74 in addition to the first electromagnetic fault injection probe 72. The second electromagnetic fault injection probe 72 is used to inject a same fault in the second circuit 78 as the one injected by the first fault injection probe 72 in the first circuit 76. This technique has the potential to successfully circumvent / thwart the mitigation approach described in relation to the use of only the first electromagnetic fault injection probe and could be successful so long as the physical location (e.g., the spatial position) at which the second fault needs to be injected is sufficiently far from the physical location of the position at which the first fault is injected. In other words, where the first circuit 76 and the second circuit 78 are spaced such that the physical distance between the first circuit 76 and the second circuit 78 is greater than the probe resolution. Here the probe resolution is defined by the minimum distance that must be provided between the first electromagnetic fault injection probe 72 and the second electromagnetic fault injection probe 74 in order that the superposition of the two fields does not cause additional or alternative changes to the state of the electronic circuits than the effects caused by each of the two electromagnetic fault injection probes operating in isolation.

[0083] In the illustrated configuration, the distance between the first electromagnetic fault injection probe 72 and the second electromagnetic fault injection probe 74 is smaller than the distance between the first circuit 76 and the second circuit 78 and it would be possible that a same fault could be injected in each of the first circuit 76 and the second circuit 78. As a result, the dual probe electromagnetic fault injection technique could potentially be used to trick the combination of the first circuit 76 and the second circuit 78.

[0084] Figure 6 schematically illustrates a circuit layout according to some configurations of the present techniques. Figure 6 illustrates a chip 80 having a first circuit 86 and a second circuit 88. The first circuit 86 may comprise a data processing pipeline according to some configurations of the present techniques. The second circuit 88 may comprise extension processing circuitry according to some configurations of the present techniques. The first circuit 86 and the second circuit 88 are physically located much closer to one another on the chip 80 and share some components with one another (illustrated by the overlap in the area of the first circuit 86 and the second circuit 88 in figure 6). As described in relation to figures 1 and 2, the data processing pipeline and the extension processing circuitry would have such a close association between the first circuit 86 and the second circuit 88.

[0085] As a result of the small physical separation between the first circuit 86 and the second circuit 88, the distance at which the faults would have to be injected in a dual probe electromagnetic fault injection attack is greatly reduced and, in the illustrated configuration, is smaller than the resolution of the first electromagnetic fault injection probe 82 and the second electromagnetic fault injection probe 84. The close association between the first circuit 86 and the second circuit 88 can therefore be used as a mitigation technique against a dual probe attack.

[0086] Figure 7 schematically illustrates decoding circuitry 92 responsive to an integrity confirmation instruction 90. Here the integrity confirmation instruction 90 is an XSTART AND CONFIRM instruction that is received by the decoding circuitry 92 and takes the form: XSTART AND CONFIRM {x0 - x7}, #imm. The XSTART AND CONFIRM instruction is a variant on the XSTART instruction described hereinabove and provides the same arguments as would be seen in a standard XSTART instruction. When decoded by the decoding circuitry, the XSTART AND CONFIRM instruction causes the content of registers x0-x7 to be retrieved from the registers 94 and passed to the extension processing circuitry 96. In this case the extension processing circuitry 96 can perform multiple types of operation (task) and the immediate value #imm (or signals based on the immediate value #imm) may be used to select between them. Unlike a standard XSTART instruction, the XSTART AND CONFIRM instruction also triggers the extension processing circuitry 96 to perform the following operations in relation to integrity confirmation information 98 stored in the extension processing circuitry 96:

[0087] • Clear the integrity confirmation information 98; and

[0088] • Cause the address and data of each store committed by the extension processing circuitry 96 to be hashed into the integrity confirmation information 98.

[0089] Figure 8 schematically illustrates the response of decoding circuitry 102 within a processing pipeline 100 to an integrity confirmation instruction 104. In the illustrated configuration the integrity confirmation instruction 104 is an XGET CONFIRM instruction. The decoding circuitry 102 is responsive to the XGET CONFIRM instruction to issue a request signal to the extension processing circuitry 106 requesting the integrity confirmation information 108 to be returned to the data processing pipeline. Asynchronously to the computational task performed by the extension processing circuitry 106 (e.g., before the XGET CONFIRM instruction was issued to the decoding circuitry 102), the data processing pipeline 100 may also perform an instance of the computational task. Results of the computational task performed by the data processing pipeline 100 are not saved to memory (to avoid duplication of the results from the extension processing circuitry 106 that are saved to memory) but confirmation data is obtained by storing an equivalent hash of the addresses and data that would otherwise have been stored to the memory into the confirmation data. Subsequent to the return of the integrity confirmation information 108 from the extension processing circuitry 106, the data processing pipeline can perform one or more validation checks to determine whether the integrity confirmation information 108 matches the confirmation data.

[0090] It will be readily apparent to the skilled person that, in alternative configurations, it may be the data processing pipeline 100 rather than the extension processing circuitry 106 that saves the data to memory. Furthermore, in some alternative configurations the confirmation data may be generated by processing circuitry other than the data processing pipeline 100, for example, one or more additional extension processing circuits or a further instance of the computational task implemented on the extension processing circuitry 106.

[0091] Figure 9 schematically illustrates an example of the generation of integrity confirmation data by the extension processing circuitry 110. The extension processing circuitry comprises an execution unit 112 which provides intermediate result data to hash circuitry 114. The intermediate result data may, for example, comprise at least a portion of addresses and data to be stored to memory. The intermediate result data is hashed using hash circuitry 114 which is passed to summation circuitry 116. The summation circuitry 116 retrieves a current value of the integrity confirmation information 118 and accumulates the result of the hash circuitry 114 into the integrity confirmation information 118. It will be readily apparent to the person skilled in the art that the circuitry illustrated in figure 9 is for exemplary purpose and that alternative circuits could be provided for accumulating the integrity confirmation information. For example, the intermediate result data could be hashed with the integrity confirmation information 118 e.g., using an XOR circuit.

[0092] Figure 10 schematically illustrates a sequence of steps carried out by a data processing pipeline and extension processing circuitry over a number of computational cycles in accordance with some configurations of the present techniques. The data processing pipeline receives an XSTART AND CONFIRM instruction in the form described in relation to figure 7. The data processing pipeline is responsive to the XSTART AND CONFIRM instruction to issue a delegation signal to the extension processing circuitry. The extension processing circuitry is responsive to receipt of the delegation signal to perform a first sequence of processing operations that comprise a computational task. The first sequence of processing operations is performed asynchronously to the processing operations carried out by the data processing pipeline.

[0093] In addition to issuing the delegation signal, the data processing pipeline causes a branch in the flow of instructions to instructions identifying a second sequence of operations that are performed by the data processing pipeline and that comprise the computational task. The computational task performed by the data processing pipeline and the extension processing pipeline is the same with the exception that the data values generated are not saved to memory (to avoid duplication of results). However, the first sequence of processing operations that are carried out by the extension processing circuitry is a different sequence of operations to the second sequence of processing operations that are carried out by the data processing pipeline. This difference is due to the different micro-architectural features that are present in the data processing pipeline, which is designed in order to achieve a wide range of different functions, and the extension processing circuitry, which is a more specialist piece of hardware designed to perform a smaller subset of computational tasks.

[0094] Subsequent to the completion of the second sequence of processing operations performed by the data processing pipeline, an XGET CONFIRM instruction (as described in relation to figure 8) is received by the data processing pipeline. The data processing pipeline issues a request for the integrity confirmation information from the extension processing circuitry. The extension processing circuitry is responsive to the request for integrity confirmation information and returns the integrity confirmation information to the data processing pipeline which is then able to compare the two in order to validate that the computational task has been performed in a manner that is consistent between the data processing pipeline and the extension processing circuitry.

[0095] It is noted that, as a more specialised piece of hardware, the extension processing pipeline is likely to complete the first sequence of operations in advance of the second sequence of operations performed by the data processing pipeline. In the illustrated configuration, the extension processing circuitry completes the first sequence of processing operations and switches to an idle state to await further delegation signals. If the data processing pipeline were to request the accumulator prior to the completion of the first sequence of processing operations by the extension processing circuitry, the extension processing circuitry may respond with a busy signal or other indication that the first sequence of processing operations is still being performed.

[0096] Figure 11 schematically illustrates a sequence of steps carried out by a data processing pipeline and extension processing circuitry over a number of computational cycles in accordance with some configurations of the present techniques. The data processing pipeline receives an XSTART AND CONFIRM instruction in the form described in relation to figure 7. The data processing pipeline is responsive to the XSTART AND CONFIRM instruction to issue a delegation signal to the extension processing circuitry. The extension processing circuitry is responsive to receipt of the delegation signal to perform a first sequence of processing operations that comprise a computational task. The first sequence of processing operations is performed asynchronously to the processing operations carried out by the data processing pipeline. Subsequent to issuing the delegation signal, in the illustrated configuration the data processing pipeline continues to perform one or more other sequences of processing operations.

[0097] Once the first sequence of processing operations is complete, the extension processing circuitry stores the integrity confirmation information in an internal buffer or register(s). The extension processing circuitry then proceeds to repeat the first sequence of processing operations to generate confirmation data. During the repeat instance of the first sequence of processing operations, the extension processing circuitry omits sending the result data to memory (to avoid repetition of results already generated during the first instance of the first sequence of processing operations and to reduce memory bandwidth usage).

[0098] At some point, subsequent to performing the one or more other sequences of processing operations, the data processing pipeline receives the XGET CONFIRM instruction which triggers the data processing pipeline to request the integrity confirmation information from the extension processing circuitry. The extension processing circuitry is responsive to the request to return the integrity confirmation information for both instances of the first sequence of processing operation. The data processing pipeline receives the integrity confirmation information and performs a comparison between the integrity confirmation information, using one of the two sets of integrity confirmation information as confirmation data.

[0099] Figure 12 schematically illustrates a sequence of steps carried out by a data processing pipeline and extension processing circuitry over a number of computational cycles in accordance with some configurations of the present techniques. The data processing pipeline receives an XSTART AND CONFIRM instruction in the form described in relation to figure 7. The data processing pipeline is responsive to the XSTART AND CONFIRM instruction to issue a delegation signal to the extension processing circuitry. The extension processing circuitry is responsive to receipt of the delegation signal to perform a first sequence of processing operations that comprise a computational task. The first sequence of processing operations is performed asynchronously to the processing operations carried out by the data processing pipeline.

[0100] Once the first sequence of processing operations is complete, the extension processing circuitry stores the integrity confirmation information in an internal buffer or register(s). As in the case of figure 11, the extension processing circuitry then proceeds to repeat the first sequence of processing operations to generate confirmation data. During the repeat instance of the first sequence of processing operations, the extension processing circuitry omits sending the result data to memory (to avoid repetition of results already generated during the first instance of the first sequence of processing operations and to reduce memory bandwidth usage). The extension processing circuitry proceeds to repeat the sequence including the buffering of integrity confirmation information multiple times until the request for integrity confirmation information is received from the data processing pipeline and at least one instance of the first sequence of operations has been completed.

[0101] In addition to issuing the delegation signal, the data processing pipeline causes a branch in the flow of instructions to instructions identifying a second sequence of operations that are performed by the data processing pipeline and that comprise the computational task. The computational task performed by the data processing pipeline and the extension processing pipeline is the same with the exception that the data values generated are not saved to memory (to avoid duplication of results). As a result, the computational task is performed multiple times, at least once by the extension processing circuitry performing the first sequence of processing operations and at least once by the data processing pipeline performing the second sequence of operations.

[0102] Subsequent to the second sequence of operations (and optional subsequent to one or more additional processing tasks performed by the data processing pipeline), the data processing pipeline receives the XGET CONFIRM instruction and issues a request for the integrity confirmation information to the extension processing circuitry. The extension processing circuitry is responsive to the request for the integrity confirmation to first perform a check that at least one instance of the first sequence of operations has been completed. If the at least one instance of the first sequence of operations has been completed, then the extension processing circuitry aborts the current repetition and returns the integrity confirmation information to the data processing pipeline. If the extension processing circuitry has not completed at least one instance of the first sequence of processing operations, then the extension processing circuitry completes the instance of the first sequence of processing operations before the integrity confirmation information for the completed repetitions is returned to the data processing pipeline. The data processing pipeline then performs a comparison between the integrity confirmation information resulting from the second sequence of processing operations and the at least one instance of the first sequence of processing operations.

[0103] In some alternative configurations the extension processing circuitry is responsive to the delegation signal to implement a variable length delay between each of the plural instances of the first sequence of processing operations. The length of the variable length delay may be based on any random or pseudo-random metadata stored by the data processing pipeline or the extension processing circuitry. For example, the length of the variable length delay may be based on a clock value, a hash of a program counter value currently being used by the data processing pipeline, and / or a hash of one or more data values stored in the extension processing circuitry.

[0104] Figure 13 schematically illustrates decoding circuitry 122 responsive to an integrity confirmation instruction 120. Here the integrity confirmation instruction 120 is an XSTART LS instruction that is received by the decoding circuitry 122 and takes the form: XSTART LS {x0 - x7}, #imm, EP0, EPl. The XSTART LS instruction is a variant on the XSTART instruction and the XSTART AND CONFIRM instruction described hereinabove and provides the same arguments as would be seen in a standard XSTART instruction along with additional parameters EP0 and EPL The additional parameters EP0 and EPl each identify an instance of extension processing circuitry. The first additional parameter EP0 identifies extension processing circuitry EP0 126 and the second additional parameter EPl identifies the extension processing circuitry EPl 130. The decoding circuitry 122 is responsive to the XSTART LS instruction to issue the delegation signal to the instances of extension processing circuitry identified in the integrity confirmation instruction 120. In particular, when decoded by the decoding circuitry, the XSTART LS instruction causes the content of registers x0-x7 to be retrieved from the registers 124 and passed to the extension processing circuitry EP0 126 and to the extension processing circuitry EPl 130. In addition to causing the extension processing circuitry EPO 126 and the extension processing circuitry EPl 130 to perform the computational task, the decoding circuitry is further configured to cause the integrity confirmation information 128 stored by the extension processing circuitry EPO 126 to be cleared and the integrity confirmation information 132 stored by the extension processing circuitry EPl 130 to be cleared. Each of the extension processing circuitry EPO 126 and the extension processing circuitry EPl 130 is responsive to the delegation signal to perform the computational task and to update the integrity confirmation information.

[0105] Figure 14 schematically illustrates the response of decoding circuitry 144 within a processing pipeline 140 to an integrity confirmation instruction 142. In the illustrated configuration the integrity confirmation instruction 142 is an XCOMP LS instruction taking the form XCOMP LS EPO, EPl where the parameters EPO and EPl identify instances of extension processing circuitry to be targeted by the XCOMP LS instruction. The decoding circuitry 144 is responsive to the XCOMP LS instruction to request integrity confirmation information 148 from the extension processing circuitry EPO 146 as identified by the parameter EPO provided in the integrity confirmation instruction 142. The extension processing circuitry EPO 146 is responsive to receipt of the request for the integrity confirmation information 148 to return the integrity confirmation information 148 to the data processing pipeline 140. The decoding circuitry 144 is also responsive to the XCOMP LS instruction to request integrity confirmation information 152 from the extension processing circuitry EPl 150 as identified by the parameter EPO provided in the integrity confirmation instruction 142. The extension processing circuitry EPl 150 is responsive to receipt of the request for the integrity confirmation information 152 to return the integrity confirmation information 152 to the data processing pipeline 140. The data processing pipeline 140 is responsive to receipt of the integrity confirmation information to perform a comparison between the integrity confirmation information 148 received from the extension processing circuitry EPO 146 and the integrity confirmation information 152 obtained from the extension processing circuitry EPl 150. The data processing pipeline may trigger an error or record an indication of a fault in response to a discrepancy between the integrity confirmation information 148 returned by the extension processing circuitry EPO 146 and integrity confirmation information 152 returned by the extension processing circuitry EPl 150. Figure 15 schematically illustrates a sequence of steps carried out by extension processing circuitry EPO and extension processing circuitry EPl in response to an XSTART LS instruction specifying EPO and EPl received by the data processing pipeline. The data processing pipeline is responsive to the XSTART LS instruction to issue a delegation signal to each of the extension processing circuitry EPO and the extension processing circuitry EPl. The extension processing circuitry EPO is responsive to receipt of the delegation signal to perform a first sequence of processing operations comprising a computational task and to store integrity confirmation information relating to the performance of the computational task. The extension processing circuitry EPl is responsive to receipt of the delegation signal to perform a second sequence of processing operations comprising a computational task and to store integrity confirmation information relating to the performance of the computational task. The first and second sequences of processing operations may be identical sequence of processing operations or different sequences of processing operations dependent on the implementation of the extension processing circuits. The extension processing circuitry EPl and the extension processing circuitry EPO operate asynchronously to the data processing pipeline and asynchronously to on another.

[0106] Subsequent to issuing the delegation signal, the data processing pipeline may perform one or more other sequences of processing operations (e.g., unrelated to the computational task being carried out by the extension processing circuitry EPO and the extension processing circuitry EPl). At some subsequent point, the data processing pipeline receives an XCOMP LS instruction specifying the extension processing circuitry EPO and the extension processing circuitry EPl. The data processing pipeline is responsive to the XCOMP LS instruction to issue a request to the extension processing circuity EPO and the extension processing circuitry EPl requesting that integrity confirmation information is returned by the extension processing circuitry EPO and the extension processing circuitry EPL The extension processing circuitry EPO and the extension processing circuitry EPl each respond by returning the integrity confirmation information which is compared by the data processing pipeline to determine if there is any discrepancy between the integrity confirmation information returned by the extension processing circuitry EPO and the extension processing circuitry EPL

[0107] Figure 16 schematically illustrates a variant on the behaviour of the data processing pipeline in response to the XSTART LS instruction specifying the extension processing circuitry EPO and the extension processing circuitry EPl. Rather than issuing the delegation signal to the extension processing circuitry EPO and the extension processing circuitry EPl without delay, the data processing pipeline issues the delegation signal to the extension processing circuitry EPO without delay and implements a delay before issuing the delegation signal to the extension processing circuitry EPl. The extension processing circuitry EPO and the extension processing circuitry EPl proceed to perform the computational task as described in relation to figure 15. The response of the data processing pipeline to the XCOMP LS instruction proceeds as described in relation to figure 15.

[0108] Figure 17 schematically illustrates a further variant on the behaviour of the data processing pipeline, the extension processing circuitry EPO and the extension processing circuitry EPl in response to the XSTART LS instruction. Whilst in figures 15 and 16, the comparison between the integrity confirmation information maintained by each of extension processing circuitry EPO and extension processing circuitry EPl was carried out by the data processing pipeline in response to an XCOMP LS instruction. In the configuration illustrated in figure 17, the extension processing circuitry EPO is triggered before the extension processing circuitry EPl, which is triggered after a delay implemented by the data processing pipeline. The extension processing circuitry EPO performs the first sequence of processing operations and stores integrity confirmation data in relation to results obtained during the performance of the first sequence of processing operations. Subsequent to completion of the first sequence of processing operations, the extension processing circuitry EPO broadcasts completion of the computational task. The extension processing circuitry EPl receives the broadcast signal from the extension processing circuitry EPO whilst it is performing the second sequence of processing operations. Once the extension processing circuitry EPl completes the second sequence of processing operations, the extension processing circuitry EPl requests the integrity confirmation information from the extension processing circuitry EPO. The extension processing circuitry EPO responds to the request for integrity confirmation information by returning the integrity confirmation to the extension processing circuitry EPl which performs the comparison between the integrity confirmation information generated by the extension processing circuitry EPO and the integrity confirmation information (confirmation data) generated by the extension processing circuitry EPl to determine if there is any discrepancy between the integrity confirmation information returned by the extension processing circuitry EPO and the extension processing circuitry EPl. In the event of a discrepancy, the extension processing circuitry EPl may store information indicative of the discrepancy, trigger an exception, or otherwise signal / broadcast the discrepancy to the data processing pipeline. The data processing pipeline may then respond either directly to the exception or subsequently retrieve the stored information / respond to a signalled / broadcast discrepancy.

[0109] It will be readily apparent to the skilled person that in alternative configurations, the delegation signal to the processing circuitry EPO may be delayed in addition, or as an alternative, to the delay of the delegation circuitry issued to extension processing circuitry EPl. Furthermore, whilst two sets of extension processing circuitry have been illustrated, in other configurations further sets of extension processing circuitry may also be provided and specified in the XSTART LS instruction.

[0110] Figure 18 schematically illustrates decoding circuitry 172 responsive to an integrity confirmation instruction 170. Here the integrity confirmation instruction 170 is an integrity check instruction XCHECK specifying an identifier of extension processing circuitry. In the illustrated configuration, the XCHECK instruction is of the form XCHECK EPID where EPID provides the identifier of extension processing circuitry EP2 176(2). Extension processing circuitry EPO 176(0), extension processing circuitry EPl 176(1) and extension processing circuitry EP2 176(2) are provided and are each configured to store integrity confirmation information 174. The decoding circuitry 172 is responsive to the XCHECK instruction to trigger the identified extension processing circuitry, in this case extension processing circuitry EP2 176(2) to indicate whether it is capable of maintaining integrity confirmation information. In the illustrated configuration, extension processing circuitry EP2 176(2) is capable of maintaining integrity confirmation information and returns a positive affirmation in response to the trigger from the decoding circuitry 172. It will be readily apparent to the skilled person that, in the event that the extension processing circuitry EP2 176(2) was unable to maintain integrity confirmation information, the extension processing circuitry EP2 176(2) may have returned a negative indication or no indication in response to the trigger from the decoding circuitry 172.

[0111] Figure 19 schematically illustrates decoding circuitry 180 responsive to an integrity confirmation instruction 182. Here the integrity confirmation instruction 182 is an integrity state modification instruction XSTATE taking the form XTATE EPID, State. Here the EPID is an identifier of the extension processing circuitry, and the parameter State is an indication of a state that the extension processing circuitry identified by the identifier EPID should be set to. Extension processing circuitry EPO 186(0), extension processing circuitry EPl 186(1) and extension processing circuitry EP2 186(2) are provided and are each configured to store integrity confirmation information 184. The apparatus stores extension state information 188 indicative of a current state of the extension processing circuitry 186. The extension state information takes one of two values. In the illustrated configuration a state value of 1 indicates that the extension processing circuitry is in a state capable of generating integrity confirmation information 184. A value of 0 indicates that the extension processing circuitry is currently in a state in which integrity confirmation information 184 will not be generated. In the illustrated configuration, the extension processing circuitry EPO 186(0) and the extension processing circuitry EPl 186(1) are each in a state in which they will generate integrity confirmation information 184. The extension processing circuitry EP2 186(2) is currently in a state in which it will not generate integrity confirmation information 184. The decoding circuitry 180 is responsive to the XSTATE instruction to trigger the state information 188 relating to the state specified by the identifier EPID to be updated to the state specified in the XSTATE instruction.

[0112] Figure 20 schematically illustrates a configuration of an apparatus in accordance with some configurations of the present techniques. The apparatus comprises a data processing pipeline 190 and extension processing circuitry 194. The data processing pipeline is provided with decoding circuitry 192. The extension processing circuitry 194 is provided with storage circuitry to store integrity confirmation information 196 and with a reduced capability processing pipeline 198. The decoding circuitry 192 may be responsive to any of the integrity confirmation instructions described above. The reduced capability pipeline 198 has a reduced capability relative to the data processing pipeline 190 but may be provided with logical circuity capable of performing any, or the majority of, processing operations that can be performed by the data processing pipeline. The reduced capability pipeline 198 can therefore be used to generate the integrity confirmation information 196 for a wide variety of possible computational tasks. The hardware provisions for the reduced capability pipeline 198 and therefore the precise sequence of operations performed by the reduced capability pipeline 198 differ from the data processing pipeline 190. Whilst such configurations provide advantages in terms of the computational tasks that can be verified, the use of the reduced capability pipeline may be inefficient in terms of circuit area, power usage and the number of cycles taken to generate the confirmation data.

[0113] Figure 21 schematically illustrates a sequence of steps carried out in accordance with some configurations of the present techniques. Flow begins at step S210 where it is determined if an extension start integrity confirmation instruction is received. If, at step S210, it is determined that the extension start integrity confirmation instruction has not been received, then flow remains at step S210. If, at step S210, it is determined that an extension start integrity confirmation instruction has been received, then flow proceeds to step S212. At step S212 the extension processing circuitry is triggered to clear integrity confirmation information before flow proceeds to step S214. At step S214, the extension processing circuitry is triggered to perform a first sequence of processing operations to perform a computational task and to maintain integrity confirmation information before flow proceeds to step S216. At step S216, instruction flow is caused to branch to instructions comprising the computational task to cause the data processing pipeline to perform a second sequence of processing operations based on the instructions.

[0114] Figure 22 schematically illustrates a simulator implementation that may be used. Whilst the earlier described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware supporting the techniques concerned, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein which is implemented through the use of a computer program. Such computer programs are often referred to as simulators, insofar as they provide a software based implementation of a hardware architecture. Varieties of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 730, optionally running a host operating system 720, supporting the simulator program 710. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations which execute at a reasonable speed, but such an approach may be justified in certain circumstances, such as when there is a desire to run code native to another processor for compatibility or re-use reasons. For example, the simulator implementation may provide an instruction execution environment with additional functionality which is not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 USENIX Conference, Pages 53 - 63.

[0115] To the extent that embodiments have previously been described with reference to particular hardware constructs or features, in a simulated embodiment, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuitry may be implemented in a simulated embodiment as computer program logic. Similarly, memory hardware, such as a register or cache, may be implemented in a simulated embodiment as a software data structure. In arrangements where one or more of the hardware elements referenced in the previously described embodiments are present on the host hardware (for example, host processor 730), some simulated embodiments may make use of the host hardware, where suitable.

[0116] The simulator program 710 may be stored on a computer-readable storage medium (which may be a non-transitory medium), and provides a program interface (instruction execution environment) to the target code 700 (which may include applications, operating systems and a hypervisor) which is the same as the interface of the hardware architecture being modelled by the simulator program 710. Thus, the program instructions of the target code 700 may be executed from within the instruction execution environment using the simulator program 710, so that a host computer 730 which does not actually have the hardware features of, for example, the apparatus 10 discussed in relation to figure 1, or the apparatus 30 described in relation to figure 2 can emulate these features.

[0117] In particular, the simulator program 710 comprises data processing pipeline logic 740 and extension processing program logic 750. The extension processing program logic 750 is configured to perform a computational task using a first sequence of processing operations different to a second sequence of processing operations required to perform the computational task using the data processing pipeline program logic 740. The data processing pipeline program logic 740 comprises decoding program logic responsive to one or more integrity confirmation instructions to trigger the extension processing program logic to perform one or more operations to maintain integrity confirmation information indicative of the first sequence of data processing operations, the integrity confirmation information suitable for comparison of the first sequence of processing operations against confirmation data obtained from one or more further sequences of processing operations that comprise the computational task

[0118] In brief overall summary there is provided an apparatus, a method, and a computer program. The apparatus is provided with a data processing pipeline configured to perform data processing, and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline. The extension processing circuitry is configured to perform the computational task asynchronously to the data processing operations. The data processing pipeline comprises decoding circuitry responsive to one or more integrity confirmation instructions to trigger the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task. The integrity confirmation information is suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

[0119] In the present application, the words “configured to. . .” are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation.

[0120] In the present application, lists of features preceded with the phrase “at least one of’ mean that any one or more of those features can be provided either individually or in combination. For example, “at least one of [A], [B] and [C]” encompasses any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), A and B in combination (without C), A and C in combination (without B), B and C in combination (without A), or A, B and C in combination. Although illustrative configurations of the invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise configurations, and that various changes, additions and modifications can be effected therein by one skilled in the art without departing from the scope of the invention as defined by the appended claims. For example, various combinations of the features of the dependent claims could be made with the features of the independent claims without departing from the scope of the present invention.

Claims

WE CLAIM:

1. An apparatus comprising: a data processing pipeline configured to perform data processing; and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline, the extension processing circuitry configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline, wherein the data processing pipeline comprises decoding circuitry responsive to one or more integrity confirmation instructions to trigger the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task, the integrity confirmation information suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

2. The apparatus of claim 1, wherein the one or more further instances comprises an instance of the computational task performed by the processing pipeline.

3. The apparatus of claim 1 or claim 2, wherein the one or more further instances of processing operations comprises a previous instance of the computational task performed by the extension processing circuitry.

4. The apparatus of any preceding claim, wherein: the one or more integrity confirmation instructions comprise an extension start integrity confirmation instruction specifying the computational task; and the decoding circuitry is responsive to the extension start integrity confirmation instruction, as one of the one or more operations, to clear the integrity confirmation information, and to generate the delegation signal to trigger the extension processing circuitry to perform the computational task.

5. The apparatus of claim 4, wherein the decoding circuitry is responsive to the extension start integrity confirmation instruction to trigger the data processing pipeline to perform a further instance of the computational task.

6. The apparatus of claim 5, wherein the extension start integrity confirmation instruction triggers a branch to instructions comprised in the further instance.

7. The apparatus of any of claims 5 to 6, wherein the extension processing circuitry is responsive to the delegation signal to perform plural instances of the computational task.

8. The apparatus of claim 7, wherein the extension processing circuitry is configured to store the integrity confirmation information for each of the plural instances.

9. The apparatus of claim 7 or claim 8, wherein the extension processing circuitry is responsive to the delegation signal to implement a variable length delay between each of the plural instances.

10. The apparatus of any of claims 7 to 9, wherein the extension processing circuitry is configured to perform the plural instances until the data processing pipeline has completed the further instance computational task.

11. The apparatus of any preceding claim, wherein: the apparatus comprises a plurality of extension processing circuits, the plurality of extension processing circuits comprising the extension processing circuitry; and the decoding circuitry is responsive to the extension start integrity confirmation instruction, to clear the integrity confirmation information for two or more extension processing circuits of the plurality of extension processing circuits, and to generate the delegation signal to trigger the two or more extension processing circuits to each perform the computational task.

12. The apparatus of claim 11, wherein the decoding circuitry is configured to implement a delay between triggering each of the two or more extension processing circuits.

13. The apparatus of claim 11 or claim 12, wherein the decoding circuitry is responsive to a compare extension processing instruction to generate control signals to cause the data processing pipeline to: trigger each extension processing circuit of the two or more extension processing circuits to return the integrity confirmation information maintained by that extension processing circuit to the data processing pipeline; and compare the integrity confirmation information returned by each extension processing circuit, and in response to a determination that the integrity confirmation information returned by at least one of the two or more extension processing circuits differs from another of the two or more extension processing circuits, to signal a fault.

14. The apparatus of any of claims 11 to 13, wherein: the decoding circuitry is responsive to a compatibility check instruction to generate compatibility check control signals; and the data processing pipeline is responsive to the compatibility check control signals to trigger at least one of the plurality of extension processing circuits to indicate whether the at least one of the plurality of extension processing circuits is capable of maintaining the integrity confirmation information.

15. The apparatus of any of claims 11 to 14, wherein: the data processing pipeline is configured to store selection information indicative of whether an integrity confirmation state is in an enabled state or a disabled state for each of the plurality of extension processing circuits; the two or more extension processing circuits comprises each of the plurality of extension processing circuits indicated as having integrity confirmation enabled; and the decoding circuitry is responsive to an integrity confirmation control instruction specifying at least one of the plurality of extension processing circuits and a new integrity confirmation state, to update the integrity confirmation state for that one of the plurality of extension processing circuits to the new integrity confirmation state.

16. The apparatus of any preceding claim, wherein:the one or more integrity confirmation instructions comprise a return integrity information instruction; and the decoding circuitry is responsive to the return integrity information instruction to trigger the extension processing circuitry, as one of the one or more operations, to return a current value of the integrity confirmation information relating to a previously delegated task to the data processing pipeline.

17. The apparatus of claim 16, wherein the extension processing circuitry is responsive to returning the current value, to disable storing of the integrity confirmation information.

18. The apparatus of any preceding claim, wherein the data processing pipeline is responsive to a determination that the integrity confirmation information differs from the confirmation data, to generate a fault signal.

19. The apparatus of claim 18, wherein generating the fault signal comprises at least one of raising an exception to interrupt a current process; triggering the current process to branch to a fault handling code sequence; and / or triggering an indication of the fault signal to be stored.

20. The apparatus of any preceding claim, wherein the extension processing circuitry comprises an accumulation register to store the integrity confirmation information, and the extension processing circuitry is configured to accumulate information indicative of result data generated during the computational task into the accumulation register.

21. The apparatus of any preceding claim, wherein the extension processing circuitry comprises a further data processing pipeline, the further data processing pipeline having a reduced processing capability relative to the data processing pipeline.

22. The apparatus of any preceding claim, wherein the extension processing circuitry and the data processing pipeline are each implemented using a different hardware configuration.

23. The apparatus of any preceding claim, wherein the extension processing circuitry is configured to perform the computational task using a first sequence of processing operations different to a second sequence of processing operations required to perform the computational task using the data processing pipeline.

24. A method of operating an apparatus comprising a data processing pipeline configured to perform data processing, and extension processing circuitry associated with the data processing pipeline and configured to perform a computational task delegated by the data processing pipeline in response to a delegation signal received from the data processing pipeline, the extension processing circuitry configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline, the method comprising: in response to one or more integrity confirmation instructions, triggering the extension processing circuitry to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task, the integrity confirmation information suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

25. A computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: data processing pipeline program logic configured to perform data processing; and extension processing program logic associated with the data processing pipeline program logic and configured to perform a computational task delegated by the data processing pipeline program logic in response to a delegation signal received from the data processing pipeline program logic, the extension processing circuitry configured to perform the computational task asynchronously to the data processing operations performed by the data processing pipeline program logic, wherein: the data processing pipeline program logic comprises decoding program logic responsive to one or more integrity confirmation instructions to trigger the extension processing program logic to perform one or more operations to validate integrity confirmation information obtained by the extension processing circuitry when performing the computational task, the integrity confirmationinformation suitable for comparison against confirmation data obtained from one or more further instances of the computational task.

Citation Information

Patent Citations

  • Debug and trace circuit in lockstep architectures, associated method, processing system, and apparatus

    EP4339783A1

  • Integrity checking

    GB2620134A

  • Fault detection system

    WO2021093931A1