Vector processing

The system addresses the limitation of hardware-dependent vector processing by switching between processing modes and registers, enhancing flexibility and efficiency across different hardware configurations.

JP7857908B2Active Publication Date: 2026-05-13ARM LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ARM LTD
Filing Date
2021-07-08
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing vector processing technologies require knowledge of the available hardware vector lengths for decoding and execution, limiting flexibility and compatibility across different hardware configurations.

Method used

A system that allows switching between processing modes using different sets of registers and instruction sets, enabling execution of instructions that are unavailable in one set by a second processing circuit, such as a coprocessor, while maintaining vector-length independence.

Benefits of technology

Enables flexible and efficient vector processing across varying hardware configurations by allowing code to utilize available vector lengths in both the main CPU and coprocessor, enhancing processing capabilities and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007857908000001
    Figure 0007857908000001
  • Figure 0007857908000002
    Figure 0007857908000002
  • Figure 0007857908000003
    Figure 0007857908000003
Patent Text Reader

Abstract

The apparatus comprises an instruction decoder that decodes processing instructions; one or more first registers; a first processing circuit that executes the decoded processing instructions in a first processing mode in which the first processing circuit is configured to execute the decoded processing instructions using the one or more first registers; and a control circuit that selectively initiates execution of the decoded processing instructions in a second processing mode in which the decoded processing instructions are selectively executed using one or more second registers, wherein the instruction decoder is configured to decode processing instructions selected from a first instruction set in the first processing mode and processing instructions selected from a second instruction set in the second processing mode, one or both of the first and second instruction sets including at least one instruction that is unavailable in the other of the first and second instruction sets, and the instruction decoder is configured to decode one or more mode change instructions for changing between the first processing mode and the second processing mode, and the first processing circuit is configured to change a current processing mode between the first processing mode and the second processing mode in response to execution of the mode change instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to vector processing.

[0002] Some data processing configurations enable vector processing operations, which include applying a single vector processing instruction to data items of a data vector having a plurality of data items at respective positions within the data vector. In contrast, scalar processing effectively operates on a single data item rather than a data vector.

[0003] Vector processing can be useful when processing operations are to be performed on many different instances of data being processed. In a vector processing configuration, a single instruction can be applied simultaneously to a plurality of data items (of a data vector). This can improve the efficiency and throughput of data processing compared to scalar processing.

[0004] It has been proposed to provide a vector processing instruction set that is “agnostic” with respect to the physical vector length provided by the hardware on which code including instructions from the instruction set is executed. An example is the instruction set defined by the so-called “Scalable Vector Extension” (SVE) or the SVE2 architecture made by Arm Ltd. However, the decoding and / or execution of at least some instructions of such an instruction set requires knowledge of the available vector lengths that are compatible with what is provided by the actual hardware.

Summary of the Invention

[0005] In an exemplary configuration, an apparatus an instruction decoder that decodes processing instructions, one or more first registers, a first processing circuit configured to execute the decoded processing instructions in a first processing mode in which the first processing circuit executes the decoded processing instructions using the one or more first registers A control circuit selectively initiates the execution of a decoded processing instruction in a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers. The instruction decoder is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for switching between a first processing mode and a second processing mode. The device is provided in which a first processing circuit is configured to change the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0006] In another exemplary configuration, the method is, Decoding processing instructions, In a first processing mode in which the processing circuit is configured to execute a decoded processing instruction using one or more first registers, the execution of the decoded processing instruction and In a second processing mode in which a decoded processing instruction is selectively executed using one or more second registers, the execution of the decoded processing instruction is selectively initiated, Decoding involves decoding a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. Decoding involves decoding one or more mode change instructions for switching between a first processing mode and a second processing mode. The method provided includes changing the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0007] In another exemplary configuration, a computer program for controlling a host data processing device to provide an instruction execution environment, An instruction decoder that decodes processing instructions, One or more first registers, In a first processing mode in which a first processing circuit is configured to execute a decoded processing instruction using one or more first registers, the first processing circuit that executes the decoded processing instruction, A control circuit selectively initiates the execution of a decoded processing instruction in a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers. The instruction decoder is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for switching between a first processing mode and a second processing mode. A computer program is provided, which configures a first processing circuit to change the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0008] Further aspects and features of this technology are defined by the attached claims. [Brief explanation of the drawing]

[0009] The present technology will be further described, merely as an example, with reference to the embodiments shown in the attached drawings. [Figure 1] A schematic diagram of the data processing device is shown. [Figure 2] Schematically shows the use of a device having a dedicated memory management unit. [Figure 3] Schematically shows a data processing device. [Figure 4] Schematically shows a storage device for vector length registers. [Figure 5] Schematically shows a storage device for vector length registers. [Figure 6] It is a schematic flowchart showing a method. [Figure 7] Schematically shows a data processing device. [Figure 8] They are schematic flowcharts showing respective methods. [Figure 9] They are schematic flowcharts showing respective methods. [Figure 10] They are schematic flowcharts showing respective methods. [Figure 11] Schematically shows a composite device. [Figure 12] Schematically shows a simulator implementation form.

Mode for Carrying Out the Invention

[0010] Before discussing the embodiments with reference to the accompanying drawings, the following embodiments will be described.

[0011] An exemplary embodiment is a device comprising an instruction decoder for decoding processing instructions, one or more first registers, a first processing circuit configured to execute the decoded processing instructions in a first processing mode in which the first processing circuit executes the decoded processing instructions using one or more first registers, and a control circuit configured to selectively start the execution of the decoded processing instructions in a second processing mode in which the decoded processing instructions are selectively executed using one or more second registers. The instruction decoder is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, and one or both of the first and second instruction sets include at least one instruction that is not available in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for changing between the first processing mode and the second processing mode. There is provided an apparatus, wherein a first processing circuit is configured to change a current processing mode between the first processing mode and the second processing mode in response to execution of a mode change instruction.

[0012] Embodiments of the present disclosure can provide a mode change caused by execution of a processing instruction between a first mode and a second mode each having a respective instruction set that is at least not completely overlapping, using respective first and second sets of registers (note that the first register and the second register may be different registers having potentially different sizes and / or capabilities, for example, some operating as matrix registers and / or at least some of the first register and the second register being vector registers having potentially different physical vector lengths).

[0013] The instruction decoder is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction unavailable in the other of the first and second instruction sets. This advantageously allows (a) an additional function to be provided in one of the processing circuits, for example, in a second processing circuit (e.g., a coprocessor) that is unavailable in the other of the processing circuits, and (b) the need to provide some basic functions in both processing circuits when the function can be adequately provided by only one of the processing circuits (e.g., some scalar processing functions are provided only in the first processing circuit). As a further example, the second processing circuit may include a matrix processing circuit, and the second instruction set includes one or more matrix processing instructions unavailable in the first instruction set.

[0014] In some embodiments, the processing instruction includes a vector processing instruction, and one or more first registers comprise one or more first vector registers having a first vector length, and the first processing circuit is configured to execute the decoded processing instruction in a first processing mode using one or more first vector registers according to a vector length less than or equal to the first vector length, and to selectively execute the decoded processing instruction in a second processing mode using one or more second vector registers having a second vector length according to a vector length less than or equal to the second vector length.

[0015] Such embodiments of the present disclosure can provide further operating modes within the context of vector processing such that a change between a first processing mode and a second processing mode can be initiated by the execution of a mode change instruction, and such a change has the effect of potentially altering the routing of the decoded processing instruction (e.g., to a second processing circuit such as a coprocessor in the second processing mode) and potentially altering the vector length applicable to the execution of the instruction in the second processing mode.

[0016] This configuration is advantageous in that it allows for the use of a coprocessor that may have a different hardware vector length than that of the main CPU (for example, the one that executes all instructions in the first processing mode), while allowing code that may be vector-length independent to utilize the vector lengths available in the main CPU in the first processing mode and in the coprocessor in the second processing mode.

[0017] In some examples, the first processing circuit can operate to execute processing instructions at the dominant exception level in a hierarchy of exception levels, each having a different level of privilege for accessing the device's processing resources. For example, the first processing circuit may be configured to detect the vector length for executing an instruction in a first processing mode and the vector length for executing an instruction in a second processing mode by executing a vector length detection processing instruction.

[0018] During operation, the first processing circuit may be configured to execute a vector length setting instruction at any exception level other than the lowest exception level to set the maximum vector length applicable to that exception level and at least one lower exception level. In other words, code executed at the lowest exception level cannot set its own maximum vector length in normal operation. For example, the first processing circuit may be configured to execute a vector length setting instruction to set the vector length used in the first processing mode to be less than or equal to the first vector length, and the vector length used in the second processing mode to be less than or equal to the second vector length. In some examples, the first processing circuit is configured to execute a vector length setting instruction at a given exception level to set the vector length for use in a given processing mode to be less than or equal to the maximum vector length of a given processing mode set by executing a vector length setting instruction at an exception level higher than the given exception level. Importantly, however, in some embodiments, the first processing circuit is configured to execute a mode change instruction at any exception level in the exception level hierarchy. The combination of these features allows code that runs at the lowest exception level, such as the so-called EL0 level in some practical examples, to initiate the modification processing mode even if it itself cannot set the maximum vector length in either the first or second processing mode.

[0019] The disclosure may be embodied as a CPU and coprocessor configuration, where the CPU provides a first processing circuit and the coprocessor provides a second processing circuit. In other examples, a single CPU may provide the first and second processing circuits and an associated set of vector registers. In any of these circumstances, the disclosure also includes an apparatus that provides one or more second vector registers having a second vector length and a second processing circuit configured to execute vector processing instructions decoded according to a vector length less than or equal to the second vector length.

[0020] In the second processing mode, there are no restrictions on the first processing circuit executing an instruction. For example, a scalar instruction may be best executed by the first processing circuit even when the device is operating in the second processing mode. To accurately handle the routing of such instructions, in an exemplary embodiment, the control circuit is configured to detect in the second processing mode whether a given processing instruction can be executed by the first or second processing circuit.

[0021] This disclosure is particularly useful for operations in which two or more connected CPUs share a common coprocessor by providing a sequence of instructions for coprocessor processing which are then queued for execution by the coprocessor (such as so-called streaming execution). In such a configuration, the device may comprise two or more instances of a first processing circuit, the instruction decoder of each instance of the first processing circuit being configured to selectively route second vector processing instructions to a common second processing circuit.

[0022] As mentioned above, in some examples, the vector processing instructions themselves do not need to depend on the vector length defined by the current processing mode, for example, using the SVE / SVE2 techniques described above.

[0023] In some exemplary configurations, the device includes a vector length register memory configured to store vector lengths less than or equal to a first vector length for use in a first processing mode, and vector lengths less than or equal to a second vector length for use in a second processing mode, and the instruction decoder is configured to access the vector length memory depending on the current processing mode.

[0024] In an exemplary configuration in which memory address translation is used, the device may include a data memory and an address translation circuit for translating between input memory addresses in an input memory address space and output memory addresses in an output memory address space, wherein the data memory is accessible according to the output memory address space. Advantageously, the control circuit may be configured, in a second processing mode, to initiate the translation of input memory addresses associated with processing instructions for execution by a second processing circuit and to provide the translated output memory addresses to the second processing circuit. This avoids the need for memory address translation to be performed by the second processing circuit.

[0025] There is no fundamental requirement that the first and second processing circuits have different native or inherent vector lengths, but in some examples, the second vector length is different from the first vector length. For example, the second vector length may be greater than the first vector length.

[0026] Another exemplary embodiment is a method, Decoding processing instructions, In a first processing mode in which the processing circuit is configured to execute a decoded processing instruction using one or more first registers, the execution of the decoded processing instruction and In a second processing mode in which a decoded processing instruction is selectively executed using one or more second registers, the execution of the decoded processing instruction is selectively initiated, Decoding involves decoding a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. Decoding involves decoding one or more mode change instructions for switching between a first processing mode and a second processing mode. The method provided includes changing the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0027] Another exemplary embodiment is a computer program for controlling a host data processing device to provide an instruction execution environment, An instruction decoder that decodes processing instructions, One or more first registers, In a first processing mode in which a first processing circuit is configured to execute a decoded processing instruction using one or more first registers, the first processing circuit that executes the decoded processing instruction, A control circuit selectively initiates the execution of a decoded processing instruction in a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers. The instruction decoder is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for switching between a first processing mode and a second processing mode. A computer program is provided, which configures a first processing circuit to change the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0028] Exemplary device Referring to the drawing, Figure 1 schematically shows a data processing device 100 comprising one or more central processing units (CPUs) or processing elements 110, 120, 130, a memory management unit (MMU) 140, and a coprocessor 150, all of which are connected to an interconnection 160 to which the main memory subsystem 170 is also connected.

[0029] During operation, each CPU can execute program instructions, which may be loaded, for example, from the memory subsystem 170 and stored in one or more local caches of their choice. The CPU can also refer to instructions for execution by the coprocessor 150. Techniques for selectively providing local execution (in the CPU) and remote execution (in the coprocessor) are described below.

[0030] Therefore, using the techniques described below, Figure 1 provides examples of two or more instances 110, 120, 130 of the first processing circuit, where the instruction decoder (described below) of each instance of the first processing circuit is configured to selectively decode processing instructions, and some of the processing instructions can be routed to a common second processing circuit 150 for execution (using the techniques described below).

[0031] Memory address translation The function of the MMU140 is to provide address translation between input memory addresses in the input memory address space (such as virtual addresses VA in the VA space) and output memory addresses in the output memory address space (such as physical addresses PA in the PA space). The reason for providing memory address translation is that different threads or processes running on the CPU can be isolated so that each operates according to its respective VA space, where virtual addresses do not give direct access to data physically stored by the memory subsystem 170. A higher-level control process, such as an operating system or hypervisor, has control over address translation and can thus provide security and isolation between different CPUs, threads, or processes, preventing one CPU, thread, or process from accessing data that is the property (reserved access) of another CPU, thread, or process.

[0032] The use of a separate MMU 140, as shown in Figure 1, is only one option. As schematically shown in Figure 2, the CPU 200 may have its own MMU function 210. Similarly, the coprocessor 220 may have its own MMU function 230. In other examples, when the CPU initiates the execution of instructions by the coprocessor, it may itself obtain any necessary address translation and provide the coprocessor with the translated output memory address (such as a physical memory address).

[0033] The above example has simply referred to a single-stage memory address translation process. However, it is possible to use multiple stages of memory address translation, such as translation from VA to Intermediate Physical Address (IPA), and then from there to PA. One reason why this multi-stage address translation can be used in some examples is that it allows the first translation stage to be controlled, for example, by the operating system, and the second stage to be controlled, for example, by a process that potentially oversees multiple operating systems, such as a hypervisor.

[0034] This technique is applicable to any of these memory address translation scenarios.

[0035] CPU example Figure 3 is a schematic diagram showing one example, CPU 300, from among CPUs 110 to 130.

[0036] The CPU is capable of both scalar and vector processing.

[0037] The general distinction between scalar and vector processing is as follows: Vector processing involves applying a single vector processing operation to multiple data items in a data vector, where each data item has its own position within the data vector. In contrast, scalar processing operates on (effectively) a single data item, not a vector.

[0038] Vector processing can be useful in situations where the same or similar processing is performed on many different instances of a data item. In a vector processing configuration, a single processing instruction can be applied to multiple data items of a data vector, and this may effectively occur simultaneously (although whether to implement concurrent or sequential execution is often left to the hardware designer). This can improve the efficiency and / or throughput of the processing configuration compared to scalar processing.

[0039] Referring to Figure 3, the Level 2 cache 305 interfaces with the memory subsystem 170 via the interconnect 160 (307). The Level 1 instruction cache 310 provides a more localized cache of processing instructions, and the Level 1 data cache 315 provides a more localized cache of data retrieved from or stored in the memory subsystem.

[0040] The fetch circuit 320 fetches program instructions from the memory subsystem 170 via various caches as shown, and provides the fetched program instructions to the decoder circuit 325. The decoder circuit 325 decodes the fetched program instructions and generates control signals to control the execution unit 330 to perform processing operations in at least a first operating mode.

[0041] Next, the first operating mode will be described.

[0042] Generally speaking, the apparatus in Figure 3 can operate in a first processing mode in which the execution unit 330 is configured to execute decoded processing instructions provided by the decoder circuit 325 (as an example of a first processing circuit).

[0043] In the second processing mode, further (second) processing circuits, not shown in Figure 3, are used for at least part of the execution of the decoded processing instructions. For example, the second processing circuit may be provided in the coprocessor 150, which is described below.

[0044] First processing mode In the first processing mode, the decoder circuit 325 provides the decoded processing instructions of the first instruction set to the issue circuit 335, which controls the issuance of the decoded instructions to the execution unit 330, which comprises a vector processor 332, a scalar processor 334, and a load / store circuit 336. Other users or circuits may be provided as part of the functionality of the execution unit 330, but it will be understood that these are not shown here for clarity in the diagram.

[0045] The vector processor 332 executes the decoded vector processing instructions using one or more vector registers 340 having a first vector length VL1. The scalar processor 334 executes the decoded scalar processing instructions using one or more scalar registers 345. In either case, the vector processor 330 or the scalar processor 334 can receive values ​​from their respective registers and write them back to the registers via the write-back circuit 350. When it is necessary to retrieve data from or store data in the memory subsystem 170, this is done by the load / store circuit 336 via the depicted cache configuration. Previously proposed coherency techniques can be applied, for example, by a coherency controller (not shown) implemented as part of the interconnect 160.

[0046] A set of one or more "ZCR" registers 355 is provided as part of a set of scalar registers 345. Using the technique described below, the ZCR registers provide the decoder 325 with a dominant vector length 360 for use when decoding the fetched instruction. The fetched instruction itself is "vector length independent," conforming to the so-called "Scalable Vector Extensions" (SVE) or SVE2 architecture as defined, for example, by Arm Ltd. That is, the program code itself can run on hardware with vector lengths selectable from at least a set of supported vector lengths, so that the same program code can run on different hardware with different examples of such vector lengths. However, the actual vector length available to the hardware on which the decoded instruction is executed is necessary to decode that instruction into appropriate control signals to control its execution.

[0047] Second processing mode In the second processing mode, at least some of the operations required to execute the decoded instructions of the second instruction set (which does not exactly match the first instruction set, i.e., at least one of the first and second instruction sets includes instructions that are not available in the other instruction set) are performed by a second processing circuit, such as the coprocessor 150.

[0048] The inability of a particular instruction to be fully executed by the first processing circuit is not a requirement of the second processing mode. The second processing mode allows instructions to be appropriately routed to either the first or second processing circuit. For example, a scalar instruction can be fully executed in a connected CPU despite the second processing mode. However, instructions can only be routed to a second processing circuit, such as a coprocessor, in the second processing mode.

[0049] In this configuration, the decoder circuit 325 receives a vector length from an alternative register or register field (circularly shown as ZCR' in Figure 3, which indicates the appropriate vector length for the second processing circuit). Note that this may (but does not have to) differ from the vector length associated with the first processing circuit and execution unit 330, and in some examples may be greater than the vector length of the first processing circuit. In such a configuration, the hardware vector lengths VL1 and VL2 associated with the first and second processing circuits may be different, and / or the maximum vector lengths allowed by their respective ZCR registers may also be different. The decoder circuit 325 decodes the fetched instructions in the second processing mode, but instead of simply routing them to the execution unit 330, the issue circuit 335 interfaces with a coprocessor interface 365 (370) that communicates with the coprocessor 150 via the interconnect 160.

[0050] It should be noted that the use of the second processing mode does not completely eliminate some contributing operations performed by the CPU300 itself, even when instruction execution refers to the coprocessor.

[0051] For example, the CPU 300 can interface with the memory management unit 140, or (if provided) its own MMU 210, to obtain any memory address translations necessary for the execution of instructions by the coprocessor, and these can be included in the communication with the coprocessor.

[0052] In another example, the coprocessor may generate results such as scalar values ​​or conditional codes that are returned to the CPU 300 via the coprocessor interface 365 for writing to scalar registers, conditional flags, etc.

[0053] In another example, CPU300 can perform several actions to prepare one or more operands for the coprocessor to act upon.

[0054] In this regard, the processing of any necessary operations in the CPU 300 and the transmission of messages defining the requested instructions to the coprocessor are handled by the issuing circuit 335 and the coprocessor interface 365, and together they provide the function of the control circuit 375 for selectively initiating the execution of a decoded processing instruction in a second processing mode in which the decoded processing instruction can be executed using (whole or at least partially) a second processing circuit having one or more second vector registers according to a second vector length, for example, a vector length defined by the register field ZCR'. The operation of the circuit 375 can be performed in response to information provided by the decoder circuit and / or in response to the current state of the SM bits or fields described below.

[0055] The ZCR and ZCR' registers or fields provide an example of a vector length register storage device configured to store vector lengths less than or equal to a first vector length for use in a first processing mode, and vector lengths less than or equal to a second vector length for use in a second processing mode, and the instruction decoder is configured to access the vector length storage device according to the current processing mode.

[0056] Therefore, in summary, Figure 3 is, An instruction decoder 325 that decodes processing instructions that potentially contain vector processing instructions, One or more first registers 340, such as a vector register 340 having a first vector length VL1, For example, in a first processing mode in which a first processing circuit is configured to execute a decoded processing instruction using one or more first registers according to a vector length less than or equal to a first vector length (if they are vector registers), the first processing circuit 330 that executes the decoded processing instruction, The system includes a control circuit 375 that selectively starts the execution of a decoded processing instruction in a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers, such as a second vector register, according to a vector length less than or equal to a second vector length, The instruction decoder 325 is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder 325 is configured to decode one or more mode change instructions (Mstart, Mstop, described later) for switching between a first processing mode and a second processing mode. The present invention provides an exemplary apparatus in which the first processing circuit is configured to change the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0057] It should be noted that the underlying techniques for providing different functionalities in the first and second modes, such as potentially different registers and instruction sets that are different or at least not entirely overlapping, apply regardless of whether the processing instructions are vector instructions and whether the registers are vector registers, matrix registers, or scalar registers.

[0058] Exception level In this example, the CPU 300 can operate to execute processing instructions and the dominant exception level in the exception level hierarchy, each having a level of privilege for accessing processing resources associated with the CPU 300.

[0059] In one example, four exception levels EL0 to EL3 may be provided, where the privileges for accessing processing resources increase with increasing numerical values ​​n in the ELn designation, with EL3 having higher privileges than EL2, and so on. The current level of privileges can only be changed when the CPU300 takes an exception or returns from an exception. In an exemplary configuration, the application code may run at EL0, the operating system at EL1, the hypervisor at EL2, and the secure firmware at EL3. However, this is merely an exemplary example. The term “privilege” can refer to privileges for accessing the memory system, and in a given example, may include privileges for modifying translation information used by the MMU140, and / or privileges for accessing other processing resources such as processing registers.

[0060] Therefore, this configuration provides an example in which the first processing circuit can operate to execute processing instructions at the dominant exception level in a hierarchy of exception levels, each having a level of privilege for accessing the processing resources of the device.

[0061] Hardware-supported vector length In principle, it may be possible to implement a system that allows for any physical (hardware) vector length, but in this embodiment, the physical vector length is constrained to a multiple of 128 bits up to a maximum of 2048 bits. In some examples, the hardware is a multiple of 128 bits related as a power of 2, i.e., 128 × 2 n It is required solely to support this. However, at a general level, it can be assumed that in at least some embodiments, there is a set of acceptable vector lengths, and also a maximum vector length that the hardware can support. A particular example of hardware may, of course, have a single associated physical vector length (e.g., 256 bits).

[0062] ZCR Register Figures 4 and 5 provide schematic examples of the ZCR register 355 in Figure 3. Generally speaking, a ZCR register is provided for each of EL1, EL2, and EL3. ZCR_EL3 is a control register for controlling the accessible vector lengths in EL3, EL2, EL1, and EL0. ZCR_EL2 is a control register for controlling the accessible vector lengths in EL2 and so-called non-secure EL1 and EL0. ZCR_EL1 is a control register for controlling the accessible vector lengths in EL1 and EL0. In all cases, the second part of the register name indicates the level of privilege or exception level required to modify the contents of the register, for example, ZCR_EL3 requests a change in exception level EL3.

[0063] These ZCR registers are hierarchical; therefore, ZCR_EL2 cannot select a vector length longer than that defined by ZCR_EL3, and similarly, ZCR_EL1 cannot select a vector length longer than that defined by ZCR_EL2 if EL2 is implemented and enabled with the same security state as EL1 (which may be controlled by SCR_EL3.EEL2 control). Also note that EL2 and EL3 are actually optional for a given implementation, and therefore the EL2 / EL3 constraints only apply if their exception levels are implemented.

[0064] In other words, the first processing circuit is configured to execute a vector length setting instruction at a given exception level to set the vector length for use in a given processing mode to be less than or equal to the maximum vector length of a given processing mode set by executing a vector length setting instruction at an exception level higher than the given exception level.

[0065] In the example in Figure 4, two sets of ZCR registers are provided, one set 400 applicable to the first processing mode and the other set 410 applicable to the second processing mode, and the appropriate control circuit 420 routes the required values ​​to the appropriate registers. In the example in Figure 5, each register 500 has two fields 510, 520 that store the relevant values ​​for the first or second processing mode.

[0066] The ZCR register can be populated with hardware-supported vector lengths using the technique described with reference to Figure 6.

[0067] In the type of SVE / SVE2 processor shown in Figure 3, the instruction set itself is vector length independent, but the ZCR register is popularized by querying the hardware to find the actual vector length implemented by the physical circuitry on which the instruction is executed. As mentioned above, the vector length defined by the ZCR register applicable to the current exception level is used by the decoding process to decode the fetched program instruction.

[0068] The process in Figure 6 may also be followed sequentially at each exception level, starting from the highest exception level EL3 down to EL1. Note that while instructions executed at EL0 cannot modify any of the ZCR registers themselves, they can detect the effect of at least the ZCR_EL1 register on the available vector length.

[0069] In step 600, a program instruction is executed, writing a value to the associated ZCR register (ZCR_EL3, assuming the process is initially executed on EL3). This could be, for example, the maximum vector length mentioned above (e.g., 2048 bits). Next, in step 610, the hardware is queryed, for example, by executing an RDVL (Read Vector Length) instruction to detect the effect of the same ZCR register on the available vector length. This prompts the hardware to write to the instruction's destination register a vector length that it can support and that is as close as possible to, but not greater than, the vector length written to the ZCR_EL3 register. In this way, the program code can determine the result of the process in step 620 with respect to having found at least one acceptable vector length that can be used. The process can be executed iteratively at a particular exception level to populate one or more sets of vector lengths that are currently allowed on the hardware. This can be done, for example, at device startup.

[0070] Once a set of one or more hardware-supported vector lengths is derived in this way, the register ZCR_EL3 can be populated with one of the allowed vector lengths.

[0071] The list of hardware-supported vector lengths is a fundamental characteristic of the design of the first and second vector processors and their vector register files. The coprocessor may communicate its set of supported vector lengths to the CPU, or it may be fixed and known to both the CPU and the coprocessor.

[0072] The RDVL instruction does not change the ZCR value itself, but merely returns the vector length (or the highest hardware-supported value compatible with it) selected by the ZCR_ELn register. The ZCR is written by a separate MSR (Move to System Register) instruction. In practice, when ZCR_ELn is written, or when the exception level changes due to an exception or exception return that applies a different set of ZCR registers, a new vector length (VL) may be calculated, used by the CPU decoder, and stored in the internal CPU memory to which the RDVL returns.

[0073] Therefore, the underlying processing of the RDVL instruction may include, for example, the following: The hardware has information available (for example, as a hardcoded storage device 342, or as part of the functionality of the vector register itself) indicating one or more hardware-supported vector lengths on which the hardware may operate. The execution of the RDVL instruction prompts the execution unit 330 to detect the hardware vector length from this source, take into account the value held by the associated ZCR register, set the VL for use now to the highest possible value while being compatible with that information and the value previously written to the ZCR register, and return the set value.

[0074] Once all of this is done, a similar process can be performed on ZCR_EL2, resulting in a population of that register with an acceptable vector length no larger than that held by ZCR_EL3. Again, the process can be performed on ZCR_EL1.

[0075] In this way, during subsequent normal operation following the initialization process described above, the ZCR_ELn register or register field contains its respective hardware-supported vector length.

[0076] An alternative field or register ZCR_ELn' may be populated by a similar technique, but operating in a second processing mode in which the hardware response may come from a second processing circuit, such as a coprocessor, rather than from the first processing circuit (execution unit 330), and reflects the vector length of the vector register described below, and other hardware capabilities associated with the second processing circuit. The hardware capabilities of the second processing circuit may be stored in the first processing circuit (similar to 342, or as part thereof), and the first and second processing circuits may be implemented in common as, for example, Figure 11 described below, or as a common system-on-chip (SoC) or network-on-chip (NoC), or, in the case of a discrete coprocessor, by querying similar hardware storage in the coprocessor.

[0077] Therefore, Figure 6 provides an example in which the first processing circuit is configured to detect the vector length for executing the instruction in the first processing mode and the vector length for executing the instruction in the second processing mode by executing a vector length detection processing instruction.

[0078] For example, the setting of the ZCR register as described above at the end of the process in Figure 6 provides an example in which the first processing circuit is configured to execute a vector length setting instruction at an exception level other than the lowest exception level to set the maximum vector length applicable to that exception level and at least one lower exception level. This is also an example in which the first processing circuit is configured to execute a vector length setting instruction to set the vector length used in the first processing mode to a first vector length or less, and the vector length used in the second processing mode to a second vector length or less.

[0079] Examples of coprocessors Figure 7 schematically shows an example coprocessor 700, such as coprocessor 160. The coprocessor communicates with one or more connected CPUs 110-130 via interconnects (705) and also provides a level 2 cache 710 and a level 1 cache 715 associated with a load / store circuit 720 to provide communication with the memory subsystem. Note that in some embodiments, depending on the overall system design, some of the caches, such as the level 2 cache, may be shared between the CPU (if only one is provided) and the coprocessor.

[0080] The second processing circuit 725, in this example, comprises a vector processor 730, a matrix processor 735, and a load / store circuit 720. This is associated with matrix registers 740 and vector registers 745, which have a physical vector length VL2 that may be the same as or different from VL1 of the first processing circuit. In some examples, VL2 may be greater than VL1.

[0081] Generally speaking, each set of registers 740 and 745, 750 in total, is provided to the coprocessor 700 for each connected CPU, which allows interleaved instructions to be executed on behalf of different CPUs without the need to write the register contents back to main memory (this is generally expected to be a slower process compared to the time it takes to execute instructions received from the connected CPUs).

[0082] The write-back circuit 760 controls the write-back of the results from the second processing circuit 725 to registers 740 and 745.

[0083] Examples of the first and second instruction sets The matrix processor 735 executes special matrix processing instructions that are unavailable for execution by the CPU execution unit 330 in this example. Similarly, some instructions that can be executed by the CPU are unavailable for execution by the second processing circuit 725, for example, collect-load and distributed-storage vector operations and / or some scalar operations that can be appropriately and efficiently executed in the connected CPU. This provides an example in which the instruction decoder is configured to decode processing instructions selected from a first instruction set in a first processing mode and processing instructions selected from a second instruction set in a second processing mode, and one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. In an example in which the second processing circuit includes a matrix processing circuit, the second instruction set may include one or more matrix processing instructions that are unavailable in the first instruction set.

[0084] Queuing and Streaming Communication from the connected CPU is processed by a queue manager 760, which receives information defining the instructions that need to be executed for a particular CPU and forms a queue of such instructions associated with the identification information of the connected CPU that requested the instruction. When an instruction reaches the appropriate stage in the queue, it is passed to an optional decoder circuit 765. In some examples, the entire decoding process takes place in the connected CPU, with the information to be communicated to the coprocessor already decoded, but in other examples, at least a partial stage of decoding may be required in the coprocessor, hence the term "optional." From there, a control signal is passed to an issue circuit 770, which controls the issuance of instructions to a second processing circuit 725.

[0085] In many exemplary instances, no further interaction is required beyond the handshake interaction to confirm the successful reception and queuing of instructions from the connected CPU. That is, the coprocessor can be instructed by the connected CPU to retrieve information from memory or registers, perform a specific operation on it, and store the result in registers or memory. No further information needs to be provided to the connected CPU that requested the instruction, as long as this instruction is completed before any subsequent instructions that may depend on it are committed. In this way, the operation of the connected CPU to provide instructions to the coprocessor can be called a so-called "streaming mode," in that it can communicate a series of instructions to the coprocessor, and the only interaction is the handshake indicating successful reception and enqueueing.

[0086] However, it is naturally possible to envision the execution of instructions by a coprocessor that produce results that need to be communicated to the connected CPU that requested the execution. Examples may include extracting a single scalar value from a vector or matrix held in the coprocessor and returned to the connected CPU that requested the instruction, or generating a conditional code that is returned to the connected CPU that requested the instruction. This type of inverse interaction can also be handled by a queue manager that can represent an interface with the CPUs connected via interconnection 160.

[0087] Either or both of the queue manager 760 and the issue circuit 770 can control which of the set of registers 750 are enabled by the second processing circuit 725 for a particular instruction execution, so that the appropriate set applicable to the connected CPU that requested the instruction execution is used.

[0088] Therefore, Figure 7 provides an example showing one or more second vector registers 745(750) having a second vector length, and a second processing circuit 725 configured to execute vector processing instructions decoded according to a vector length less than or equal to the second vector length.

[0089] Example of communication between a connected CPU and coprocessor The connected CPU can provide the coprocessor with at least some of the following information, for example, in the form of data processing transactions or packets addressed to the coprocessor: • Identification information for the connected CPU (to enable the selection of the correct set of registers in the coprocessor, to enable queue arbitration, and / or to provide a return address for handshake or response signals). • Coprocessor identification information (if there is any possible ambiguity). • Identification information for the instruction to be executed (which may be an opcode that is at least partially decoded by the coprocessor, or may be in the decoded form) • Identification information and / or values ​​for one or more operands (e.g., one or more PAs, one or more VAs or IPAs that the coprocessor needs to convert, identification information for one or more coprocessor registers) • The value of the selected scalar register required by the instruction.

[0090] As described above, the CPU can provide the translated PA to the coprocessor, thereby avoiding the need for the coprocessor itself to have MMU functionality or for an action to be initiated by a separate MMU. This provides an example in which the device comprises a data memory 170 and address translation circuits 140, 210 for translating between input memory addresses in the input memory address space and output memory addresses in the output memory address space, wherein the data memory is accessible according to the output memory address space. The control circuit 375 is configured to initiate the translation of input memory addresses associated with processing instructions for execution by a second processing circuit in a second processing mode and to provide the translated output memory addresses to the second processing circuit.

[0091] Switching between the first processing mode and the second processing mode. A pair of instructions are provided for the CPU to execute at any exception level (from EL0 upwards) in order to switch between the first and second processing modes. In other words, the first processing circuit is configured to execute mode change instructions at any exception level in the exception level hierarchy. These instructions are,

[0092] Mstart: This switches the operation to a second processing mode, i.e., the execution unit 330, in response to the Mstart instruction, does at least the following: To indicate a second processing mode, set a “Streaming Mode” (SM) control bit, register, or field (for example, in the decoder, or in the execution unit, or in a scalar register so that it can be referenced by various units of the CPU). • Choose to use the ZCR_ELn' variant for the ZCR register / field. • Enable the instruction set associated with the second processing circuit. The decoder circuit 325 is controlled to decode the instructions in the instruction set into appropriate control signals that are then transferred to the second processing circuit. • Control the control circuit 375 to selectively initiate the execution of the decoded instruction by the second processing circuit.

[0093] Mstop: This switches the operation to a first processing mode, i.e., the execution unit 330, in response to the Mstop instruction, does at least the following: • De-set the "Streaming Mode" (SM) control bit, register, or field to indicate the first processing mode. • Choose to use the ZCR_ELn variant for the ZCR register / field. • Enable the instruction set associated with the first processing circuit. The decoder circuit 325 is controlled to decode the instructions in the instruction set into appropriate control signals that are issued to the first processing circuit. • Control the control circuit 375 to issue the decoded instruction to the first processing circuit.

[0094] As mentioned above, these instructions are executable at any exception level. Therefore, in contrast to setting the ZCR register (which cannot be done at EL0), they provide techniques that are also available at EL0 to potentially change the dominant vector length at which the instructions are executed.

[0095] It should be noted that a change from one vector length to another is not necessarily a change between the physical vector lengths VL1 and VL2 of the first and second processing circuits, but rather a change between the dominant allowable vector lengths defined by the ZCR registers for the first and second processing circuits, which are appropriate for the dominant exception level.

[0096] It should be noted that the exception levels discussed here are the dominant exception levels for a particular connected CPU.

[0097] The use of these instructions is summarized by the schematic flowchart in Figure 8. Here, the execution of the Mstart instruction 800 results in various outcomes 810, including accessing the second processing mode register or register field ZCR_ELn', decoding the second set of processing mode instructions, and, where appropriate, routing the decoded instructions to the coprocessor. The execution of the Mstop instruction 820 results in outcome 830, including accessing the first processing mode register or register field ZCR_ELn, decoding the first set of processing mode instructions, and executing the decoded instructions in the first processing circuit. Although not shown in Figure 8, setting / unsetting the SM bits is also done as described above.

[0098] Example of operation in the second processing mode The illustrative flowchart in Figure 9 assumes that the starting point 900 in the second processing mode, i.e., the Mstart instruction, has already been executed. In the configuration of Figure 9, the operation by the CPU connected to the left of the dashed line 905 is schematically shown, and the operation by the coprocessor is depicted to the right of the dashed line 905.

[0099] In step 910, the connected CPU decoder circuit decodes the next fetched instruction based on the dominant vector length appropriate to the dominant exception level defined by the register or register field ZCR_ELn'. If necessary, in step 915, the CPU retrieves one or more operands associated with the decoded instruction.

[0100] Even in the second processing mode, some instructions, such as scalar instructions, may not require execution by the coprocessor. This relates to the terminology used to “selectively” initiate execution by the second processing circuit. Detection of whether this is necessary is performed in step 920 (for example, by comparing the current instruction with a predetermined set of instructions in this category), and if not, control moves to step 925, where the instruction is issued for local execution by the first processing circuit. Control moves to step 930, and the process is repeated for the next instruction. This provides an example in which the control circuit is configured to detect whether a given processing instruction is executable by either the first or second processing circuit in the second processing mode.

[0101] However, if at least part of the execution is required by a second processing circuit, control is transferred to step 935, where information defining the instruction is communicated to the coprocessor. In step 940, it is detected whether any results are required in the connected CPU. If the answer is "yes," the connected CPU halts in step 945 to wait for these results before passing control to step 930. If the answer is "no," control proceeds directly to step 930.

[0102] The test performed in step 940 concerns a simple mapping of instruction types. Some instruction types produce the expected results when returned to the connected CPU, while others (the so-called instructions that can operate in streaming mode as described above) do not.

[0103] Referring to the operation in the coprocessor, step 950 includes receiving communication from the connected CPU, queuing it, and acknowledging it.

[0104] In step 952, the next instruction is dequeued from the top of the queue and decoded (in step 955), or further decoded (as described above) if necessary. In step 960, the instruction is executed in the second processing circuit. If it does not need to be sent back to the connected CPU, the process can be terminated in the coprocessor, but if a result is needed, these are returned in step 965. Control then returns to step 952 to dequeue the next instruction in the queue.

[0105] Therefore, it should be noted that the loop 952-965 can operate independently of step 950, i.e., in the streaming operation mode as described above, except that, naturally, instances of step 950 occasionally require at least some population in the queue. Execution in the coprocessor by loops 952-965 continues to process each instruction from the queue as long as the queue contains instructions (if the queue is empty, step 952 simply pauses or stops functioning until an instruction enters the queue). Step 960 operates whenever the CPU communicates to the coprocessor an instruction to join the queue.

[0106] Method overview Figure 10 is a schematic flowchart illustrating a method that includes the following: (In step 1000) the process instruction is decoded, In a first processing mode (in step 1010), the processing circuit is configured to execute a decoded processing instruction using one or more first registers, and the decoded processing instruction is executed. (In step 1020) in a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers, the execution of the decoded processing instruction is selectively initiated, Decoding involves decoding a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. Decoding involves decoding one or more mode change instructions for switching between a first processing mode and a second processing mode. The method provided includes changing the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction.

[0107] Single CPU Embodiment It is not a requirement that the second processing circuit be implemented by a separate processing device such as the coprocessor mentioned above. In fact, it can be provided by a second processing circuit embodied within a single CPU, and thus this CPU provides a first and second processing circuit having potentially different associated vector lengths and vector registers of different sizes.

[0108] Figure 11 schematically shows a configuration in which the processing unit or CPU 1100 includes the circuit 1110 of Figure 3 connected to the circuit 1120 of Figure 7. In at least some examples, the type of interconnection described above is not required, i.e., a direct connection can be used between the control circuit of the circuit in Figure 3 and the queue manager of the circuit in Figure 7, but in other examples such as SoC or NoC, the form of interconnection described above can be provided.

[0109] It should be noted that in this embodiment, even a "queue manager" configuration may not be necessary. The second processing circuit may be provided within the CPU resources, and some or all of the instructions associated with the second mode may simply be executed using the CPU vector processing resources, but only using a second vector length (which may be the same as or different from the first vector length).

[0110] Simulator embodiment Figure 12 shows a simulator implementation that may be used with respect to a CPU alone, or a combination of a CPU and a coprocessor, or the configuration of Figure 11, as described above. While the above embodiments implement the present invention in terms of devices and methods for operating specific processing hardware that supports the art, it is also possible to provide an instruction execution environment according to the embodiments described herein, which are implemented through the use of a computer program. Such a computer program is often referred to as a simulator, insofar as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 1230 by optionally running a host operating system 1220 that supports the simulator program 1210. In some configurations, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that run at reasonable speeds, but such techniques may be justified in situations where it is desirable to run native code on a different processor for reasons of compatibility or reuse. For example, a simulator implementation may provide an instruction execution environment with additional features not supported by the host processor hardware, or it may provide an instruction execution environment typically associated with a different hardware architecture. An overview of the simulation is provided in "Some Efficient Architecture Simulation Techniques," Robert Bedichek, 1990 Winter USENIX Conference, pp. 53-63.

[0111] While embodiments have been described with reference to specific hardware components or functions, in simulated embodiments, equivalent functionality can be provided by appropriate software components or functions. For example, a particular circuit may be implemented as computer program logic in a simulated embodiment. Similarly, memory hardware such as registers or caches may be implemented as software data structures in a simulated embodiment. In configurations where one or more of the hardware elements referenced in the above embodiments reside in host hardware (e.g., host processor 1230), some simulated embodiments may use the host hardware where appropriate.

[0112] The simulator program 1210 may include, for example, instruction decoding program logic 1212, register emulation program logic 1214, and address space mapping program logic 1216, and may be stored in a computer-readable storage medium (which may be a non-temporary medium), and provides a program interface (instruction execution environment) to the target code 1200 (which may include an application, operating system, and hypervisor), and the program interface is the same as the application program interface of the hardware architecture modeled by the simulator program 1210. Thus, program instructions of the target code 1200, including the functions described above, may be executed from within the instruction execution environment using the simulator program 1210, thereby allowing a host computer 1230 that does not actually possess the hardware functions of the device discussed above to emulate these functions.

[0113] Therefore, the configuration in Figure 12 provides an example of a computer program for controlling a host data processing device to provide an instruction execution environment, which includes the following, when used to simulate the operation described with reference to Figure 3, for example: An instruction decoder that decodes processing instructions, One or more first registers, In a first processing mode in which a first processing circuit is configured to execute a decoded processing instruction using one or more first registers, the first processing circuit that executes the decoded processing instruction, A control circuit selectively initiates the execution of a decoded processing instruction in a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers. The instruction decoder is configured to decode a processing instruction selected from a first instruction set in a first processing mode and a processing instruction selected from a second instruction set in a second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for switching between a first processing mode and a second processing mode. An instruction execution environment is provided, in which the first processing circuit is configured to change the current processing mode between a first processing mode and a second processing mode in response to the execution of a mode change instruction. (Summary of the invention)

[0114] In this application, the term "configured to..." is used to mean that an element of the device has a configuration that enables it to perform a defined operation. In this context, "configuration" means the configuration or interconnection of hardware or software. For example, the device may have dedicated hardware to provide the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to..." does not mean that any modifications must be made to the device element to provide the defined operation.

[0115] While exemplary embodiments of the present invention are described in detail herein with reference to the accompanying drawings, it should be understood that the present invention is not limited to these precise embodiments, and various changes and modifications can be made in these embodiments by those skilled in the art without departing from the scope and spirit of the invention as defined by the appended claims.

Claims

1. It is a device, An instruction decoder that decodes processing instructions, One or more first registers, A first processing circuit, wherein in a first processing mode, the first processing circuit is configured to execute the decoded processing instruction using one or more first registers, the first processing circuit executes the decoded processing instruction, In a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers, the system includes a control circuit that selectively starts the execution of the decoded processing instruction, The instruction decoder is configured to decode a processing instruction selected from a first instruction set in the first processing mode and a processing instruction selected from a second instruction set in the second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for switching between the first processing mode and the second processing mode. The first processing circuit is configured to change the current processing mode between the first processing mode and the second processing mode in response to the execution of a mode change instruction. The aforementioned processing instruction includes a vector processing instruction, The one or more first registers comprises one or more first vector registers having a first vector length, The first processing circuit is configured to execute the decoded processing instruction using one or more first vector registers in the first processing mode, according to a vector length less than or equal to the first vector length. A device that, in the second processing mode, selectively executes the decoded processing instruction using one or more second vector registers having the second vector length, according to a vector length less than or equal to the second vector length.

2. The apparatus according to claim 1, wherein the first processing circuit is operable to execute processing instructions in the dominant exception level of an exception level hierarchy, each having a level of privilege for accessing the processing resources of the apparatus.

3. The apparatus according to claim 2, wherein the first processing circuit is configured to detect the vector length for executing the instruction in the first processing mode and the vector length for executing the instruction in the second processing mode by executing a vector length detection processing instruction.

4. The apparatus according to claim 2 or 3, wherein the first processing circuit is configured to execute a vector length setting processing instruction at an exception level other than the lowest exception level to set the maximum vector length applicable to that exception level and at least one lower exception level.

5. The apparatus according to claim 4, wherein the first processing circuit is configured to execute the vector length setting processing command to set the vector length used in the first processing mode to be less than or equal to the first vector length, and to set the vector length used in the second processing mode to be less than or equal to the second vector length.

6. The apparatus according to claim 5, wherein the first processing circuit is configured to execute the vector length setting processing instruction at a given exception level to set the vector length for use in a given processing mode to be less than or equal to the maximum vector length of the given processing mode set by executing the vector length setting processing instruction at an exception level higher than the given exception level.

7. The apparatus according to any one of claims 2 to 6, wherein the first processing circuit is configured to execute the mode change instruction at any exception level of the exception level hierarchy.

8. One or more second vector registers having the second vector length, A second processing circuit configured to execute the decoded vector processing instruction according to a vector length less than or equal to the second vector length, The apparatus according to any one of claims 1 to 7, comprising:

9. The apparatus according to any one of claims 1 to 8, wherein the vector processing instruction is independent of the vector length defined by the current processing mode.

10. The apparatus according to any one of claims 1 to 9, comprising a vector length register storage device configured to store vector lengths less than or equal to the first vector length for use in the first processing mode, and to store vector lengths less than or equal to the second vector length for use in the second processing mode, wherein the instruction decoder is configured to access the vector length register storage device according to the current processing mode.

11. The apparatus according to any one of claims 1 to 10, wherein the second vector length is different from the first vector length.

12. The apparatus according to claim 11, wherein the second vector length is greater than the first vector length.

13. The apparatus according to claim 8, wherein the control circuit is configured to detect whether a given processing instruction can be executed by the first processing circuit or the second processing circuit in the second processing mode.

14. The circuit comprises two or more instances of the first processing circuit, The instruction decoder of each instance of the first processing circuit is configured to selectively route processing instructions to a common second processing circuit. The apparatus according to any one of claims 1 to 13.

15. The second processing circuit comprises a matrix processing circuit and one or more matrix registers. The apparatus according to claim 8 or 14, wherein the second instruction set includes one or more matrix processing instructions that are unavailable in the first instruction set.

16. Data memory and An address translation circuit for converting between an input memory address in an input memory address space and an output memory address in an output memory address space, wherein the data memory is accessible according to the output memory address space, comprising: The control circuit is configured to initiate the conversion of an input memory address associated with a processing instruction for execution by the second processing circuit in the second processing mode, and to provide the converted output memory address to the second processing circuit. The apparatus according to claim 8, claim 14, or claim 15.

17. It is a method, Decoding processing instructions, In a first processing mode in which the first processing circuit is configured to execute the decoded processing instruction using one or more first registers, the execution of the decoded processing instruction is performed. In a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers, the execution of the decoded processing instruction is selectively initiated, The decoding includes decoding a processing instruction selected from a first instruction set in the first processing mode and a processing instruction selected from a second instruction set in the second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The decoding includes decoding one or more mode change instructions for switching between the first processing mode and the second processing mode. The aforementioned action involves changing the current processing mode between the first processing mode and the second processing mode in response to the execution of a mode change instruction. Includes, The aforementioned processing instruction includes a vector processing instruction, The one or more first registers comprises one or more first vector registers having a first vector length, The first processing circuit is configured to execute the decoded processing instruction using one or more first vector registers in the first processing mode, according to a vector length less than or equal to the first vector length. A method for selectively executing the decoded processing instruction in the second processing mode using one or more second vector registers having the second vector length, according to a vector length less than or equal to the second vector length.

18. A computer program that controls a host data processing device to provide an instruction execution environment, An instruction decoder that decodes processing instructions, One or more first registers, In a first processing mode in which a first processing circuit is configured to execute the decoded processing instruction using one or more first registers, the first processing circuit that executes the decoded processing instruction, In a second processing mode in which the decoded processing instruction is selectively executed using one or more second registers, the system includes a control circuit that selectively starts the execution of the decoded processing instruction, The instruction decoder is configured to decode a processing instruction selected from a first instruction set in the first processing mode and a processing instruction selected from a second instruction set in the second processing mode, wherein one or both of the first and second instruction sets include at least one instruction that is unavailable in the other of the first and second instruction sets. The instruction decoder is configured to decode one or more mode change instructions for switching between the first processing mode and the second processing mode. The first processing circuit is configured to change the current processing mode between the first processing mode and the second processing mode in response to the execution of a mode change instruction. The aforementioned processing instruction includes a vector processing instruction, The one or more first registers comprises one or more first vector registers having a first vector length, The first processing circuit is configured to execute the decoded processing instruction using one or more first vector registers in the first processing mode, according to a vector length less than or equal to the first vector length. A computer program that, in the second processing mode, selectively executes the decoded processing instruction using one or more second vector registers having the second vector length, according to a vector length less than or equal to the second vector length.