Processor system
Patent Information
- Application Number
- PCT/EP2026/058494
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058494_01102026_PF_FP_ABST
Abstract
Description
PROCESSOR SYSTEMFIELD OF THE INVENTION
[0001] The present disclosure relates to processor systems, and more particularly to a processor system, method, and computer program for improving processing performance through instruction translation and optimization using a processor interface unit.BACKGROUND
[0002] Processor systems have become increasingly complex and powerful over time, with advancements in microarchitecture design and instruction set architectures (ISAs) enabling improved performance and capabilities. Modern processors typically execute instructions according to a specific ISA, which defines the supported instructions, registers, memory addressing modes, and other aspects of the processor's programming model.
[0003] As processor designs have evolved, there has been a proliferation of different ISAs and microarchitectures tailored for various applications and use cases. This diversity can create challenges when developing software that needs to run efficiently across different processor types or when migrating existing code to new architectures. Additionally, the increasing complexity of modern ISAs and microarchitectures can make it difficult to fully utilize all of a processor's capabilities through conventional programming models.
[0004] One issue that arises is the potential mismatch between high-level algorithms or programming constructs and the low-level instructions supported by a particular processor architecture. This can lead to inefficiencies in code execution, as compilers may not always generate optimal instruction sequences for complex operations. Another challenge is that some desirable features or instructions may be available in one processor architecture but not in others, complicating software portability.
[0005] The growing demand for improved performance in areas such as artificial intelligence, cybersecurity, autonomous systems, data analysis, and real-time decision making has also highlighted limitations in how conventional processor systems handle certain types of computations. Vector and matrix operations, for instance, are common in many of these domains but may not be natively supported or optimally implemented in all processor architectures.
[0006] Additionally, there is a need to facilitate edge processing (applications on the periphery of a data processing ecosystem) within reduced performance / power consumption / cost limits, especially when the latency and / or security risk of exporting data to a data centre is unacceptable.
[0007] It has been appreciated that a processor system is needed that overcomes one or more of these problems.SUMMARY
[0008] The following disclosure provides various processor systems which comprise one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.
[0009] In other words, embodiments of this disclosure provide hardware which can act as a functional wrapper around traditional processor cores and / or processor elements, and which can enhance the functionality of the traditional processor cores and / or processor elements by, for example, modifying, translating or otherwise converting instructions before the instructions are passed to the core(s) and / or elements. This has many useful applications, including:
[0010] - Abbreviated instructions for common behaviours, where the processor interface unit is configured to decompress an abbreviated instruction into a corresponding set of standard instructions to cause the processor cores to perform a required operation. This can, for example, decrease the number of instructions which need to be retrieved from a remote memory, and reduce the risk that instruction retrieval times can act as a bottleneck to execution speed.
[0011] - Optimisation of instructions for a given processor core configuration. In contrast to softwarebased compilation, the processor interface unit can perform optimisation immediately prior to execution and can take into account the most detailed available picture of the available execution resources. This can include re-ordering of standard instructions in order to optimize usage of available resources. In some embodiments, the processor interface unit even monitors available resources in real-time and / or directly interacts with internal resources of processor cores.
[0012] - Performing calculations to handle a range of common engineering problems such as neural nets and / or Al, filtering and signal processing, cryptography, Position, Navigation and Timing (PNT),motor control, etc. For example, embodiments may be configured to efficiently perform one or more of:
[0013] — Matrix multiplication (such as matrix dot product), element-wise multiplication (such as matrix Hadamard product), or fast fourier transform,
[0014] — Calculations for various machine learning, statistical analysis and / or searching techniques including LSTM, CNN, Transformer, Probabilistic Graphical Models, Reinforcement Learning, Linear / Logistic Regression, Decision Tree, K means clustering, Gaussian Mixture Model, K nearest neighbours, Monte Carlo, Hidden Markov Model, Random Forest, Naive Bayes Classifier, Linear Discriminant Analysis, Breadth-First Search (BFS), Depth-First Search (DFS)
[0015] — Sorting and / or filtering algorithms
[0016] — Image processing algorithms
[0017] — C / C++ data structures and algorithms (DSA).
[0018] This may provide measurable benefits such as: fewer clock cycles taken to execute an algorithm or operation; fewer clock cycles taken to execute a repetitive stream or sequence of algorithms or operations; reduced power consumption by the one or more processor core(s) for the execution of algorithms or operations; reduced power consumption by the overall processor system for the execution of algorithms or operations; and / or reduced additional functionality (such as additional silicon area) required to increase the performance in a processor system.
[0019] Furthermore, customised embodiments of this disclosure can provide some of the benefits of ASIC design, while still using existing, mass-produced processor core components.
[0020] In particular, according to a first aspect, the following disclosure provides a processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.
[0021] For example, the configuration of the one or more processor cores may comprise a plurality of resources including processing resources, memory resources (e.g. registers, caches and any otherprocessor core memory) and / or communication pathway resources (e.g. input / output interfaces, or inter-core connections) in the one or more processor cores.
[0022] In more detail, in some embodiments of the first aspect, generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more processor cores; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more processor cores to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more processor cores.
[0023] Furthermore, in some embodiments of the first aspect, the processor interface unit is configured to: iteratively repeat the steps of determining a flow of processing sub-operations and allocating resources, and evaluate an expected performance characteristic for the one or more processor cores performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources. Such iterative improvement may be used on test examples of operations, in order to train a model which can be used to more immediately determine an improved flow of processing sub-operations and allocation of resources, for future operations. In other words, when a suitably trained model is available, the processor may additionally or alternatively be configured to use a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.
[0024] Furthermore, in some embodiments of the first aspect, the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more processor cores. Such access may, for example, be exposed by an existing commercially available processor core package. Alternatively, such access may be provided by a custom package in which the processor interface unit is at least partly integrated with the one or more processor cores.
[0025] According to a second aspect, the following disclosure provides a processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions; and control the one or more processor cores to perform theprocessing operation based on the one or more second processing instructions, wherein: the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, and the first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.
[0026] According to a third aspect, the following disclosure provides a processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive a first processing instruction associated with a processing operation, the first processing instruction comprising an instruction prefix and an instruction body; determine an algorithm for generating one or more second instructions based on the instruction prefix; generate one or more second processing instructions based on the algorithm and the instruction body; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.
[0027] In some embodiments, the processor interface unit is configured to maintain an internal state, and may receive state management instructions which manage this internal state. In embodiments of the third aspect, this internal state may be used when determining the algorithm based on the instruction prefix and the internal state. As one example, a prefix may indicate that the processor interface unit should use a vector processing algorithm (i.e. the processor interface unit should generate second processing instructions for vector processing). As further examples, a prefix may indicate that the interface unit should use a matrix processing algorithm, a tensor processing algorithm, or more generally any non-scalar format for data processing. This may, for example, be useful in order to select different algorithms based on a contextual state of the processor interface unit and / or of the one or more processor cores. This may also be useful in order to, for example, provide different externally-controlled general modes of the processor interface unit, wherein different modes may be associated with different algorithms for the same instruction prefix.
[0028] Each of the first, second and third aspects achieves benefits on their own, but any of these aspects can also be combined to provide even greater combined benefits.
[0029] Embodiments of the processor interface unit are preferably implemented in hardware which is as close as possible to the processor core(s), such as in a same chip or chiplet, or on a same circuit board. Nevertheless, in many embodiments, the processor interface unit is not integrated, or not fully integrated, with the processor core(s), and may be initially manufactured as a separate self-contained component, separate from the one or more processor core(s). For example, in some embodiments, the processor interface unit may be a coprocessor associated with the processor core(s).
[0030] Any of the methods performed by the processor interface unit may themselves be defined using executable instructions, such as a computer program, which may be stored in a memory.
[0031] In particular, according to a fourth aspect, the following disclosure provides a method performed by a processor interface unit, in a processor system comprising one or more processor cores and the processor interface unit, the method comprising: receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; and controlling the one or more processor cores to perform the processing operation based on the one or more second processing instructions.
[0032] According to a fifth aspect, the following disclosure provides a method performed by a processor interface unit, in a processor system comprising one or more processor cores and the processor interface unit, the method comprising: receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions; and controlling the one or more processor cores to perform the processing operation based on the one or more second processing instructions, wherein: the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, and the first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.
[0033] According to a sixth aspect, the following disclosure provides a method performed by a processor interface unit, in a processor system comprising one or more processor cores and theprocessor interface unit, the method comprising: receiving a first processing instruction associated with a processing operation, the first processing instruction comprising an instruction prefix and an instruction body; determining an algorithm for generating one or more second instructions based on the instruction prefix; generating one or more second processing instructions based on the algorithm and the instruction body; and controlling the one or more processor cores to perform the processing operation based on the one or more second processing instructions.
[0034] According to a seventh aspect, the following disclosure provides a computer program comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any one or more of the fourth, fifth and sixth aspects.
[0035] According to an eighth aspect, the following disclosure provides a non-transitory storage medium storing instructions which, when executed by one or more processors, cause the processors to perform a method according to any one or more of the fourth, fifth and sixth aspects.
[0036] According to a ninth aspect, the following disclosure provides a data signal comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any one or more of the fourth, fifth and sixth aspects.BRIEF DESCRIPTION OF THE FIGURES
[0037] Embodiments of the invention will be described, by way of example, with reference to the following drawings.
[0038] Fig. 1 is a block diagram schematically illustrating a processor system according to an embodiment;
[0039] Fig. 2 is a flow chart schematically illustrating a method performed by a processor interface unit, according to an embodiment;
[0040] Fig. 3 is a flow chart schematically illustrating a method performed by the processor interface unit, according to an embodiment;
[0041] Fig. 4 is a block diagram schematically illustrating various optional modules of the processor system, according to embodiments;
[0042] Figs. 5A to 5C are flow charts schematically illustrating further details of methods performed by the processor interface according to embodiments;
[0043] Fig. 6 is a block diagram schematically illustrating a processor system according to a partially integrated embodiment;
[0044] Fig. 7 is a block diagram schematically illustrating a processor system according to a partially integrated embodiment;
[0045] Fig. 8 is a block diagram schematically illustrating an example application of embodiments;
[0046] Fig. 9 is a block diagram schematically illustrating an example configuration for a re-ordering module;
[0047] Fig. 10 is a block diagram schematically illustrating a processor system according to an embodiment;
[0048] Fig. 11 is a flow chart schematically illustrating a method performed by a processor interface unit, according to an embodiment;
[0049] Fig. 12 is a flow chart schematically illustrating a further method performed in a processor system, according to an embodiment; and
[0050] Fig. 13 is a block diagram schematically illustrating further example embodiments of a processor system.
[0051] Common reference numerals are used throughout the figures to indicate similar features. DETAILED DESCRIPTION
[0052] The applicant has developed a system concept called VISC (Versatile Intrinsic Structured Computing), aspects of which are schematically explained in the following detailed disclosure. VISC uses a combination of new instructions and new hardware to improve the performance of microprocessor architectures. The new instructions are called "Abstracted VISC Instructions" (AVIs). These instructions can invoke functional elements of VISC's hardware in a manner optimised to a combination of one or more of: an algorithm or operation being executed; the data being processed; and the resource utilisation and availability in an underlying microprocessor architecture. The new hardware includes a number of functional elements which may be used individually or in combinations in order to accelerate certain functions defined by special or standard instructions. The elements may be logically positioned as a front-end pre-processor to one or more microprocessor core(s) and / or integrated within a microprocessor's internal pipeline. The aspects of VISC can be used together with various microprocessor configurations, to deliver performance improvements with minimal disruption to established software and hardware design techniques and methodology.
[0053] FIG. 1 illustrates a block diagram of a processor system. The processor system includes a processor interface unit 100 and one or more processor core(s) 200.
[0054] The processor core(s) may be commercially available standard processor cores. For example, each of the processor cores may be configured to use an instruction set architecture, ISA, such as open-source architectures like RISC-V or Power, or traditional proprietary architectures like x86 or Arm. Alternatively, each of the processor cores may be configured to use an applicationspecific architecture like MIPS or XMOS. Furthermore, each processor core may be configured to use a generic architecture like a neural processor architecture or a digital signal processor architecture. In embodiments where there is more than one processor core 200, the processor cores may be homogeneous or heterogeneous - in other words, the processor interface unit may be configured to interact with multiple cores using a same ISA or a mixture of cores each using a respective ISA. For example, in an embodiment, the processor cores 200 may include a RISC-V compatible core and an Arm core.
[0055] Some embodiments of a processor system may comprise managed processor elements, in addition or alternative to the processor core(s), or including the processor core(s). The managed processor elements may, for example, comprise compute units. The compute units may be specialised or may be general-purpose. The managed processor elements may be controlled according to the techniques described herein for controlling one or more processor core(s).
[0056] Each managed processor element may be configured to perform mathematical operations and / or may comprise one or more arithmetic logic units.
[0057] Additionally or alternatively, each managed processor element may comprise one or more interface elements suitable for communicating with other compute unit(s). For example, compute units may comprise one or more logic gates multiplexers or demultiplexers for sending or receiving data to or from other compute units.
[0058] The managed processor elements may be homogeneous. Alternatively, different managed processor elements may have different features and / or different arrangements.
[0059] In some examples, the managed processor units may be interconnected such that the plurality of managed processor elements performs operations as an interconnected fabric. In some examples, the managed processor units may be configured to perform operations as a systolic array. In some examples, the managed processor elements may be configured to operate as a field-programmable processing array.
[0060] The processor interface unit 100 can be integrated into a system design in various ways. In some examples, the processor interface unit 100 can be integrated as a front-end accelerator / translator / interpreter to one or more standard microprocessor cores. In other examples, the processor interface unit 100 can be fully integrated within the microprocessor architecture.
[0061] In some examples, the processor interface unit 100 may be configured to control one or more interconnections between pairs of processor cores and / or between pairs of managed processor elements. For example, the processor interface unit 100 may be configured to control interconnections between processor cores and / or between managed processor elements such that the processor cores and / or the managed processor elements perform operations as an interconnected fabric, or even as a systolic array or as a field-programmable processing array.
[0062] In an example, the processor interface unit 100 and the processor cores (and / or managed processor elements) 200 may be configured to operate together to provide a block capable of reconfigurable mathematics, to provide hardware-level efficiency improvements in applications such as performing calculations to handle a range of common engineering problems such as neural nets and / or Al, filtering and signal processing, cryptography, Position, Navigation and Timing (PNT), motor control, etc.
[0063] In some implementations, the processor interface unit 100 can be provided on a separate functional silicon dice, known as a chiplet, connected on a substrate to one or more microprocessor chiplets comprising the one or more processor core(s). Alternatively, the processor interface unit 100 can be provided on a separate functional chip connected on a printed circuit board to one or more microprocessor chips comprising the one or more processor core(s).
[0064] The processor interface unit 100 is configured to receive first processing instructions 110. The first processing instructions 110 may include a mixture of abstracted instructions 31 and / or core instructions 32. The first processing instructions 110 also include data such as operands within abstracted instructions 31 or core instructions 32.
[0065] Herein, it should be understood that core instructions 32 are instructions which can be received and executed by one or more of the processor cores 200 (and / or one or more of the managed processor elements).
[0066] On the other hand, abstracted instructions 31 are instructions intended for the processor interface unit 100. Abstracted instructions 31 may have different possible values, or a differentinstruction structure, from core instructions 32. For example, abstracted instructions 31 may have a shorter or longer bit length than core instructions 32. Additionally or alternatively, abstracted instructions may include a type header at the start of the instruction indicating that a payload body of the instruction is intended for the processor interface unit 100.
[0067] In some examples, the abstracted instructions 31 may be regarded as conventional instructions according to a known instruction set, such as RISC-V instructions. On the other hand, the core instructions 32 sent to one or more of the processor cores 200 (and / or one or more of the managed processor elements) may be instructions specially designated for controlling processing operations and / or state memory at the processor cores 200 (and / or one or more of the managed processor elements), and may be instructions of an application-specific instruction set or general computation instructions.
[0068] Various types of abstracted instruction 31 are discussed below and these may, for example, include state management instructions for managing a state of the processor interface unit 100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions or tensor instructions), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.
[0069] The processor interface unit 100 may be configured to receive instructions passively, at whatever rate they arrive at an input interface of the processor interface unit 100. Alternatively, the processor interface unit 100 may be configured to fetch instructions from a configured static or dynamic location. For example, on first powering up the processor system, the processor interface unit 100 may fetch BIOS instructions, or fetch instructions from a predetermined memory.Subsequently, the processor interface unit 100 may fetch instructions from various locations according to a dynamic state of the processor interface unit 100.
[0070] The processor interface unit 100 processes the first processing instructions 110 and generates second processing instructions 120. The second processing instructions typically comprise core instructions 32 which can be executed by one or more of the processor core(s) 200. In some embodiments, the second processing instructions may also comprise operand data 33, which is data to be provided to the one or more processor core(s) 200 outside of core instructions 32.
[0071] For example, when the first processing instructions comprise a core instruction 32, generating second processing instructions may simply comprise passing the core instruction 32 as a second processing instruction for the processor core(s) to execute. Additionally, the processor interface unit may perform some management operations, such as selecting which of a plurality of processor cores 200 should execute the core instruction 32.
[0072] On the other hand, when the first processing instructions comprise an abstracted instruction 31 , the process of generating second processing instructions 120 may be more complex. For example, generating the second processing instructions 120 may comprise any one or more of: identifying an algorithm for generating the second processing instructions 120; allocating resources at the processor core(s); organising vector, matrix, tensor and / or any type of non-scalar or SIMD processing operations; or re-ordering a sequence of generated instructions in order to more-optimally use available resources at the one or more processor cores.
[0073] The processor interface unit 100 provides the second processing instructions 120 to the processor core(s) 200 for execution.
[0074] For example, the processor interface unit 100 may be connected to an instruction input interface for each of the one or more processor core(s) 200, and may be configured to provide each of the second processing instructions 120 to one or more of the processor core(s) 200. For example, the processor core(s) 200 may be standard cores which are not adapted for, or in any way aware of, the processor interface unit 100.
[0075] The connection(s) between the processor interface unit 100 and the processor core(s) 200 may be exclusive connections, such that only the processor interface unit 100 can provide instructions to the one or more processor core(s) 200.
[0076] Alternatively, the processor interface unit 100 may share control of the one or more processor cores, and the processor system may additionally provide a direct interface for externally receiving core instructions 32 and providing the externally-received core instructions 32 to the one or more processor core(s) 200. However, this shared control is not preferred, as this reduces the extent to which the processor interface unit 100 can predict and / or manage available resources in the one or more processor core(s) 200.
[0077] As a compromise, the processor system may be configured to support dynamically switching between exclusive control and shared control of the one or more processor core(s). For example, theprocessor system may comprise a direct interface for externally receiving core instructions 32 as mentioned above, but the processor interface unit 100 may be configured with the capability to disconnect the direct interface in an exclusive mode, and to connect the direct interface in a nonexclusive mode.
[0078] Outputs from the processor core(s) 200 may be handled in a conventional manner. For example, outputs from the processor core(s) 200 may be stored in a volatile or non-volatile memory.
[0079] In some embodiments, the processor interface unit 100 may be configured to monitor or intercept outputs from the processor core(s) 200. For example, the processor interface unit 100 may be configured to generate further second processing instructions 120 based on outputs from the processor core(s) 200. In other words, in some embodiments, the processor system is configured to perform a sequence of inter-related steps without any communication outside of the processor system. This may, for example, be used to accelerate iterative calculations. This may be especially applicable when the processor interface unit 100 receives a first processing instruction 110 which is an algorithmic instruction. Such contained processing may also improve security, by reducing communication on an external data bus outside of the processor system, and enabling algorithms such as encryption to be performed with fewer opportunities for interception of sensitive data such as encryption keys.
[0080] In many embodiments, the processor interface unit 100 works in a real-time streaming configuration, simultaneously performing all of the functions of receiving first processing instructions, generating second processing instructions and controlling the processor core(s) 200. The processor interface unit 100 may also buffer selected first processing instructions 110 and / or second processing instructions 120, in order change the order of instructions or change the grouping of instructions, in order to more optimally use available resources at the one or more processor cores.
[0081] To summarise, the processor interface unit 100 is configured to act as an interface between the incoming first processing instructions 110 and the processor core(s) 200, translating and optimizing the instruction flow. This can enable various effects including: reducing the data quantity of first processing instructions supplied to the processor system (for example by using algorithmic instructions) and optimizing utilization of resources at the one or more processor core(s) 200.
[0082] In some embodiments, a processor interface unit 100 and one or more processor core(s) 200 may be a subsystem of a greater processor system.
[0083] For example, a processor system may comprise multiple processor interface units 100, each configured to control a respective one or more processor core(s) or managed processor elements. Additionally or alternatively a processor system may additionally comprise one or more conventionally-accessible processor(s) which are configured to receive instructions directly, without a processor interface unit 100.
[0084] In some processor systems, a subsystem comprising the processor interface unit 100 and the one or more processor core(s) 200 may be configured to act as a coprocessor together with one or more further subsystems. For example, the processor system may be configured to use a “main” processor subsystem for some common operations and to use the processor interface unit 100 and the one or more processor core(s) 200 to perform some specialised operations.
[0085] Preferably, the processor interface unit 100 is also supported through a development tool chain including compiler and assembler support for producing functions or programs comprising first processing instructions 110, suitable for the processor interface unit 100. This tool chain may allow developers to efficiently create and optimize code which takes advantage of the capabilities provided by the processor interface unit 100 and the processor core(s) 200. For example, this may enable mathematicians to take algorithms to code even if they are unfamiliar with a target microprocessor architecture, because the processor interface unit 100 can manage in hardware some details which the developer would have to consider in software for traditional microprocessor code.
[0086] Embodiments of the invention may also enable the creation of computer programs with reduced size, using abstracted instructions 31 and optionally also including core instructions 32, by comparison to a corresponding computer program consisting solely of core instructions 32.Furthermore, the use of abstracted instructions may reduce a total number of instructions to be read into the processor system, thereby potentially reducing latency and / or power consumption.
[0087] FIG. 2 is a flowchart schematically illustrating an outline method performed by the processor interface unit 100.
[0088] At step S201 , the processor interface unit 100 receives the first processing instructions 110 associated with a processing operation. The first processing instructions 110 may include a mixture of abstracted instructions 31 and / or core instructions 32. The first processing instructions 110 also include data such as operands within abstracted instructions 31 or core instructions 32.
[0089] Herein, it should be understood that core instructions 32 are instructions which can be received and executed by one or more of the processor cores 200.
[0090] On the other hand, abstracted instructions 31 are instructions intended for the processor interface unit 100. Abstracted instructions 31 may have different possible values, or a different instruction structure, from core instructions 32. For example, abstracted instructions 31 may have a shorter or longer bit length than core instructions 32. Additionally or alternatively, abstracted instructions 31 may include a type header at the start of the instruction indicating that a payload body of the instruction is intended for the processor interface unit 100.
[0091] Various types of abstracted instruction 31 are discussed below and these may, for example, include state management instructions for managing a state of the processor interface unit 100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions, tensor instructions or any other non-scalar or SIMD instruction format), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.
[0092] The processor interface unit 100 may be configured to receive instructions passively, at whatever rate they arrive at an input interface of the processor interface unit 100. Alternatively, the processor interface unit 100 may be configured to fetch instructions from a configured static or dynamic location. For example, on first powering up the processor system, the processor interface unit 100 may fetch BIOS instructions, or fetch instructions from a predetermined memory.Subsequently, the processor interface unit 100 may fetch instructions from various locations according to a dynamic state of the processor interface unit 100.
[0093] At step S202, the processor interface unit 100 generates the second processing instructions 120 based on the received first processing instructions 110. The second processing instructions typically comprise core instructions 32 which can be executed by one or more of the processor core(s) 200. In some embodiments, the second processing instructions may also comprise operand data 33, which is data to be provided to the one or more processor core(s) 200 outside of core instructions 32.
[0094] For example, when the first processing instructions comprise a core instruction 32, generating second processing instructions may simply comprise passing the core instruction 32 as a second processing instruction for the processor core(s) to execute. Additionally, the processor interface unitmay perform some management operations, such as selecting which of a plurality of processor cores 200 should execute the core instruction 32.
[0095] On the other hand, when the first processing instructions comprise an abstracted instruction 31 , the process of generating second processing instructions 120 may be more complex. For example, generating the second processing instructions 120 may comprise any one or more of: identifying an algorithm for generating the second processing instructions 120; allocating resources at the processor core(s); organising vector, matrix, tensor and / or any type of non-scalar or SIMD processing operations; or re-ordering a sequence of generated instructions in order to more-optimally use available resources at the one or more processor cores.
[0096] At step S203, the processor interface unit 100 controls the processor core(s) 200 to perform the processing operation based on the second processing instructions 120.
[0097] FIG. 3 is a flowchart schematically illustrating a method for processing instructions performed by the processor interface unit 100, according to an embodiment.
[0098] The method begins at a step S301 , where the processor interface unit 100 receives a first processing instruction 110. The first processing instruction 110 is an abstracted instruction 31 and / or a core instruction 32. The first processing instruction 110 also include data such as an operand within the abstracted instruction 31 or core instruction 32.
[0099] Herein, it should be understood that core instructions 32 are instructions which can be received and executed by one or more of the processor cores 200.
[0100] On the other hand, abstracted instructions 31 are instructions intended for the processor interface unit 100. Abstracted instructions 31 may have different possible values, or a different instruction structure, from core instructions 32. For example, abstracted instructions 31 may have a shorter or longer bit length than core instructions 32. Furthermore, different sub-types of abstracted instruction 31 may have different bit lengths. Additionally or alternatively, abstracted instructions 31 may include a type header at the start of the instruction indicating that a payload body of the instruction is intended for the processor interface unit 100.
[0101] Various sub-types of abstracted instruction 31 are discussed below and these may, for example, include state management instructions for managing a state of the processor interface unit 100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions, tensor instructions or any other non-scalar or SIMDinstruction format), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.
[0102] At a step S302, the processor interface unit 100 determines an instruction type of the first processing instruction 110.
[0103] For example, the processor interface unit may determine whether the first processing instruction 110 is a core instruction 32 or an abstracted instruction 31. This may, for example, be determined based on a value or a structure of the first processing instruction 110. For example, an abstracted instruction 31 may be detected based on a value, a bit length or a type header of the first processing instruction 110. Similarly, a core instruction 32 may be detected based on a value, a bit length or other structure of a core instruction 32.
[0104] The processor interface unit 100 may, for example, store a set of properties associated with different instruction types. The processor interface unit 100 may determine whether the received first processing instruction 110 matches the stored properties of any of the instruction types. For each instruction type, these stored properties may comprise a partial or complete instruction set architecture (ISA) specification.
[0105] If the first processing instruction is a core instruction 32 then, at step S303, the processor interface unit 100 passes the instruction unchanged as a second processing instruction 120 to the processor core(s) 200. Additionally, the processor interface unit 100 may perform some management operations for the core instruction 32, such as selecting which of a plurality of processor cores 200 should execute the core instruction 32.
[0106] On the other hand, if the first processing instruction 110 is an abstracted instruction 31 , the method proceeds to step S304, where the processor interface unit 100 generates one or more second processing instructions 120. The second processing instructions typically comprise core instructions 32 which can be executed by one or more of the processor core(s) 200. In some embodiments, the second processing instructions may also comprise operand data 33, which is data to be provided to the one or more processor core(s) 200 outside of core instructions 32. Some example implementations of step S304 are illustrated using Figs. 4 and 5A to 5C as discussed below.
[0107] Following the generation of the second processing instructions 120, the method moves to a step S305, where the processor interface unit 100 passes the second processing instruction(s) 120 to the processor core(s) 200.
[0108] According to the method of Fig. 3, the processor interface unit 100 can pass through core instructions 32, as well as handling abstracted instructions 31. This enables the processor system to be used as a standard processor system having standard cores, in a case where no specific abstracted instructions 31 have been generated. In other words, the processor system can run compiled code which was compiled without being specifically intended for a processor system according to the invention, increasing flexibility. At the same time, the processor system can provide the described advantages - including reduced program size, and increased execution efficiency - for programs which include abstracted instructions 31.
[0109] FIG. 4 is a block diagram schematically illustrating a processor interface unit 400 having modules according to some more detailed embodiments. The processor interface unit 400 may be used in processor systems, similar to the processor interface unit 100 of Fig. 1. According to various embodiments, different combinations of these modules may be included or omitted, depending on the requirements of a given application of the technology. The functions of the modules are independent, so each module can be included or omitted independently. The modules are only shown together in a single figure for brevity.
[0110] In some embodiments, the processor interface unit 400 includes a first instruction buffer module 410 for storing received first processing instructions 110. Including a buffer enables the processor interface unit 400 to receive multiple first processing instructions 110, and to generate second processing instructions 120 based on the multiple first processing instructions. For example, a first abstracted instruction 31 of the first processing instructions 110 may be a vector processing instruction, indicating that the processor interface unit 400 should control the one or more processor cores to perform vector processing, and a second abstracted instruction 31 of the first processing instructions may contain data for the vector processing.
[0111] In some embodiments, the processor interface unit 400 includes a fetch module 415.
[0112] The fetch module 415 may be configured to fetch first processing instructions 110. For example, the processor interface unit 100 may be configured to fetch instructions from a configured static or dynamic location. For example, on first powering up the processor system, the processor interface unit 100 may fetch BIOS instructions, or fetch instructions from a predetermined memory. Subsequently, the processor interface unit 100 may fetch instructions from various locations according to a dynamic state of the processor interface unit 100.
[0113] Additionally or alternatively, the fetch module 415 may be configured to retrieve further instructions or data based on previously processed first processing instructions 110. For example, where first processing instruction 110 is an algorithm instruction or a vector processing instruction, the fetch module 415 may retrieve any additional data required for the algorithm or for the vector processing, and buffer the additional data in a memory at the processor interface unit 400. Thus, when the one or more processor core(s) 200 are subsequently controlled to perform steps of the algorithm or steps of the vector processing, any required additional data is ready to be swiftly provided to the one or more processor core(s), and the processing can be performed efficiently by the one or more processor core(s), without data fetch delays in the middle of the processing at the one or more processor core(s). The above example is equally applicable to matrix processing, tensor processor, and other non-scalar and / or SIMD processing types, as alternatives to vector processing.
[0114] In some embodiments, the processor interface unit 400 includes a second instruction buffer module 420 for storing generated second processing instructions 120 before they are sent to the one or more processor core(s). For example, the second instruction buffer module 420 may be used to control instruction timing, so that instructions are sent to a suitable processor core 200 which has available resources. As another example, storing second processing instructions in the second instruction buffer module 420 may enable re-ordering of the second processing instructions after further second processing instructions have been generated based on subsequent first processing instruction(s). This may, for example, help to make more efficient usage of available processor core resources.
[0115] In some embodiments, the processor interface unit 400 comprises an interface state module 430 configured to maintain state information for the processor interface unit 400.
[0116] This interface state may, for example, include any of: information about ongoing instructions which are being executed in the one or more processor cores 200; information about planned instructions which are due to be executed in the one or more processor cores 200; information about iterative processes in which an output from the one or more processor cores 200 may be used when generating a further second processing instruction 120; information about processes which require the fetching of additional data; information about an algorithm which is currently being executed or which is due to be executed.
[0117] The processor interface unit 400 may, for example, use any of this interface state information when controlling the fetch module 415, the vector execution module 440, the SIMD module 450, the algorithm module 460, or the re-ordering module 490.
[0118] In some embodiments, the processor interface unit 400 comprises a vector execution module 440 configured to handle vector processing in the processor system. Herein, it should be understood that vector processing comprises performing a same instruction yt= f xj multiple times, for each of multiple operand values (a "vector" of operand values x), in order to produce multiple result values (a "vector" of output values y. The vector processing unit is configured to manage execution resources for vector processing such that, when sufficient resources are available at the one or more processor core(s) 200, multiple copies of the instruction for different operands are executed simultaneously using different processing resources. This managed vector processing operation can be contrasted with a for-loop methodology found in known vector processing systems, in which a loop is executed for successive executions of the instruction using a same processing resource, meaning that the processing time scales with the number of operand values in the "vector", and the vector elements of the output from the instruction become available sequentially. Furthermore, even when parallel execution is not possible for all elements of the "vector" of operand values, the vector execution module 440 may be configured to start a downstream operation on elements of the vector result z = g y^, even before the vector result has been fully calculated, if the processor interface unit 400 has received first processing instructions corresponding to such a sequence of operations (for example in the case of an algorithmic instruction). Additionally or alternatively, the processor interface unit 400 may comprise a matrix execution module or a tensor execution module configured for performing a same instruction for each of a multi-dimensional collection of operand values.
[0119] In some embodiments, the processor interface unit 400 comprises a SIMD module 450 configured to control SIMD (single-instruction multiple-data) execution resources in the one or more processor core(s) 200. The SIMD module has similar uses to the vector execution module, but makes use of SIMD execution resources in the one or more processor core(s) 200 which support parallel execution for multiple operands, on a hardware level. For example, the SIMD module may rearrange a data sequence according to a configuration of the one or more processor core(s). More specifically, the SIMD module may rearrange a data sequence to fit available SIMD resources, according to their respective capacities for multiple data in each execution step. Additionally, in a scenario where a datalength (such as a data word length) is shorter than a corresponding register length in an available SIMD resource, the SIMD module 450 may adjust alignment of data in registers to improve execution order and efficiency for execution across a broader set of operand data.
[0120] In some embodiments, the processor interface unit 400 comprises an algorithm module 460 configured to control execution of any of one or more predetermined algorithms. As examples, the predetermined algorithms may include frequently used mathematical operations such as Fast Fourier Transform, Discrete Cosine Transform, matrix operations, encoding operations or encryption operations, or ML model operations performed while training or operating an ML model. The predetermined algorithms need not be basic elementary algorithms, and can be complex algorithms. The predetermined algorithms may, in some cases, be tailored to individual customers, depending on which algorithms are frequently needed for their use case.
[0121] Each predetermined algorithm may simply comprise reading a predetermined sequence of second processing instructions from a memory, and adding any required operands. Alternatively, the pre-defined algorithm may comprise rules for generating the second processing instructions, without the sequence of second processing instructions being fully defined in advance. This may, for example, be necessary for more complex branching algorithms, or for feedback algorithms (such as iterative algorithms) where the total length of the algorithm may not be known in advance.
[0122] In addition to predetermined algorithms, the algorithm module 460 may support dynamically updated algorithms. For example, the algorithm module may be configured to optimise a predetermined algorithm over time based on monitored performance when the algorithm is used in the processor system. Such optimisation may comprise evaluation of candidate modified versions of the predetermined algorithm, in order to determine whether any of the modified versions provide improved performance. Candidate modified versions of the predetermined algorithm may be generated randomly, or may be generated using a trained machine learning model. Using a machine learning model may have the benefit of attempting modifications in parts of the algorithm which are likely to be improvable (such as loops which increase memory usage over time), while reducing the risk of modifications which are likely to break the algorithm.
[0123] For example, the algorithm module 460 may be implemented together with an algorithm library module 470 which stores predetermined algorithms for generating second processing instructions 120.
[0124] When controlling the one or more processor core(s) to execute second processing instructions 120 corresponding to a predetermined algorithm, the processor interface unit 400 may use a local memory, such as the interface state module 430, in order to perform the complete algorithm with minimal instances of storing data in an external memory or retrieving data from an external memory, thus substantially eliminating a slow operation type, and increasing the execution speed and efficiency for the predetermined algorithm. Furthermore, the processor interface unit 400 may use the fetch module 415 in order to anticipate any fetch operations so that data is retrieved before it is required, and the processor core(s) 200 do not need to wait for fetch operations any more than necessary.
[0125] In some embodiments, the processor interface unit 400 is configured to process specific subtypes of abstracted instruction 31 , which may feed into the control of the previously described modules.
[0126] As a first example, prefix instructions may be processed as a sub-type of abstracted instruction 31. Prefix instructions may be used to set up certain repetitive operations such as vector processing. Prefix instructions may for example include a vector processing instruction indicating that a subsequent instruction should be repeated for each of a plurality of operands. This can be contrasted with known systems which define vector instructions as separate instructions in the ISA-in other words, in this application, the combination of a prefix instruction with a subsequent instruction can enable vector processing for any subsequent instruction, without requiring the processor interface unit 400 to be configured to distinctly process a defined vector instruction (multiple-data instruction) for each single-data instruction. Prefix instructions may be shorter than other abstracted instructions 31 (in terms of bit length). The processor interface unit 400 may use a vector processing instruction as a trigger to perform operations using the vector execution module 440, and / or to allocate space in a memory (for example in the interface state module) for vector processing. This can be contrasted with known systems which have a bank of registers dedicated to vector operations, where those registers are not used for any other purpose, leading to inefficient usage of resources.
[0127] As a second example, algorithmic instructions may be processed as a sub-type of abstracted instruction 31. An algorithmic instruction may comprise a reference to a predetermined algorithm in the algorithm library module 470. For example, the algorithmic instruction may comprise a pointer to a memory location in the algorithm library module 470, in a case where the contents of the algorithmlibrary module 470 are fixed and known to external entities producing first processing instructions 110 - such as compilers configured to produce algorithmic instructions. For example, a manufacturer of a processor interface unit 100 or a processor system according to this application may supply, to customers, an SDK (Software Development Kit) for the processor interface unit. This SDK may include the available algorithms which are stored in the algorithm library module 470, and which can be efficiently incorporated into programs executed on the processor system. The algorithm library module 470 (and the corresponding SDK) may, in some cases, be tailored to individual customers, depending on which algorithms are frequently needed for their use case. In some embodiments, the processor interface module 400 may additionally support adding algorithms to the algorithm library module 470 and / or removing algorithms from the algorithm library module 470. For example, after a processor interface module 400 has been manufactured (such as when the processor interface module 400 is installed in a specific computer system) a user may wish to add, modify or remove algorithms, in order to customise the processor interface module 400 to be maximally beneficial for their specific use case. In other words, while the application refers to "predetermined" algorithms, these algorithms may be "predetermined" at any stage prior to an instance of executing first processing instructions 110. For example, the processor system may be used to perform operations corresponding to a first set of first processing instructions 110 with a first set of predetermined algorithms in the algorithm library module 470, then modified to replace the first set of predetermined algorithms with a second set of predetermined algorithms in the algorithm library module 470, and then used again to perform operations corresponding to a second set of first processing instructions 110 with the second set of predetermined algorithms in the algorithm library module 470. The system may be configured to support changes to the algorithms stored in the algorithm library module 470 as a regular, generally available operation. Alternatively, system may be configured to support changes to the algorithms stored in the algorithm library module 470 in the form of a firmware update. As mentioned above, the predetermined algorithms need not be elementary algorithms, and may be complex - nevertheless, each predetermined algorithm may advantageously be identified using a single algorithmic instruction.
[0128] In some examples, algorithmic instructions may comprise a prefix and a body, where the prefix comprises a reference to a predetermined algorithm, and the body comprises data to be used as one or more parameters in the algorithm. In such cases, an algorithmic instruction may beimplemented as two or more of the first processing instructions 110 - an algorithm prefix instruction, and one or more algorithm body instructions. When the algorithm prefix instruction is received, the processor interface unit 400 may store information in the interface state module 430 indicating that one or more algorithm body instructions is expected subsequently in the first processing instructions 110. Alternatively, the algorithmic instruction may be implemented as a single instruction including both of a prefix and a body.
[0129] As a third example, management instructions may be processed as a sub-type of abstracted instruction 31. Management instructions may be used to configure various settings and behaviours of the processor interface unit 400, and data indicating settings may be stored in the interface state module 430. Examples of management instructions include:
[0130] - Step
[0131] A Step instruction is a management instruction which may be used to increment or decrement a memory offset. For example, the memory offset may indicate a location in a memory resource (such as a register) of a processing core 200. More specifically, the memory offset may indicate a register file offset. The processor interface unit 400 may additionally or alternatively be configured to automatically increment a memory offset. However, manually controlling a memory offset may be useful in some scenarios, such as processing multiple first processing instructions 110 without incrementing a register file offset.
[0132] - Set Vector Length
[0133] A Set Vector Length instruction is a management instruction which may be used to set a vector length for vector processing handled by the vector execution module 440. In other words, the vector execution module 440 may be configured to read a vector length setting and, after receiving a vector processing instruction and a subsequent instruction in the first processing instructions 110, generate a number of second processing instructions 120 based on the vector length setting and based on the subsequent instruction. For example, if the subsequent instruction is a core instruction 32, the vector execution module 440 may generate (vector length) copies of the core instruction 32.
[0134] The Set Sequence Length instruction may also be used to set a vector operating mode, such as setting an "ignore vector processing instructions" mode, a mode in which a register file offset is automatically incremented for each element of a vector of operations, or a mode in which a register file offset is not incremented for each of (vector length) operations. That last mode may be used to"copy" a core instruction 32 for each operation, by using the same register location multiple times, without actually generating copies of the core instruction 32.
[0135] - Store Algorithm
[0136] A Store Algorithm instruction is a management instruction which may be used to indicate that subsequent first processing instructions 110 should be stored in the algorithm library module 470. For example, this may be used to add, remove, or update the predetermined algorithms stored in the algorithm library module 470.
[0137] - Configure Re-order
[0138] A Configure Re-order instruction is a management instruction which may be used to configure the re-ordering module 490. For example, this may be used to set sequence generator rules for rearranging instructions and operations.
[0139] In some embodiments, the processor interface unit 400 comprises a processor core state module 480 that maintains information about the state of the processor core(s) 200. This information may, for example, be used by the reordering module 490 to make informed decisions about instruction re-ordering or may be used by the second instruction buffer module 420 in order to make decisions about when to send second processing instructions 120 from the processor interface unit 400 to the one or more processor cores 200. Additionally, the processor core state module 480 may store a history of the state of the processor core(s) 200. For example, the history may comprise, for each of a plurality of time steps, associated data of a state of the processor core(s) 200 and data a state of the processor interface unit 400. By using such processor core state information (and optionally using historical state information), the processor system can help to reduce inefficiencies such as waiting for busy processing resources, or errors or delays due to simultaneous attempts to access a same memory resource, which can in turn lead to stall events, wasted clock cycles or wasted power.
[0140] As discussed below, in embodiments where the processor interface unit has read-access and / or write-access to resources within the one or more processor core(s), the processor interface unit may directly monitor availability of resources in the one or more processor core(s).
[0141] Alternatively, the processor core state module 480 may predict usage of resources within the one or more processor core(s), based on the configuration of the one or more processor core(s). For example, commercially-available processor cores may include a known combination of processing resources (such as decoders, arithmetic logic units, floating point units, SIMD units or adders),memory resources (such as registers, caches and any other processor core memory) and / or communication pathway resources (such as input / output interfaces, or inter-core connections). The processor interface unit may use a known configuration for each of the one or more processor cores, together with a knowledge of the instructions (and data) which have been input to the processor cores, in order to predict a current state of the resources in the one or more processor cores.
[0142] The processor core(s) state may, for example, include any of: a configuration of the one or more processor cores 200; a monitored current state of the one or more processor cores 200; a predicted current state of the one or more processor cores 200; data or statistics regarding historical performance of the one or more processor cores 200.
[0143] The processor interface unit 400 may, for example, use any of this processor core(s) state information when controlling the fetch module 415, the vector execution module 440, the SIMD module 450, the algorithm module 460, or the re-ordering module 490.
[0144] In some embodiments, the processor interface unit comprises a re-ordering module 490 configured to re-arrange instructions and operations. For example, the re-ordering module may be configured to apply a set of sequence generator rules for re-arranging instructions and operations. Additionally or alternatively, the re-ordering module may comprise one or more sequence generators, an example of which is described below with reference to Fig. 9.
[0145] Such re-ordering of instructions is useful in a wide variety of cases, to improve execution speed and / or improve efficient usage of processing resources. In particular, in many cases, it is useful for the re-ordering module 490 to re-order a sequence of second processing instructions 120 based on a configuration of the one or more processor core(s) 200. Additionally, such re-ordering may ensure that operands can be accessed using a linear order of data access operations, reducing or avoiding the need for branched data access when second processing instructions 120 are executed in the core(s) 200.
[0146] As a first case, the re-ordering module 490 may be configured to re-order execution of a sequence of second processing instructions 120 depending on a predicted availability of resources in the one or more processor core(s).
[0147] Based on actual and / or predicted knowledge of resource usage in the one or more processor cores, the processor interface unit may try to optimise an order of execution for the second processing instructions 120 at the one or more processor cores 200.
[0148] As a second case, the re-ordering module 490 may be configured to determine a linear or non-linear sequence for loading operand data into registers. As discussed further below, in embodiments where the processor interface unit has write-access to resources within the processor cores, the re-ordering module may determine a sequence for loading operand data into registers within the one or more processor cores. Additionally or alternatively, the re-ordering module 490 may be configured to determine a sequence for loading operand data into registers of the processor interface unit, such as for inclusion in second processing instructions 120.
[0149] As a third case, the re-ordering module 490 may be configured to calculate register offsets or other memory locations for storing and / or retrieving operand data, to enable a computed execution on the operands in a linear or non-linear order.
[0150] As a fourth case, the re-ordering module 490 may be configured to calculate register offsets or other memory locations for storing and / or retrieving output data following a calculation at the one or more processor cores. Such output destinations may be configured using additional second processing instructions 120 or operands within the second processing instructions.
[0151] As a fifth case, the re-ordering module may be configured to re-order operands in registers for efficient execution of SIMD operations, whether as part of vector processing or in other contexts.
[0152] In some embodiments, the processor interface unit 400 also includes a memory management module and / or a garbage collection module (not shown), in order to manage any data that it stores. For example, a memory management module may allocate memory space for a matrix calculation, until a result is ready to be output from the processor system. As another example, data from intermediate steps of an algorithm, which may be stored while the processor interface unit 400 is controlling execution of an algorithm, is preferably cleared up as soon as it is no longer required for the algorithm.
[0153] In some embodiments, the processor interface unit 400 also includes a pointer alignment module and / or a programme counter module. Misalignment due to a branch instruction or a jump instruction can sometimes require re-alignment of pointers and memory address offsets, or flushing of data pipelines. Such misalignments may be detected or predicted based on a state of the processor core(s) or based on the second processing instructions 120. By quickly detecting, or even predicting, misalignments, some interruptions and latency can be avoided, improving overall performance.
[0154] The modules within the processor interface unit 400 interact to optimize instruction execution. For example, the first instruction buffer module 410 may distribute first processing instructions 110 to the appropriate processing modules (e.g., vector execution module 440, SIMD module 450, algorithm module 460) based on instruction type. The re-ordering module 490 may then coordinate with these processing modules to optimize instruction execution order based on information from the processor core state module 480 and the interface state module 430.
[0155] Through these interactions, both in terms of the individual modules, and the combined effects of the multiple modules, the processor interface unit 400 can control the processing core(s) 200 to more efficiently perform tasks indicated by the first processing instructions 110.
[0156] Referring again to Fig. 3, the processor interface unit 100, 400 (which at least has functionality as described for Fig. 1 , and may further have functionality as described for Fig. 4) can perform various functions within the step of "generating one or more second processing instructions" (step S304). Examples of these detailed functions will now be described with reference to Figs. 5Ato 5C.
[0157] FIG. 5A is a flow chart schematically illustrating a first embodiment S304-A of step S304 of Fig. 3.
[0158] At a step S501 , the processor interface unit 100, 400 determines a configuration of the processor core(s) 200. This configuration may include information about available processing resources, memory resources, and communication pathways within the processor core(s) 200.
[0159] This configuration may be provided in advance to the processor interface unit 100, 400. For example, the processor interface unit 100, 400 this configuration may be stored in a processor core(s) state module 480, and step S501 may simply comprise accessing the predetermined configuration. For example, commercially-available processor cores may include a known combination of processing resources (such as decoders, arithmetic logic units, floating point units, SIMD units or adders), memory resources (such as registers, caches and any other processor core memory) and / or communication pathway resources (such as input / output interfaces, or inter-core connections). A known configuration for each of the one or more processor cores may be stored in the processor core(s) state module 480.
[0160] Additionally or alternatively, the processor interface unit 100, 400 may have access to monitor the one or more processor core(s) and determine the configuration of the processor core(s) based on observations.
[0161] Then, at an optional step S502, the processor interface unit 100 determines a current state of the processor core(s) 200. This state information may, for example, include details about the current utilization of resources and the status of ongoing operations.
[0162] As discussed below, in embodiments where the processor interface unit has read-access and / or write-access to resources within the one or more processor core(s), the processor interface unit may directly monitor availability of resources in the one or more processor core(s).
[0163] Alternatively, the processor core state module 480 may predict usage of resources within the one or more processor core(s), based on the configuration of the one or more processor core(s). The processor interface unit may use a known configuration for each of the one or more processor cores, together with a knowledge of the instructions (and data) which have been input to the processor cores, in order to predict a current state of the resources in the one or more processor cores.
[0164] The process then moves to a step S503, where the processor interface unit 100, 400 determines a flow of processing sub-operations. This flow represents the sequence of operations required to execute the first processing instructions 110. In other words, before generating second processing instructions - which are real-world executable instructions designed for the specific execution context in the one or more processor core(s) - the processor interface unit 100, 400 may perform a functional or conceptual assessment of the required operations, and break this functional or conceptual assessment into a series of steps.
[0165] At a step S504, the processor interface unit 100, 400 allocates resources in the processor core(s) 200 to perform each sub-operation. This allocation is based on the determined configuration and current state of the processor core(s) 200, as well as the required processing for each suboperation.
[0166] Finally, at step S505, the processor interface unit 100 generates the second processing instructions 120 based on the determined flow of sub-operations, and based on the resources allocated for each sub-operation. This generation process may, for example, involve the re-ordering module 490 to optimize the instruction sequence based on the available resources and current processor state.
[0167] Further to the steps illustrated in Fig. 5A, the determining of sub-operations (step S503) and the allocation of resources (step S504) may be an iterative process. For example, the processor interface unit 100, 400 may be configured to predict a performance characteristic (e.g. total processing resources occupied over time, or total time to reach a final result) of a given flow of processing sub-operations and a given allocation of resources. The processor interface unit may use such a predicted performance characteristic in order to compare multiple candidate implementations, each comprising a flow of processing sub-operations and an allocation of resources. The processor interface unit 100, 400 may use such comparisons in order to iteratively locate an improved implementation which may be the best available implementation, or may be an implementation which exceeds a threshold performance characteristic. Alternatively, rather than iteration, the processor interface unit 100, 400 may simply compare a set of candidate implementations in parallel, and select the best of the candidate implementations.
[0168] Such optimisation strategies may comprise re-ordering processing sub-operations, or even re-ordering operations at the level of the received first processing instructions 110. For example, when the first processing instructions 110 comprise an abstracted instruction 31 followed by a core instruction 32, the steps may in some circumstances be re-ordered so that functionality corresponding to the core instruction 32 is executed by the one or more processor core(s) 200 before functionality corresponding to the abstracted instruction 31 by the one or more processor core(s) 200.
[0169] Of course, such optimisations as described in the preceding paragraphs may be difficult to perform in real-time. As an alternative, a model may be trained using machine learning for determining the flow of processing sub-operations (S503) and allocating resources (S504) based on the first processing instructions 110 and the processor core configuration (S501) - and optionally based on the processor cores current state (S502). Such a pre-trained model may be used to implement steps S503 and S504. A training phase for training the model may be performed for a range of training scenarios with different first processing instructions, different processor core configurations and different processor core current states. This training phase may involve iterative and / or parallel evaluation of candidate implementations, as discussed in the preceding paragraphs.
[0170] As a further extension to the method of Fig. 5A, the processor interface unit 100, 400 may be configured to store determined flows of processing sub-operations (S503) and / or store resource allocations (S504) for future reference. For example, if a sequence of operations is performedrepeatedly, the flow of sub-operations the resources may be determined and allocated a first time the sequence is performed. On the other hand, in subsequent times when the sequence is performed, steps S503 and S504 may be efficiently implemented by reading the determined sub-operations and the allocated resources from memory. In other words, the efficiency gains of using a processor interface unit 100 may increase over time, as the processor system learns how to efficiently handle regular sequences of operations.
[0171] FIG. 5B is a flow chart schematically illustrating a second embodiment S304-B of step S304 of Fig. 3. The embodiments S304-A and S304-B provide different but compatible functions, and they may be implemented separately or together.
[0172] At step S511 , the processor interface unit 100, 400 determines an architecture of the processor core(s) 200.
[0173] The processor core(s) 200 may be commercially available standard processor cores. For example, each of the processor cores may be configured to use an instruction set architecture, ISA, such as open-source architectures like RISC-V or Power, or traditional proprietary architectures like x86 or Arm. Alternatively, each of the processor cores may be configures to use an applicationspecific architecture like MIPS or XMOS. Furthermore, each processor core may be configured to use a generic architecture like a neural processor architecture or a digital signal processor architecture.
[0174] The determination at step S511 may involve identifying, for each of the one or more processor core(s), an instruction set architecture (ISA) for instructions which can be externally received by the processor core(s). In other words, the processor interface unit 100, 400 may determine a set of available instruction types which can be used as second processing instructions 120 for communicating with each of the one or more processor core(s). In embodiments where there is more than one processor core 200, the processor cores may be homogeneous or heterogeneous - in other words, the processor interface unit may be configured to interact with multiple cores using a same ISA or a mixture of cores each using a respective ISA.
[0175] Additionally or alternatively, this determination at step S511 may involve determining, for each of the one or more processor core(s), an architecture for instructions and data which can be used internally within the processor core. For example, the architecture may be a microarchitecture of the processor core. This may be particularly applicable in embodiments where the processor interface unit 100, 400 has read-access and / or write-access to resources within the one or more processorcore(s) - for example in embodiments where the processor interface unit 100, 400 is partially or fully integrated with the processor core(s) in a single package (such as a chip or a chiplet).
[0176] Following the architecture determination, at step S512, the processor interface unit 100 generates the second processing instructions 120 according to the architecture of the processor core(s) 200.
[0177] This step S512 may involve generating some ISA-specific instructions that can be received and executed by the processor core(s) 200.
[0178] As previously mentioned, some of the first processing instructions may be core instructions 32 suitable for execution in a processor core. Nevertheless, the processor interface unit 100, 40 may control instructions to be executed in a core which uses a different ISA from the received core instruction 32. Different ISAs may have some instruction types which are the same in both architectures, but in general a first ISA will include at least one instruction type that is not included in a second, different ISA, or the first ISA will lack at least one instruction type that is included in the second, different ISA. In such cases, step S512 may comprise processing a first processing instruction 110 which is a core instruction 32 conforming to a first ISA, and generating a second processing instruction 120 according to a second ISA different from the first ISA, where the second processing instruction 120 is sent to a processor core 200 instead of the received first processing instruction 110. By this mechanism, the processor system may control "execution" of received core instructions 32 even when the received core instructions do not match an ISA of an available processor core.
[0179] This step S512 may also involve generating some second processing instructions which are designed to be inserted using access to the internal structure of a processor core. For example, some of the second processing instructions may conform to a microarchitecture used within the processor core, or may be configured to be read and / or executed by a specific processing resource within the processor core. The second processing instructions 120 may, for example, include opcodes which do not need to be decoded or translated in a processor core prior to execution.
[0180] FIG. 5C is a flow chart schematically illustrating a third embodiment S304-C of step S304 of Fig. 3. The embodiments S304-A and S304-C provide different but compatible functions, and they may be implemented separately or together. The embodiments S304-B and S304-C provide differentbut compatible functions, and they may be implemented separately or together. Furthermore, in some embodiments, all three of S304-A, S304-B and S304-C are implemented together.
[0181] At an optional initial step S521 , the processor interface unit 100, 400 receives a state management instruction and updates an interface state of the processor interface unit 100, 400. This interface state may be maintained by an interface state module 430 and used to inform subsequent instruction processing decisions.
[0182] At step S522 (which may follow step S521 , or which may be performed as a first part of step S304 in Fig. 3), the processor interface unit 100 determines an algorithm based on an algorithm prefix of a first processing instruction 110 and optionally based on the interface state. This step may involve the algorithm module 460 working in conjunction with the algorithm library 470 to retrieve an appropriate algorithm for generating second processing instructions 120.
[0183] At S523, the processor interface unit 100 generates the second processing instructions 120 based on the determined algorithm and an algorithm body of the first processing instructions 110. As previously described, an algorithmic instruction may comprise an algorithm prefix and the algorithm body in a single first processing instruction 110, or the algorithm prefix and algorithm body may be split between multiple first processing instructions 110. This generation process may utilize various components of the processor interface unit 100, such as the vector execution module 440 or the SIMD module 450, depending on the nature of the algorithm and the instruction.
[0184] FIG. 6 is a block diagram schematically illustrating a processor system with a partially-integrated configuration. The processor system includes the processor interface unit 100 and the processor core(s) 200.
[0185] Fig. 6 is similar to Fig. 1 , with some additional features, and like features are indicated with like reference numerals. The existing features of Fig. 1 will not be repeated unnecessarily.Additionally, the processor interface unit 100 may comprise some or all of the modules described with reference to Fig. 4.
[0186] As illustrated in Fig. 6, in some embodiments, some or all of the one or more processor core(s) 200 may include an interface connection 210.
[0187] The interface connection 210 may, for example, be a hardware modification which enables the processor interface unit 100 to read information from, and / or write information to, at least one of the memory resources and / or processing resources within the processor core. Such an interfaceconnection may be used to monitor a current processor core state for the one or more processor cores 200, which the processor interface unit 100 can use in its various methods as discussed above. For example, a monitored current processor core state may be stored in a processor core(s) state module 480 as discussed above.
[0188] Additionally or alternatively, some types of processor core 200 may include an interface for monitoring a processor core state, as a standard feature of their outputs, without requiring any modification - this may be used as the interface connection 210. In other words, some degree of monitoring a processor core state may be possible, even without any hardware integration between the processor interface unit 100, 400 and the one or more processor core(s) 200. For example, some Arm processors include a performance monitoring unit PMU which includes counters for certain event types which occur within the Arm core, and which may be used to at least partially understand a state of the Arm core.
[0189] The interface connection 210 may facilitate taking into account an internal state of the processor core(s) 200 when operating the processor interface unit 100. For example, the processor interface unit 100 may receive information about the current state of the processor core(s) 200 through the interface connection 210. This information may include details about resource utilization, execution status, or other relevant parameters of the processor core(s) 200.
[0190] By receiving this state information, the processor interface unit 100 can make informed decisions when generating the second processing instructions 120. For instance, the processor interface unit 100 may adjust the order or timing of instructions based on the current resource availability in the processor core(s) 200. This can lead to more efficient execution of instructions and improved overall system performance.
[0191] The interface connection 210 may support bidirectional communication. In other words, the interface connection 210 may also support write access to one or more internal resources (e.g. processing resources or memory resources) of a processor core. This may enable finer control of the processor core(s), and may enable more efficient execution of processing operations.
[0192] The configuration of Fig. 6 may correspond to a partial or full integration of the processor interface unit 100 and the processor core(s) 200, where at least one of the processor core(s) 200 is not a standard commercially available core and is modified for use together with a processor interface unit 100. For example, in embodiments where the processor interface unit 100 is partially or fullyintegrated with the processor core(s) 200, the processor interface unit 100 may comprise an integrated element integrated within an integrated processor core of the one or more processor cores. For example, the integrated element of the processor interface unit may be arranged on a same chip, or a same chiplet, as the integrated processor core. Alternatively, the configuration of Fig. 6 may be a non-integrated configuration, where the processor core(s) 200 are not modified specifically for use with a processor interface unit 100.
[0193] FIG. 7 is a block diagram schematically illustrating a partially or fully integrated processor system, in an embodiment where the processor interface unit 100 has write access to individual processing resources within a processor core 201 of the one or more processor cores 200.
[0194] As shown in Fig. 7, in an example scenario, the processor interface unit 100 controls the processor core 201 to perform a processor operation using a pipeline comprising multiple resources 220-1 to 220-n. Passed information (which may be unchanged, or may be computed or otherwise modified by each resource) is passed from resource to resource along the pipeline ultimately leading to an output 34 from the processor core 201. Each resource may, for example, comprise one or more processing resources (such as decoders, arithmetic logic units, floating point units, SIMD units or adders) and / or one or more memory resources (such as registers, caches and any other processor core memory)
[0195] At each stage of the pipeline, the processor interface unit 100 can provide an instruction 32 and / or data 33 to the respective resource 220-1 to 220-n. For example, if the resource 220-2 were a memory resource such as a register, the processor interface unit 100 could store relevant data 33 in the register.
[0196] Although the pipeline 220-1 to 220-n is linear in this example, processor cores often have branching and / or looping arrangements of resources, and the processor interface unit 100 may accordingly be configured to control any or all of the resources in the processor core 201.
[0197] Such an integrated scenario, where the processor interface unit 100 has access to resources within a processor core, enables further improvement of processing of instructions according to criteria such as execution time and resource usage.
[0198] For example, referring again to Figs. 4 and 5A, the processor interface unit 100 may use the re-ordering module 490 to optimize the pipeline of resources 220-1 to 220-n and a corresponding sequence of instructions and data provided to the resources. As another example, the direct access toindividual resources may be used together with the vector execution module 440 and the SIMD module 450 to process vector and SIMD operations efficiently within the resources of the one or more processor cores.
[0199] As mentioned above, one example application of the techniques described herein is to facilitate edge processing, using constrained processing resources, where efficiency is very important.
[0200] FIG. 8 is a block diagram schematically illustrating an example data processing sequence that may be performed in an edge processing scenario, using the processor system described herein. The sequence includes multiple stages for processing and analysing data from input to output.
[0201] The sequence begins with a stage 810, comprising digital signal processing applied to input received from sensors and data sources. The digital signal processing may involve manipulating sensor signals to generate meaningful data streams. Such digital signal processing may benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may utilize the vector execution module 440 and SIMD module 450 to efficiently process multiple data points in parallel during this stage.
[0202] Following the digital signal processing, the sequence moves to a stage 820, comprising data fusion from multiple inputs. Data fusion may involve combining data from different types and viewpoints to create a constant 'picture' of situational awareness for the edge system. Such fusion processes are commonly mathematically intense and often involve 3D transformations. Such data fusion may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may receive algorithmic instructions and employ the algorithm module 460 and algorithm library 470 to perform complex mathematical operations required for data fusion.
[0203] After data fusion, the sequence proceeds to a stage 830, comprising data analysis. The data analysis stage may involve looking for changes or anomalies in the created 'picture' and using mathematical algorithms to detect objects or features. Such data analysis may again benefit from various of the above-described techniques. For example, in an implementation, the reordering module 480 of the processor interface unit 100 may optimize the execution of analysis algorithms across the processor core(s) 200.
[0204] The sequence then moves to a stage 840, where decision making is performed. Such decision making may, for example, comprise using Al and machine learning techniques. This stagemay involve using algorithms to react appropriately and immediately to changing data patterns. Such decision-making may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may receive abstracted instructions 31 which efficiently represent complex Al and machine learning operations, thus reducing the total size of a program comprising instructions for the Al and machine learning operations, and reducing the data rate at which instructions need to be read into the processor system.
[0205] From stage 840, the sequence branches into two paths. The first path comprises outputting of system control signals. These control signals may, for example, be used to modify physical aspects of the system's behavior based on the decision-making results.
[0206] The second path flows through three additional processing stages for data storage and communication functions.
[0207] A stage 850 comprises data encoding to minimize size for storage or transmission. In many scenarios, a paper-trail of situations and responses must be stored and / or communicated with other elements of the system. The transmitted or stored data often needs to be compact, so encoding or compression algorithms are executed. Such data encoding may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may use specialized instructions and algorithms, re-ordering of instructions, and or knowledge of available resources at the one or more processor core(s), to efficiently compress data.
[0208] A stage 860 comprises data encryption for privacy and security. It is often desirable to prevent theft, interception or manipulation of data in the ecosystem by an adversary. Such threats can be mitigated using encryption to protect privacy and security, which often requires significant processing resources. Algorithms for creating keys, authenticating users and encrypting data are increasingly complex, and the data stream may need to be continuous and real-time. Such encryption may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may be configured to receive abstracted instructions corresponding to dedicated encryption algorithms stored in the algorithm library 470, and may be configured to generate second processing instructions according to those algorithms to secure the data.
[0209] A stage 870 comprises handling data communication with the ecosystem. This stage may involve preparing data for transmission using various communication protocols, which may for example involve error checking and correction. Such communication may again benefit from variousof the above-described techniques. For example, in an implementation, the processor interface unit 100 may be specifically configured with algorithms for generating signals according to the signal structure requirements of a communication protocol, especially where such structures are used repeatedly over the course of regular communication. Thus, a program for communication according to the communication protocol may contain fewer, smaller abstract instructions referring to the algorithms stored at the processor interface unit 100, and an overall data size of the program may be reduced.
[0210] The output from stage 870 provides data for storage or communication within the larger system ecosystem.
[0211] This example data processing sequence demonstrates how the described processor system may be used to efficiently handle complex, multi-stage data processing tasks. The processor interface unit 100 enables optimized execution of various algorithms and operations across the different stages, potentially improving overall system performance and efficiency. Furthermore, where multiple stages of a data processing pipeline are known at an early stage, a suitable processor interface unit 100 may be configured to efficiently perform multiple stages of the data processing pipeline by any combination of: using algorithms stored at the processor interface unit, using vector processing, using re-ordering of instructions to more efficiently utilize the associated processor core resources.
[0212] Fig. 9 is a block diagram schematically illustrating an example configuration for a re-ordering module 490.
[0213] Referring to Fig. 9, the re-ordering module may comprise a series of operational layers, wherein outputs from a preceding layer are passed as inputs to a next layer.
[0214] The operational layers may perform various functions, including: counting clock pulses, mixing, modifying or skipping values, and combining values.
[0215] In the example of Fig. 9, a first layer 910 comprises a series of counters COUNTERO, COUNTER1 and COUNTER2, each of which counts up to a respective maximum MAX0, MAX1 , MAX2 before overflowing and incrementing a next counter in the series (except for the last counter COUNTER2 which does not have an overflow destination). The first counter COUNTERO is incremented by a clock signal Clk (which may be generated in the processor interface unit, in a processor core 200, or elsewhere). Using three counters may often be useful, for example, when performing operations over a set of data points in a 3D grid.
[0216] A second layer 920 is configured to permute outputs from the first layer 910. In other words, the second layer may change an order of the counters, and change which counter is least / most significant. A third layer 930 is configured to invert outputs from the second layer 920. For example, the third layer may change a counting direction (decreasing vs increasing). A fourth layer 940 is configured to skip outputs from the third layer 930. For example, the fourth layer may be used to ignore a single counter, by setting it to zero. These are just examples, and such mixing, modifying and / or skipping can be used to produce a wide variety of mathematical sequences.
[0217] A fifth layer 950 is configured combine values from the fourth layer 940. More specifically, in this example, the fifth layer includes two multiply units and an adder, which combine the outputs from the fourth layer 940 to generate an output sequence Seq. In this example, the fifth layer calculates elements of the sequence based onSeqi = AS_Out0i + AS MaxO * AS_0utli + AS MaxO * AS_Maxl * AS_0ut2i
[0218] See the below examples for settings and generated sequences:Example MaxO Maxi Max2 Permute Skip Seq Order CounterA 3 2 2 0,2,1 1 0,0, 2, 2, 4, 4, 1 ,1, 3, 3, 5, 5 B 3 2 2 0,2,1 3 0,1, 0,1, 0,1, 2, 3, 2, 3, 2, 3 C 3 2 2 0,1,2 3 0,1, 2, 3, 4, 5, 0,1, 2, 3, 4, 5
[0219] In the above, the "Permute Order" column indicates the order in which counters are arranged by the permute layer, in increasing significance order.
[0220] The example sequences may, for example, be relevant for a matrix multiplication operation, where a 3x2 matrix is multiplied by a 2x2 matrix. Using a generated sequence, the matrices may be stored in a row-first order in a memory resource of a core 200 and accessed in the required by using the generated sequence to control a memory location offset before a corresponding second processing instruction 120 is decompressed
[0221] A series of operational layers, such as the example of Fig. 9, may be denoted as a "sequence generator". The re-ordering module may comprise one sequence generator, or a plurality of sequence generators. Some or all of the sequence generators may be implemented in hardware. Additionally,the re-ordering module 490 may comprise programmable re-ordering hardware such that different sequence generators can be configured. For example, when performing certain types of mathematical operation (and potentially working together with the algorithm module 460), the processor interface unit 400 may be configured to configure one or more sequence generators according to predetermined settings for the mathematical operation.
[0222] Fig. 10 is a block diagram schematically illustrating a processor system according to an embodiment.
[0223] In this embodiment, a processor system comprises a first processor subsystem 1010 including a management unit 1100 and an evaluation engine 1200. The management unit 1100 is configured to act as a processor interface unit. The evaluation engine 1200 comprises a plurality of managed processor elements.
[0224] The managed processor elements may, for example, comprise compute units. The compute units may be specialised or may be general-purpose. The managed processor elements may be controlled according to the techniques described herein for controlling one or more processor core(s).
[0225] Each managed processor element may be configured to perform mathematical operations and / or may comprise one or more arithmetic logic units. In such embodiments, the evaluation engine 1200 may be described as a maths function engine.
[0226] Additionally or alternatively, each managed processor element may comprise one or more interface elements suitable for communicating with other compute unit(s). For example, compute units may comprise one or more logic gates multiplexers or demultiplexers for sending or receiving data to or from other compute units.
[0227] The managed processor elements may be homogeneous. Alternatively, different managed processor elements may have different features and / or different arrangements.
[0228] In some examples, the managed processor units may be interconnected such that the plurality of managed processor elements performs operations as an interconnected fabric. In some examples, the managed processor units may be configured to perform operations as a systolic array. In some examples, the managed processor elements may be configured to operate as a field-programmable processing array.
[0229] As shown in Fig. 10, the management unit 1100 and the evaluation engine 1200 may be configured to exchange information in a bidirectional flow. The management unit 1100 is configured togenerate one or more second processing instructions and send the second processing instructions to the evaluation engine 1200.
[0230] The management unit 1100 is further configured to provide input data to the evaluation engine 1200, for use when evaluating the second processing instructions. The input data may be provided as part of the second processing instructions, or may be transferred by the management unit 1100 to the evaluation engine 1200 separately from second processing instructions, or may be actively fetched by the evaluation engine 1200.
[0231] The evaluation engine 1200 may be further configured to provide output data to the management unit 1100. The output data may comprise a result and / or a flag associated with evaluating one or more of the second processing instructions. For example, if the managed processor elements of the evaluation engine 1200 include one or more arithmetic logic units, the output data may comprise a result of a mathematical operation, and / or may comprise an overflow flag (a.k.a. a carry bit) indicating that an arithmetic logic operation overflowed an upper or lower limit of numbers which can be represented in the managed processor element
[0232] The example of Fig. 10 additionally illustrates how a first processor subsystem 1010 may be combined into a larger processor system. In the example of Fig. 10, the management unit 1100 is communicatively coupled to a communications bus 1030. The communications bus 1030 may be used to communicate between the first processor subsystem 1010 and a memory (such as a local flash memory, DRAM or SRAM, or an external memory such as an SSD or HDD) and / or between the first processor subsystem 1010 and one or more peripherals such as user interface devices and / or network communication devices.
[0233] The first processor subsystem 1010 is configured to receive instructions and data 1110 from the bus 1030 and to output processing results to the bus 1030. The instructions and data 1110 may comprise first processing instructions. The first processing instructions may include conventional instructions according to a standard instruction set e.g. RISC-V. In such embodiments, the management unit 1100 may control the evaluation engine 1200 to evaluate the conventional instructions with an efficiency equal or greater than execution of the conventional instructions using a conventional processor core. The first processing instructions may additionally or alternatively include abstracted instructions, such as high level instructions for selecting a stored algorithm at the management unit 1100 or for otherwise configuring behaviour of the management unit 1100.
[0234] Furthermore, the communications bus 1030 may be used to communicate between the first processor subsystem 1010 and one or more further processor subsystems 1020. Each of the further processor subsystems 1020 may be similar to the first processor subsystem, or may be different. For example, a second processor subsystem 1020 may comprise a coprocessor configured to cooperate with the first processor subsystem 1010 in order to complete instructions. In one example, a second processor subsystem comprises a conventional CPU, and the system is configured to obtain first processing instructions (e.g. by retrieving them from a memory, or receiving then from an external entity) and allocate each of the first processing instructions to the first processor subsystem 1010 or the second processor subsystem 1020.
[0235] In a processor system comprising multiple processor subsystems, two or more of the processor subsystems may share access to a shared memory (via direct connection or via a bus).
[0236] The system may use various factors in order to determine a processor subsystem to evaluate each first processor instruction. For example, the system may analyze one or more (and preferably a plurality) of upcoming first instructions in order to determine whether each first instruction is part of a repeated mathematical operation, such as matrix multiplication. If the first processor subsystem 1010 is provided with efficient hardware or an efficient algorithm for completing the mathematical operation, then the first instruction may be sent to the first processor subsystem for evaluation. On the other hand, if the first instruction is an isolated instruction which can be straightforwardly evaluated by a conventional CPU, then the first instruction may be passed to a second processor subsystem 1020.
[0237] In embodiments comprising multiple processor subsystems, the subsystems may be adapted to communicate with each other. For example, a result of a processing operation by one subsystem may be communicated to another subsystems as an input for a further processing operation.Additionally, subsystems may communicate synchronization signals and / or interrupts to each other in order to cooperatively perform a set of first processing instructions. In some embodiments, two or more of the processor subsystems may be coprocessors.
[0238] Fig. 11 is a flow chart schematically illustrating a method performed by management unit 1100, according to an embodiment.
[0239] At step S1101 , the management unit 1100 receives the first processing instructions 1110 associated with a processing operation. The first processing instructions 1110 may include a mixtureof abstracted instructions, standard instructions, data for processing, and control instructions for configuring the behaviour of the first processor subsystem 1010.
[0240] Abstracted instructions are instructions intended for the management unit 1100. Abstracted instructions may have different possible values, or a different instruction structure, from standard instructions, or they may have the same possible values and structure.
[0241] Various types of abstracted instruction are discussed above and are equally applicable to the embodiment of Fig. 10. These may, for example, include state management instructions for managing a state of the management unit 1100, instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions, tensor instructions or any other non-scalar or SIMD instruction format), and / or algorithm instructions indicating that the management unit should control the managed processor elements to perform a multi-instruction algorithm.
[0242] The management unit 1100 may be configured to receive instructions passively, at whatever rate they arrive at an input interface of the management unit 1100. Alternatively, the management unit 1100 may be configured to fetch instructions from a configured static or dynamic location.
[0243] At step S1102, the management unit 1100 generates the second processing instructions 1120 based on the received first processing instructions 1110. The second processing instructions typically comprise instructions for operations to be performed by one or more of the managed processor elements.
[0244] For example, when the first processing instructions comprise a standard instruction, generating second processing instructions may comprise breaking the standard instruction into a sequence of one or more special operations which can be performed by managed processor elements. Additionally, the processor interface unit may perform some management operations, such as selecting which of a plurality of managed processor elements should perform each second processing instruction and / or controlling how multiple of the managed processor elements should perform a plurality of second processing instructions.
[0245] In other examples, generating the second processing instructions may comprise any one or more of: identifying an algorithm for generating the second processing instructions; allocating resources at the managed processor elements; organising vector, matrix, tensor and / or any type of non-scalar or SIMD processing operations; or re-ordering a sequence of generated instructions - orgenerating a new sequence of instructions - in order to more-optimally use available resources at the one or more managed processor elements.
[0246] The management unit 1100 may additionally generate configurator data for configuring the evaluation engine 1200. As discussed further below, the configurator data may be used to configure behaviour of individual managed processor elements 1210 and / or interconnections between pairs of managed processor elements.
[0247] At step S1103, the management unit 1100 controls the managed processor elements to perform the processing operation based on the second processing instructions 1120 (and optionally based on the configurator data). In other examples, configurator data may be sent to the evaluation engine 1200 separately, in advance of sending second processing instructions 1120 and / or while processing operations are ongoing at the evaluation engine 1200.
[0248] Fig. 12 is a flow chart schematically illustrating a further method performed in a processor system, according to an embodiment.
[0249] The method begins at a step S1201 , where the system receives a first processing instruction 1110. The first processing instruction 1110 may comprise an abstracted instruction and / or a standard instruction. The first processing instruction 1110 also include configurator data for configuring behaviour of the processor interface unit.
[0250] Various sub-types of abstracted instruction are discussed below and these may, for example, include state management instructions for managing a state of the management unit 1100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions, tensor instructions or any other non-scalar or SIMD instruction format), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.
[0251] At a step S1202, the system determines an instruction type of the first processing instruction.
[0252] For example, the system may determine whether the first processing instruction comprises a standard instruction and / or an abstracted instruction. This may, for example, be determined based on a value or a structure of the first processing instruction 1110. For example, an abstracted instruction may be detected based on a value, a bit length or a type header of the first processing instruction 1110. Similarly, a standard instruction may be detected based on a value, a bit length or other structure of a standard instruction.
[0253] The system may, for example, store a set of properties associated with different instruction types. The system may determine whether the received first processing instruction 1110 matches the stored properties of any of the instruction types. For each instruction type, these stored properties may comprise a partial or complete instruction set architecture (ISA) specification.
[0254] If the first processing instruction is a standard instruction then, at step S1203, the system may use the bus 1030 to passes the instruction to a coprocessor 1020 without using the management unit 1100 or the associated evaluation engine 1200.
[0255] On the other hand, if the first processing instruction 1110 is an abstracted instruction or another instruction which is suitable for the first processor subsystem 1010, the method proceeds to step S1204, where the management unit 1100 generates one or more second processing instructions 1120. The second processing instructions typically comprise standard instructions or specialised instructions which can be performed by one or more managed processor elements of the evaluation engine 1200. In some embodiments, the second processing instructions may also comprise operand data, which is data to be provided to the evaluation engine 1200 outside of standard instructions. Additionally, the management unit 1100 may perform some management operations, such as selecting which one or more of a plurality of managed processor elements should help to perform the second processing instruction.
[0256] Following the generation of the second processing instructions 1120, the method moves to a step S1205, where the management unit 1100 passes the second processing instructions 1120 to the evaluation engine 1200.
[0257] The management unit 1100 can preferably pass through some instructions, as well as handling abstracted instructions. This enables the processor system to be used as a standard processor system, in a case where no specific abstracted instructions have been generated. In other words, the processor system can run compiled code which was compiled without being specifically intended for a processor system according to the invention, increasing flexibility. At the same time, the processor system can provide the described advantages - including reduced program size, and increased execution efficiency - for programs which include abstracted instructions.
[0258] Fig. 13 is a block diagram schematically illustrating further example embodiments of a processor system comprising a management unit 1100 and an evaluation engine 1200.
[0259] As shown in Fig. 13, the management unit 1100 may have a variety of modules. Different combinations of these modules may be included or omitted, depending on the requirements of a given application of the technology. The functions of the modules are independent, so each module can be included or omitted independently. The modules are only shown together in a single figure for brevity.
[0260] The management unit 1100 may comprise one or more memory modules 1135.
[0261] In some embodiments, the memory modules 1135 includes a memory controller 1115. The memory controller may be configured to store and / or obtain instructions and other data received or retrieved by the management unit 1100. In some embodiments, the memory controller 1115 may be a direct memory access controller which can directly receive first processing instructions 1110 and other data as discussed above, for example receiving instructions and / or data from a communication bus 1030. The memory controller 1115 may be further configured to transfer second instructions and / or data to the evaluation engine 1200 as and when required.
[0262] In some embodiments, the memory modules 1135 include a data memory 1125. The data memory may serve a range of purposes including storing first processing instructions and data received or retrieved by the management unit 1100, storing second processing instructions and other data generated by the management unit 1100, and storing one or more results output from the evaluation engine 1200 during or after evaluation of the second processing instructions. The data memory 1125 is preferably a relatively fast memory such as SRAM or, more preferably, a cache memory.
[0263] In some embodiments, the memory modules 1135 include an algorithm library module 1170. The algorithm library module 1170 may be configured to store one or more predetermined algorithms which can be performed using the evaluation engine 1200. As examples, the predetermined algorithms may include frequently used mathematical operations such as Fast Fourier Transform, Discrete Cosine Transform, matrix operations, encoding operations or encryption operations, or ML model operations performed while training or operating an ML model. The predetermined algorithms need not be basic elementary algorithms, and can be complex algorithms. The predetermined algorithms may, in some cases, be tailored to individual customers, depending on which algorithms are frequently needed for their use case.
[0264] Each predetermined algorithm may simply comprise reading a predetermined sequence of second processing instructions from a memory, and adding any required operands. Alternatively, thepre-defined algorithm may comprise rules for generating the second processing instructions, without the sequence of second processing instructions being fully defined in advance. This may, for example, be necessary for more complex branching algorithms, or for feedback algorithms (such as iterative algorithms) where the total length of the algorithm may not be known in advance.
[0265] In addition to predetermined algorithms, the algorithm library module 1170 may support dynamically updated algorithms. For example the algorithm library module 1170 may be configured to receive firmware updates as part of first processing instructions or data which are received by or retrieved by the management module 1100.
[0266] Additionally or alternatively, the algorithm library module may be configured to optimise a predetermined algorithm over time based on monitored performance when the algorithm is used in the processor system. Such optimisation may comprise evaluation of candidate modified versions of the predetermined algorithm, in order to determine whether any of the modified versions provide improved performance. Candidate modified versions of the predetermined algorithm may be generated randomly, or may be generated using a trained machine learning model. Using a machine learning model may have the benefit of attempting modifications in parts of the algorithm which are likely to be improvable (such as loops which increase memory usage over time), while reducing the risk of modifications which are likely to break the algorithm.
[0267] When controlling the one or more managed processor elements 1210 to perform second processing instructions corresponding to a predetermined algorithm, the management unit 1100 may use a local memory, such as the data memory 1125, in order to perform the complete algorithm with minimal instances of storing data in an external memory or retrieving data from an external memory, thus substantially eliminating a slow operation type, and increasing the processing speed and efficiency for the predetermined algorithm. Furthermore, the management module 1100 may use the memory controller 1115 in order to anticipate any fetch operations so that data is retrieved before it is required, and the evaluation engine 1200 does not need to wait for fetch operations any more than necessary.
[0268] For example, a manufacturer of a management unit 1100 or a processor subsystem 1010 or a processor system according to this application may supply, to customers, an SDK (Software Development Kit) for the management unit. This SDK may include the available algorithms which are stored by default in the algorithm library module 1170, and which can be efficiently incorporated intoprograms executed on the processor system. The algorithm library module 1170 (and the corresponding SDK) may, in some cases, be tailored to individual customers, depending on which algorithms are frequently needed for their use case. In some embodiments, the management unit 1100 may additionally support adding algorithms to the algorithm library module 1170 and / or removing algorithms from the algorithm library module 1170. For example, after a management unit 1100 has been manufactured (such as when the management unit 1100 is installed in a specific computer system) a user may wish to add, modify or remove algorithms, in order to customise the processor subsystem 1010 to be maximally beneficial for their specific use case. In other words, while the application refers to "predetermined" algorithms, these algorithms may be "predetermined" at any stage prior to an instance of executing first processing instructions 1110. For example, the processor system may be used to perform operations corresponding to a first set of first processing instructions 1110 with a first set of predetermined algorithms in the algorithm library module 1170, then modified to replace the first set of predetermined algorithms with a second set of predetermined algorithms in the algorithm library module 1170, and then used again to perform operations corresponding to a second set of first processing instructions 1110 with the second set of predetermined algorithms in the algorithm library module 1170. The system may be configured to support changes to the algorithms stored in the algorithm library module 1170 as a regular, generally available operation. Alternatively, system may be configured to support changes to the algorithms stored in the algorithm library module 1170 in the form of a firmware update. As mentioned above, the predetermined algorithms need not be elementary algorithms, and may be complex - nevertheless, each predetermined algorithm may advantageously be identified using a single algorithmic instruction.
[0269] The management unit 1100 may comprise one or more optimization modules 1165.
[0270] In some embodiments, the optimization modules 1165 include a vector execution module 1140 configured to manage vector processing in the evaluation engine 1200. Herein, it should be understood that vector processing comprises performing a same instruction yt=multiple times, for each of multiple operand values (a "vector" of operand values x), in order to produce multiple result values (a "vector" of output values y. The vector processing unit is configured to manage execution resources for vector processing such that, when sufficient resources are available at the one or more managed processor elements 1210, multiple copies of the instruction for different operands are executed simultaneously using different processing resources. This managed vectorprocessing operation can be contrasted with a for-loop methodology found in known vector processing systems, in which a loop is executed for successive executions of the instruction using a same processing resource, meaning that the processing time scales with the number of operand values in the "vector", and the vector elements of the output from the instruction become available sequentially. Furthermore, even when parallel execution is not possible for all elements of the "vector" of operand values, the vector execution module 1140 may be configured to start a downstream operation on elements of the vector result z= even before the vector result has been fully calculated, if the management unit 1100 has received first processing instructions corresponding to such a sequence of operations (for example in the case of an algorithmic instruction). Additionally or alternatively, the management unit 1100 may comprise a matrix execution module or a tensor execution module configured for performing a same instruction for each of a multi-dimensional collection of operand values.
[0271] In some embodiments, the management unit 1100 comprises a SIMD module configured to control SIMD (single-instruction multiple-data) execution resources in the one or more managed processor elements 1210. Operations of the SIMD module may overlap with, or be merged with, operations of a vector execution module or a multi-dimensional execution module as described above. The SIMD module makes use of SIMD execution resources in the one or more managed processor elements 1210 which support parallel execution for multiple operands, on a hardware level. For example, the SIMD module may rearrange a data sequence according to a configuration of the one or more managed processor elements 1210. More specifically, the SIMD module may rearrange a data sequence to fit available SIMD resources, according to their respective capacities for multiple data in each execution step. Additionally, in a scenario where a data length (such as a data word length) is shorter than a corresponding register length in an available SIMD resource, the SIMD module may adjust alignment of data in registers to improve execution order and efficiency for execution across a broader set of operand data.
[0272] In some embodiments, the optimization modules 1165 include an ordering module 1190 configured to arrange or re-arrange instructions and operations. For example, the ordering module may be configured to apply a set of sequence generator rules for re-arranging instructions and operations. Additionally or alternatively, the ordering module may comprise one or more sequence generators, an example of which is described above with reference to Fig. 9.
[0273] Such ordering of instructions is useful in a wide variety of cases, to improve execution speed and / or improve efficient usage of processing resources. In particular, in many cases, it is useful for the ordering module 1190 to order a sequence of second processing instructions 1120 based on a configuration of the one or more managed processor elements 1210. Additionally, such ordering may ensure that operands can be accessed using a linear order of data access operations, reducing or avoiding the need for branched data access when second processing instructions 1120 are executed in the managed processor element(s) 1120.
[0274] As a first case, the ordering module 1190 may be configured to order processing of a sequence of second processing instructions 1120 depending on a predicted availability of resources in the one or more managed processor element(s) 1120.
[0275] Based on actual and / or predicted knowledge of resource usage in the one or more managed processor element(s) 1120, the management unit may try to optimise an order of execution for the second processing instructions 1120 at the one or more managed processor element(s) 1120.
[0276] As a second case, the ordering module 1190 may be configured to determine a linear or nonlinear sequence for loading operand data into registers. In embodiments where the processor interface unit has write-access to resources within the evaluation engine 1200, the ordering module may determine a sequence for loading operand data into registers within the evaluation engine or for loading data into the data memory 1125.
[0277] As a third case, the ordering module 1190 may be configured to calculate register offsets or other memory locations for storing and / or retrieving operand data, to enable a computed execution on the operands in a linear or non-linear order, without requiring any rearrangement of data in memory.
[0278] As a fourth case, the ordering module 1190 may be configured to calculate register offsets or other memory locations for storing and / or retrieving output data following a calculation at the one or more managed processor elements 1210. Such output destinations may be configured using additional second processing instructions 1120 or operands within the second processing instructions.
[0279] As a fifth case, the ordering module may be configured to order operands in registers for efficient execution of SIMD operations, whether as part of vector processing or in other contexts.
[0280] The management unit 1100 may comprise one or more higher level modules for controlling the previously described operations and modules.
[0281] For example, higher level modules may include a configuration module 1180.
[0282] The configuration module 1180 may, for example, comprise a processor core such as a RISC-V processor core.
[0283] In some embodiments, the configuration module 1180 is configured to maintain information about the state of the managed processor elements 1210. This information may, for example, be used by the ordering module 1190 to make informed decisions about instruction ordering or may be used in order to make decisions about when to send second processing instructions 1120 from the management unit 1100 to the evaluation engine 1200. Additionally, the configuration module 1180 may store a history of the state(s) of the managed processor elements 1210. For example, the history may comprise, for each of a plurality of time steps, associated data of a state of the managed processor element(s) and data a state of the management unit 1100. By using such managed processor element state information (and optionally using historical state information), the management unit 1100 can help to reduce inefficiencies such as waiting for busy processing resources, or errors or delays due to simultaneous attempts to access a same memory resource, which can in turn lead to stall events, wasted clock cycles or wasted power.
[0284] In embodiments where the management unit 100 has read-access and / or write-access to resources within the one or more managed processor elements, the management unit may directly monitor availability of resources in the one or more managed processor elements.
[0285] Alternatively, the configuration module 1180 may predict usage of resources within the one or more managed processor elements, based on the configuration of the one or more managed processor elements.
[0286] The managed processor elements state may, for example, include any of: a configuration of the one or more managed processor elements 1210; a monitored current state of the one or more managed processor elements; a predicted current state of the one or more managed processor elements; data or statistics regarding historical performance of the one or more managed processor elements.
[0287] The configuration module 1180 may be further adapted to configure individual managed processor elements 1210 and / or interconnections between managed processor elements within the evaluation engine 1200. As mentioned above, each managed processor element may comprise one or more specialised or general-purpose compute units and / or one or more interface elements (such as logic gates, multiplexers or demultiplexers) suitable for communicating with other managedprocessor elements in the evaluation engine 1200. The configuration module 1180 may be configured to configure a compute behaviour of each managed processor element and / or interconnections between pairs of managed processor elements. Such configuration may be performed on a perelement basis, or may be performed on a grouped or collective basis. For example, the configuration module 1180 may configure operations and interconnections within the evaluation engine 1200 such that a group of managed processor elements 1210 performs operations as a systolic array. More generally, the configuration module 1180 may be configured to configure the evaluation engine as a field-programmable processing array (i.e. a configure the managed processor elements 1210 as an interconnected network, where each node of the network may comprise hardware for general or specialised computation).
[0288] Some of these configuration operations for configuring the evaluation engine 1200 may be externally controlled, for example when the management unit 1100 receives first processing instructions 1110 or data comprising the previously described configuration data illustrated in Fig. 11. Additionally or alternatively, the configuration operations may be performed by sending configurator data to the evaluation engine 1200 (together with, or separately from, second processing instructions 1120). The evaluation engine 1200 may include an interface for receiving and enacting configurator data, and / or the management unit 1100 may directly configure the affected elements and / or interconnections.
[0289] The management unit 1100 may, for example, use any of the managed processor elements state information when controlling the memory controller 1115, the vector execution module 1140, the algorithm library module 1170, or the ordering module 1190.
[0290] The configuration module 1180 may additionally or alternatively be configured to perform management operations for one or more functions or modules of the management unit 1100. Some management operations may be externally controlled, for example when the management unit 1100 receives first processing instructions 1110 or data comprising a management code. Such external management information may include the previously described configuration data illustrated in Fig. 11.
[0291] Management operations may be used to configure various settings and behaviours of the management unit 1100, and data indicating settings may be stored in a local memory such as the data memory 1125. Examples of management operations include:
[0292] -Step
[0293] A Step operation which may be used to increment or decrement a memory offset. For example, the memory offset may indicate a location in a memory resource (such as a register) of a managed processor element 1210. More specifically, the memory offset may indicate a register file offset. The management unit 1100 may additionally or alternatively be configured to automatically increment a memory offset. However, manually controlling a memory offset may be useful in some scenarios, such as processing multiple first processing instructions 1110 without incrementing a register file offset.
[0294] - Set Vector Length
[0295] A Set Vector Length operation may be used to set a vector length for vector processing handled by the vector execution module 1140. In other words, the vector execution module 1140 may be configured to read a vector length setting and, after receiving a vector processing instruction and a subsequent instruction in the first processing instructions 1110, generate a number of second processing instructions 1120 based on the vector length setting and based on the subsequent instruction. For example, if the subsequent instruction is a standard instruction, the vector execution module 1140 may generate (vector length) copies of the standard instruction.
[0296] A Set Vector Length operation may also be used to set a vector operating mode, such as setting an "ignore vector processing instructions" mode, a mode in which a register file offset is automatically incremented for each element of a vector of operations, or a mode in which a register file offset is not incremented for each of (vector length) operations. That last mode may be used to "copy" a standard instruction for each operation, by using the same register location multiple times, without actually generating copies of the standard instruction.
[0297] - Store Algorithm
[0298] A Store Algorithm operation may be used to indicate that a set of first processing instructions 1110 should be stored in the algorithm library module 1170. For example, this may be used to add, remove, or update the predetermined algorithms stored in the algorithm library module 1170.
[0299] - Configure Ordering
[0300] A Configure Ordering operation may be used to configure the ordering module 1190. For example, this may be used to set sequence generator rules for arranging or re-arranging instructions and operations.
[0301] The configuration module 1180 may be further configured to maintain state information for the management unit 1100.
[0302] This state information may, for example, include any of: information about ongoing instructions which are being executed in the evaluation engine 1200; information about planned instructions which are due to be executed in the evaluation engine 1200; information about iterative processes in which an output from the evaluation engine 1200 may be used when generating a further second processing instruction 1120; information about processes which require the fetching of additional data; information about an algorithm which is currently being executed or which is due to be executed.
[0303] The management unit 1100 may, for example, use any of this state information when controlling the memory controller 1115, the vector execution module 1140, the algorithm library module 1170, or the ordering module 1190.
[0304] In some embodiments, the management unit 1100 also includes a memory management module and / or a garbage collection module (not shown), in order to manage any data that it stores. For example, a memory management module may allocate memory space for a matrix calculation, until a result is ready to be output from the processor system. As another example, data from intermediate steps of an algorithm, which may be stored while the management unit 1100 is controlling processing of an algorithm, is preferably cleared up as soon as it is no longer required for the algorithm.
[0305] In some embodiments, the management unit 1100 also includes a pointer alignment module and / or a programme counter module. Misalignment due to a branch instruction or a jump instruction can sometimes require re-alignment of pointers and memory address offsets, or flushing of data pipelines. Such misalignments may be detected or predicted based on a state of the managed processor elements 1210 or based on the second processing instructions 1120. By quickly detecting, or even predicting, misalignments, some interruptions and latency can be avoided, improving overall performance.
[0306] The subject-matter of the application also includes the following numbered feature combinations.1. A processor system comprising one or more managed processor elements and a processor interface unit, wherein the processor interface unit is configured to:receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions; andcontrol the one or more managed processor elements to perform the processing operation based on the one or more second processing instructions.Managed element(s) structure and control2. A processor system according to feature combination 1 , wherein each of the one or more managed processor elements is a compute unit configured to perform mathematical operations. 3. A processor system according to feature combination 1 or feature combination 2, wherein the one or more managed processor elements is a plurality of managed processor elements.4. A processor system according to feature combination 3, wherein the processor interface unit is configured to further configure how the processing operation will be distributed across two or more of the managed processor elements.5. A processor system according to feature combination 3 or feature combination 4, wherein the processor interface unit is further configured to control one or more interconnections between pairs of the plurality of managed processor elements.6. A processor system according to feature combination 5, wherein the processor interface unit is configured to control a plurality of interconnections between pairs of the plurality of managed processor elements, such that the plurality of managed processor elements performs operations as an interconnected fabric.7. A processor system according to feature combination 6, wherein the processor interface unit is configured to control the plurality of managed processor elements and the plurality of interconnections between pairs of the plurality of managed processor elements such that the plurality of managed processor operations performs operations as a systolic array.8. A processor system according to any of feature combinations 4 to 7, wherein the processor interface unit is configured to control the plurality of managed processor elements to operate as a field-programmable processing array.Coprocessor9. A processor system according to any of feature combinations 1 to 8, wherein the processor interface unit and the one or more managed processor elements are together configured to operate as a coprocessor in cooperation with an unmanaged processor element.10. A processor system according to feature combination 9, wherein the unmanaged processor element comprises one or more processor core(s).11 . A processor system according to feature combination 9 or feature combination 10, wherein:the processor interface unit and the one or more managed processor elements are connected to a bus; and / orthe processor interface unit and the one or more managed processor elements are configured to access a shared memory.Different Instruction Sets / Different Ordering12. A processor system according to any of feature combinations 1 to 11 , wherein an order of the plurality of second processing instructions is based on a configuration of the one or more managed processor elements.13. A processor system according to feature combination 12, wherein the configuration of the one or more managed processor elements comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in or between the one or more managed processor elements.14. A processor system according to feature combination 13, wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more managed processor elements; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more managed processor elements to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more managed processor elements.15. A processor system according to feature combination 14, wherein the processor interface unit is configured to: iteratively repeat the steps of determining a flow of processing sub-operations and allocating resources, and evaluate an expected performance characteristic for the one or more managed processor elements performing the processing operation using the flow of processing suboperations and the allocated resources, in order to determine an improved flow of processing suboperations and allocation of resources.16. A processor system according to feature combination 13, wherein the processor interface unit is configured to use a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.17. A processor system according to any of feature combinations 12 to 16, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more managed processor elements.18. A processor system according to any of feature combinations 1 to 17, wherein the one or more second processing instructions comprises at least one of the first processing instructions.18a. A processor system according to feature combination 18, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.19. A processor system according to feature combination 18a, wherein the at least two processing instructions comprise a memory read instruction.20. A processor system according to any of feature combinations 1 to 19, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more managed processor elements or a microarchitecture of the one or more managed processor elements, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.Exclusive control21. A processor system according to any of feature combinations 1 to 20, wherein the processor interface is configured to operate with exclusive control of the one or more managed processor elements.22. A processor system according to feature combination 21 , wherein the one or more managed processor elements comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.23. A processor system according to feature combination 21 , wherein the exclusive control of the one or more managed processor elements comprises disconnecting a dynamically switchable connection between the one or more managed processor elements and another component of the system.Hardware arrangements24. A processor system according to any of feature combinations 1 to 23, wherein the processor interface unit comprises a non-integrated element separate from the one or more managed processor elements.25. A processor system according to feature combination 24, comprising a plurality of chips, wherein the non-integrated element of the processor interface unit is arranged in a first chip, and the one or more managed processor elements are arranged in one or more other chips separate from the first chip.26. A processor system according to feature combination 24, comprising a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the one or more managed processor elements are arranged in one or more chiplets separate from the first chiplet.27. A processor system according to any of feature combinations 1 to 26, wherein the processor interface unit comprises an integrated element integrated within a managed processor element of the one or more managed processor elements.28. A processor system according to feature combination 27, wherein the integrated element of the processor interface unit and the one or more managed processor elements are arranged in a chip.29. A processor system according to feature combination 27, wherein the integrated element of the processor interface unit and the one or more managed processor elements are arranged in a chiplet.Instruction transforming30. A processor system according to any of feature combinations 1 to 29, wherein the processor interface unit is configured to:receive one or more first instructions comprising configuration data and data to be processed; determine an algorithm for generating one or more second instructions based on the configuration data; andgenerate one or more second processing instructions for the managed processor elements to process the data based on the algorithm.31. A processor system according to feature combination 30, wherein the processor interface unit is configured to maintain an internal state, and the algorithm is determined based on the configuration data and the internal state.32. A processor system according to feature combination 31 , wherein the processor interface unit is further configured to receive a state management instruction for managing the internal state.33. A processor system according to any of feature combinations 30 to 32, wherein the one or more second processing instructions comprise a memory read instruction.34. A processor system according to any of feature combinations 30 to 33, wherein the algorithm is a vector processing algorithm, a matrix processing algorithm, a tensor processing algorithm or a non-scalar processing algorithm.35. A processor system according to any of feature combinations 30 to 34, wherein the processor interface unit comprises a memory storing one or more algorithms, and the algorithm for generating the one or more second instructions is determined from the stored one or more algorithms.36. A processor system according to any of feature combinations 30 to 35, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.37. A processor system according to feature combination 35 or feature combination 36, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.38. A processor system according to any of feature combinations 35 to 37, wherein the processor interface unit is configured to dynamically update a stored algorithm or a stored sequence of instructions.39. A processor system according to feature combination 38, wherein the processor interface unit comprises a machine learning model for dynamically updating the stored algorithm or the stored sequence of instructions.Management40. A processor system according to any of feature combinations 1 to 39, wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.Integrated functionality41. A processor system according to any of feature combinations 1 to 40, wherein the one or more managed processor elements comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.42. A processor system according to feature combination 41 , wherein the one or more second processing instructions comprise a machine code instruction, and the processor interface unit is configured to input the machine code instruction directly to a register or functional unit of the one or more managed processor elements.43. A processor system according to feature combination 42, wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the processor interface unit is configured to modify or rearrange data in a register of the one or more managed processor elements while the one or more cores execute the SIMD instruction.44. A processor system according to feature any of feature combinations 41 to 43, wherein the one or more registers or functional units comprise a parallel execution unit.45. A processor system according to feature combination 30 and according to feature combination 41 , wherein the determined algorithm comprises:monitoring a state of the one or more managed processor elements to obtain a current state; andgenerating a second processing instruction based on the current state, or re-ordering at least two second processing instructions based on the current state.46. A processor system according to feature combination 45, wherein the current state comprises a resource utilization.47. A processor system according to feature combination 35 and according to feature combination 41 , wherein the processor interface unit is configured to:monitor a state of the one or more managed processor elements over time to obtain a state history; anddynamically update the stored algorithm or the stored sequence of instructions based on the state history.48. A processor system according to feature combination 47, wherein the state history comprises a resource utilization history.49. A method performed by a processor interface unit, in a processor system comprising one or more managed processor elements and the processor interface unit, the method comprising:receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions; andcontrolling the one or more managed processor elements to perform the processing operation based on the one or more second processing instructions.50. A method according to feature combination 49, wherein each of the one or more managed processor elements is a compute unit configured to perform mathematical operations.51 . A method according to feature combination 49 or feature combination 50, wherein the one or more managed processor elements is a plurality of managed processor elements.52. A method according to feature combination 51 , further comprising configuring how the processing operation will be distributed across two or more of the managed processor elements. 53. A method according to feature combination 51 or feature combination 52, further comprising controlling one or more interconnections between pairs of the plurality of managed processor elements.54. A method according to feature combination 53, comprising controlling a plurality of interconnections between pairs of the plurality of managed processor elements, such that the plurality of managed processor elements performs operations as an interconnected fabric.55. A method according to feature combination 54, comprising controlling the plurality of managed processor elements and the plurality of interconnections between pairs of the plurality of managed processor elements such that the plurality of managed processor operations performs operations as a systolic array.56. A method according to any of feature combinations 52 to 55, comprising controlling the plurality of managed processor elements to operate as a field-programmable processing array.57. A method according to any of feature combinations 49 to 56, wherein the processor interface unit and the one or more managed processor elements are together configured to operate as a coprocessor in cooperation with an unmanaged processor element.58. A method according to feature combination 57, wherein the unmanaged processor element comprises one or more processor core(s).59. A method according to feature combination 57 or feature combination 58, wherein:the processor interface unit and the one or more managed processor elements are connected to a bus; and / orthe processor interface unit and the one or more managed processor elements are configured to access a shared memory.60. A method according to any of feature combinations 49 to 59, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more managed processor elements.61 . A method according to feature combination 60, wherein the configuration of the one or more managed processor elements comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in or between the one or more managed processor elements.62. A method according to feature combination 61 , wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more managed processor elements; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the oneor more managed processor elements to perform each sub-operation of the flow of processing suboperations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more managed processor elements.63. A method according to feature combination 62, comprising: iteratively repeating the steps of determining a flow of processing sub-operations and allocating resources, and evaluating an expected performance characteristic for the one or more managed processor elements performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.64. A method according to feature combination 62, comprising using a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.65. A method according to any of feature combinations 61 to 64, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more managed processor elements .66. A method according to any of feature combinations 49 to 65, wherein the one or more second processing instructions comprises at least one of the first processing instructions.67. A method according to feature combination 66, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.68. A method according to feature combination 67, wherein the at least two processing instructions comprise a memory read instruction.69. A method according to any of feature combinations 49 to 68, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more managed processor elements or a microarchitecture of the one or more managed processor elements, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.70. A method according to any of feature combinations 49 to 69, wherein the processor interface is configured to operate with exclusive control of the one or more managed processor elements. 71. A method according to feature combination 70, wherein the one or more managed processor elements comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.72. A method according to feature combination 70, wherein the exclusive control of the one or more managed processor elements comprises disconnecting a dynamically switchable connection between the one or more managed processor elements and another component of the system. 73. A method according to any of feature combinations 49 to 72, wherein the processor interface unit comprises a non-integrated element separate from the one or more managed processor elements.74. A method according to feature combination 73, wherein the processor system comprises a plurality of chips, wherein the non-integrated element of the processor interface unit is arranged in a first chip, and the one or more managed processor elements are arranged in one or more other chips separate from the first chip.75. A method according to feature combination 73, wherein the processor system comprises a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the one or more managed processor elements are arranged in one or more chiplets separate from the first chiplet.76. A method according to any of feature combinations 49 to 75, wherein the processor interface unit comprises an integrated element integrated within a managed processor element of the one or more managed processor elements.77. A method according to feature combination 76, wherein the integrated element of the processor interface unit and the one or more managed processor elements is arranged in a chip. 78. A method according to feature combination 76, wherein the integrated element of the processor interface unit and the one or more managed processor elements is arranged in a chiplet.79. A method according to any of feature combinations 49 to 78, comprising:receiving one or more first instructions comprising configuration data and data to be processed;determining an algorithm for generating one or more second instructions based on the configuration data; andgenerating one or more second processing instructions for the managed processor elements to process the data based on the algorithm.80. A method according to feature combination 79, further comprising maintaining an internal state of the processor interface unit, and determining the algorithm based on the configuration data and the internal state.81. A method according to feature combination 80, further comprising receiving a state management instruction for managing the internal state.82. A method according to any of feature combinations 79 to 81 , wherein the one or more second processing instructions comprise a memory read instruction.83. A method according to any of feature combinations 79 to 82, wherein the algorithm is a vector processing algorithm, a matrix processing algorithm, a tensor processing algorithm or a non-scalar processing algorithm.84. A method according to any of feature combinations 79 to 83, wherein the processor interface unit comprises a memory storing one or more algorithms, and method comprises determining the algorithm for generating the one or more second instructions from the stored one or more algorithms.85. A method according to any of feature combinations 79 to 84, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.86. A method according to feature combination 84 or feature combination 85, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.87. A method according to any of feature combinations 84 to 86, further comprising dynamically updating a stored algorithm or a stored sequence of instructions.88. A method according to feature combination 87, further comprising using a machine learning model to dynamically update the stored algorithm or the stored sequence of instructions.89. A method according to any of feature combinations 49 to 88, wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.90. A method according to any of feature combinations 49 to 89, wherein the one or more managed processor elements comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.91. A method according to feature combination 90, wherein the one or more second processing instructions comprise a machine code instruction, and the method comprises inputting the machine code instruction directly to a register or functional unit of the one or more managed processor elements.92. A method according to feature combination 91 , wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the method comprises modifying or rearranging data in a register of the one or more managed processor elements while the one or more managed processor elements execute the SIMD instruction.93. A method according to feature any of feature combinations 90 to 92, wherein the one or more registers or functional units comprise a parallel execution unit.94. A method according to feature combination 79 and according to feature combination 90, wherein the determined algorithm comprises:monitoring a state of the one or more managed processor elements to obtain a current state; andgenerating a second processing instruction based on the current state, or re-ordering at least two second processing instructions based on the current state.95. A method according to feature combination 94, wherein the current state comprises a resource utilization.96. A method according to feature combination 84 and according to feature combination 90, further comprising:monitoring a state of the one or more managed processor elements over time to obtain a state history; anddynamically updating the stored algorithm or the stored sequence of instructions based on the state history.97. A method according to feature combination 96, wherein the state history comprises a resource utilization history.98 A computer program comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of feature combinations 49 to 97.99. A non-transitory storage medium storing instructions which, when executed by one or more processors, cause the processors to perform a method according to any of feature combinations 49 to 97.100. A data signal comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of feature combinations 49 to 97.
[0307] Features of any of the examples or embodiments outlined above may be combined to create additional examples or embodiments without losing the intended effect. It should be understood that the description of an embodiment or example provided above is by way of example only, and various modifications could be made by one skilled in the art. Furthermore, one skilled in the art will recognise that numerous further modifications and combinations of various aspects are possible. Accordingly, the described aspects are intended to encompass all such alterations, modifications, and variations that fall within the scope of the appended claims.
Claims
CLAIMS1. A processor system comprising a plurality of managed processor elements and a processor interface unit, wherein the processor interface unit is configured to:receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions;configure how the processing operation will be distributed across two or more of the managed processor elements; andcontrol the two or more of the managed processor elements to perform the processing operation based on the one or more second processing instructions.
2. A processor system according to claim 1 , wherein each of the plurality of managed processor elements is a compute unit configured to perform mathematical operations.
3. A processor system according to claim 1 or claim 2, wherein the processor interface unit is further configured to control one or more interconnections between pairs of the plurality of managed processor elements.
4. A processor system according to claim 3, wherein the processor interface unit is configured to control a plurality of interconnections between pairs of the plurality of managed processor elements, such that the plurality of managed processor elements performs operations as an interconnected fabric.
5. A processor system according to claim 4, wherein the processor interface unit is configured to control the plurality of managed processor elements and the plurality of interconnections between pairs of the plurality of managed processor elements such that the plurality of managed processor operations performs operations as a systolic array.
6. A processor system according to any of claims 3 to 5, wherein the processor interface unit is configured to control the plurality of managed processor elements to operate as a field-programmable processing array.
7. A processor system according to any of claims 1 to 6, wherein the processor interface unit and the plurality of managed processor elements are together configured to operate as a coprocessor in cooperation with an unmanaged processor element.
688. A processor system according to claim 7, wherein the unmanaged processor element comprises one or more processor core(s).
9. A processor system according to claim 7 or claim 8, wherein:the processor interface unit and the plurality of managed processor elements are connected to a bus; and / orthe processor interface unit and the plurality of managed processor elements are configured to access a shared memory.
10. A processor system according to any of claims 1 to 9, wherein an order of the plurality of second processing instructions is based on a configuration of the plurality of managed processor elements.
11. A processor system according to claim 10, wherein the configuration of the plurality of managed processor elements comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in or between the plurality of managed processor elements.
12. A processor system according to claim 11 , wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more managed processor elements; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more managed processor elements to perform each sub-operation of the flow of processing suboperations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the plurality of managed processor elements.
13. A processor system according to claim 12, wherein the processor interface unit is configured to: iteratively repeat the steps of determining a flow of processing sub-operations and allocating resources, and evaluate an expected performance characteristic for the one or more managed processor elements performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.6914. A processor system according to claim 11 , wherein the processor interface unit is configured to use a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.
15. A processor system according to any of claims 10 to 14, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the plurality of managed processor elements .
16. A processor system according to any of claims 1 to 14, wherein the one or more second processing instructions comprises at least one of the first processing instructions.
17. A processor system according to claim 16, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.
18. A processor system according to claim 17, wherein the at least two processing instructions comprise a memory read instruction.
19. A processor system according to any of claims 1 to 18, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the plurality of managed processor elements or a microarchitecture of the plurality of managed processor elements , andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.
20. A processor system according to any of claims 1 to 19, wherein the processor interface is configured to operate with exclusive control of the plurality of managed processor elements .
21. A processor system according to claim 20, wherein the plurality of managed processor elements comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.7022. A processor system according to claim 20, wherein the exclusive control of the plurality of managed processor elements comprises disconnecting a dynamically switchable connection between the plurality of managed processor elements and another component of the system.
23. A processor system according to any of claims 1 to 22, wherein the processor interface unit comprises a non-integrated element separate from the plurality of managed processor elements.
24. A processor system according to claim 23, comprising a plurality of chips, wherein the nonintegrated element of the processor interface unit is arranged in a first chip, and the plurality of managed processor elements are arranged in one or more other chips separate from the first chip.
25. A processor system according to claim 23, comprising a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the plurality of managed processor elements are arranged in one or more chiplets separate from the first chiplet.
26. A processor system according to any of claims 1 to 25, wherein the processor interface unit comprises an integrated element integrated within a managed processor element of the plurality of managed processor elements.
27. A processor system according to claim 26, wherein the integrated element of the processor interface unit and the plurality of managed processor elements are arranged in a chip.
28. A processor system according to claim 26, wherein the integrated element of the processor interface unit and the plurality of managed processor elements are arranged in a chiplet.
29. A processor system according to any of claims 1 to 28, wherein the processor interface unit is configured to:receive one or more first instructions comprising configuration data and data to be processed; determine an algorithm for generating one or more second instructions based on the configuration data; andgenerate one or more second processing instructions for the managed processor elements to process the data based on the algorithm.
30. A processor system according to claim 29, wherein the processor interface unit is configured to maintain an internal state, and the algorithm is determined based on the configuration data and the internal state.7131. A processor system according to claim 30, wherein the processor interface unit is further configured to receive a state management instruction for managing the internal state.
32. A processor system according to any of claims 29 to 31 , wherein the one or more second processing instructions comprise a memory read instruction.
33. A processor system according to any of claims 29 to 32, wherein the algorithm is a vector processing algorithm, a matrix processing algorithm, a tensor processing algorithm or a non-scalar processing algorithm.
34. A processor system according to any of claims 29 to 33, wherein the processor interface unit comprises a memory storing one or more algorithms, and the algorithm for generating the one or more second instructions is determined from the stored one or more algorithms.
35. A processor system according to any of claims 29 to 34, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.
36. A processor system according to claim 34 or claim 35, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.
37. A processor system according to any of claims 34 to 36, wherein the processor interface unit is configured to dynamically update a stored algorithm or a stored sequence of instructions.
38. A processor system according to claim 37, wherein the processor interface unit comprises a machine learning model for dynamically updating the stored algorithm or the stored sequence of instructions.
39. A processor system according to any of claims 1 to 38, wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.
40. A processor system according to any of claims 1 to 39, wherein the plurality of managed processor elements comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.
41. A processor system according to claim 40, wherein the one or more second processing instructions comprise a machine code instruction, and the processor interface unit is configured to72input the machine code instruction directly to a register or functional unit of the plurality of managed processor elements .
42. A processor system according to claim 41 , wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the processor interface unit is configured to modify or rearrange data in a register of the plurality of managed processor elements while the one or more cores execute the SIMD instruction.
43. A processor system according to feature any of claims 40 to 42, wherein the one or more registers or functional units comprise a parallel execution unit.
44. A processor system according to claim 29 and according to claim 40, wherein the determined algorithm comprises:monitoring a state of the plurality of managed processor elements to obtain a current state; andgenerating a second processing instruction based on the current state, or re-ordering at least two second processing instructions based on the current state.
45. A processor system according to claim 44, wherein the current state comprises a resource utilization.
46. A processor system according to claim 34 and according to claim 40, wherein the processor interface unit is configured to:monitor a state of the plurality of managed processor elements over time to obtain a state history; anddynamically update the stored algorithm or the stored sequence of instructions based on the state history.
47. A processor system according to claim 46, wherein the state history comprises a resource utilization history.
48. A method performed by a processor interface unit, in a processor system comprising a plurality of managed processor elements and the processor interface unit, the method comprising:receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions;73configuring how the processing operation will be distributed across two or more of the managed processor elements; andcontrolling the two or more of the managed processor elements to perform the processing operation based on the one or more second processing instructions.
49. A method according to claim 48, wherein each of the plurality of managed processor elements is a compute unit configured to perform mathematical operations.
50. A method according to claim 48 or claim 49, further comprising controlling one or more interconnections between pairs of the plurality of managed processor elements.
51. A method according to claim 50, comprising controlling a plurality of interconnections between pairs of the plurality of managed processor elements, such that the plurality of managed processor elements performs operations as an interconnected fabric.
52. A method according to claim 51 , comprising controlling the plurality of managed processor elements and the plurality of interconnections between pairs of the plurality of managed processor elements such that the plurality of managed processor operations performs operations as a systolic array.
53. A method according to any of claims 48 to 52, comprising controlling the plurality of managed processor elements to operate as a field-programmable processing array.
54. A method according to any of claims 48 to 53, wherein the processor interface unit and the plurality of managed processor elements are together configured to operate as a coprocessor in cooperation with an unmanaged processor element.
55. A method according to claim 54, wherein the unmanaged processor element comprises one or more processor core(s).
56. A method according to claim 54 or claim 55, wherein:the processor interface unit and the plurality of managed processor elements are connected to a bus; and / orthe processor interface unit and the plurality of managed processor elements are configured to access a shared memory.
57. A method according to any of claims 48 to 56, wherein an order of the plurality of second processing instructions is based on a configuration of the plurality of managed processor elements.
58. A method according to claim 57, wherein the configuration of the plurality of managed processor elements comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in or between the plurality of managed processor elements.
59. A method according to claim 58, wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the plurality of managed processor elements; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the plurality of managed processor elements to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing suboperations and the allocated resources in the plurality of managed processor elements.
60. A method according to claim 59, comprising: iteratively repeating the steps of determining a flow of processing sub-operations and allocating resources, and evaluating an expected performance characteristic for the plurality of managed processor elements performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.
61. A method according to claim 59, comprising using a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.
62. A method according to any of claims 58 to 61 , wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the plurality of managed processor elements.
63. A method according to any of claims 48 to 62, wherein the one or more second processing instructions comprises at least one of the first processing instructions.
64. A method according to claim 63, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.
65. A method according to claim 64, wherein the at least two processing instructions comprise a memory read instruction.
66. A method according to any of claims 48 to 65, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the plurality of managed processor elements or a microarchitecture of the plurality of managed processor elements, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.
67. A method according to any of claims 48 to 66, wherein the processor interface is configured to operate with exclusive control of the plurality of managed processor elements.
68. A method according to claim 67, wherein the plurality of managed processor elements comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.
69. A method according to claim 67, wherein the exclusive control of the plurality of managed processor elements comprises disconnecting a dynamically switchable connection between the plurality of managed processor elements and another component of the system.
70. A method according to any of claims 48 to 69, wherein the processor interface unit comprises a non-integrated element separate from the plurality of managed processor elements.
71. A method according to claim 70, wherein the processor system comprises a plurality of chips, wherein the non-integrated element of the processor interface unit is arranged in a first chip, and the plurality of managed processor elements are arranged in one or more other chips separate from the first chip.
72. A method according to claim 70, wherein the processor system comprises a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the plurality of managed processor elements are arranged in one or more chiplets separate from the first chiplet.7673. A method according to any of claims 48 to 72, wherein the processor interface unit comprises an integrated element integrated within a managed processor element of the plurality of managed processor elements.
74. A method according to claim 73, wherein the integrated element of the processor interface unit and the plurality of managed processor elements is arranged in a chip.
75. A method according to claim 73, wherein the integrated element of the processor interface unit and the plurality of managed processor elements is arranged in a chiplet.
76. A method according to any of claims 48 to 75, comprising:receiving one or more first instructions comprising configuration data and data to be processed;determining an algorithm for generating one or more second instructions based on the configuration data; andgenerating one or more second processing instructions for the managed processor elements to process the data based on the algorithm.
77. A method according to claim 76, further comprising maintaining an internal state of the processor interface unit, and determining the algorithm based on the configuration data and the internal state.
78. A method according to claim 77, further comprising receiving a state management instruction for managing the internal state.
79. A method according to any of claims 76 to 78, wherein the one or more second processing instructions comprise a memory read instruction.
80. A method according to any of claims 76 to 79, wherein the algorithm is a vector processing algorithm, a matrix processing algorithm, a tensor processing algorithm or a non-scalar processing algorithm.
81. A method according to any of claims 76 to 80, wherein the processor interface unit comprises a memory storing one or more algorithms, and method comprises determining the algorithm for generating the one or more second instructions from the stored one or more algorithms.
82. A method according to any of claims 76 to 81 , wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.7783. A method according to claim 81 or claim 82, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.
84. A method according to any of claims 81 to 83, further comprising dynamically updating a stored algorithm or a stored sequence of instructions.
85. A method according to claim 84, further comprising using a machine learning model to dynamically update the stored algorithm or the stored sequence of instructions.
86. A method according to any of claims 48 to 85, wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.
87. A method according to any of claims 48 to 86, wherein the plurality of managed processor elements comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.
88. A method according to claim 87, wherein the one or more second processing instructions comprise a machine code instruction, and the method comprises inputting the machine code instruction directly to a register or functional unit of the plurality of managed processor elements.
89. A method according to claim 88, wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the method comprises modifying or rearranging data in a register of the plurality of managed processor elements while the plurality of managed processor elements execute the SIMD instruction.
90. A method according to any of claims 87 to 89, wherein the one or more registers or functional units comprise a parallel execution unit.
91. A method according to claim 76 and according to claim 87, wherein the determined algorithm comprises:monitoring a state of the plurality of managed processor elements to obtain a current state; andgenerating a second processing instruction based on the current state, or re-ordering at least two second processing instructions based on the current state.
92. A method according to claim 91 , wherein the current state comprises a resource utilization.
93. A method according to claim 81 and according to claim 87, further comprising:78monitoring a state of the plurality of managed processor elements over time to obtain a state history; anddynamically updating the stored algorithm or the stored sequence of instructions based on the state history.
94. A method according to claim 93, wherein the state history comprises a resource utilization history.
95. A computer program comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of claims 48 to 94.
96. A non-transitory storage medium storing instructions which, when executed by one or more processors, cause the processors to perform a method according to any of claims 48 to 94.
97. A data signal comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of claims 48 to 94.79