Processor system

WO2026202116A1PCT designated stage Publication Date: 2026-10-01RED SEMICONDUCTOR INTERNATIONAL LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/058495
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure EP2026058495_01102026_PF_FP_ABST
    Figure EP2026058495_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A processor system comprising one or more processor cores (200) and a processor interface unit (100), wherein the processor interface unit is configured to: receive one or more first processing instructions (110) associated with a processing operation; generate one or more second processing instructions (120) based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.
Need to check novelty before this filing date? Find Prior Art

Description

PROCESSOR SYSTEMFIELD OF THE INVENTION

[0001] The present disclosure relates to processor systems, and more particularly to a processor system, method, and computer program for improving processing performance through instruction translation and optimization using a processor interface unit.BACKGROUND

[0002] Processor systems have become increasingly complex and powerful over time, with advancements in microarchitecture design and instruction set architectures (ISAs) enabling improved performance and capabilities. Modern processors typically execute instructions according to a specific ISA, which defines the supported instructions, registers, memory addressing modes, and other aspects of the processor's programming model.

[0003] As processor designs have evolved, there has been a proliferation of different ISAs and microarchitectures tailored for various applications and use cases. This diversity can create challenges when developing software that needs to run efficiently across different processor types or when migrating existing code to new architectures. Additionally, the increasing complexity of modern ISAs and microarchitectures can make it difficult to fully utilize all of a processor's capabilities through conventional programming models.

[0004] One issue that arises is the potential mismatch between high-level algorithms or programming constructs and the low-level instructions supported by a particular processor architecture. This can lead to inefficiencies in code execution, as compilers may not always generate optimal instruction sequences for complex operations. Another challenge is that some desirable features or instructions may be available in one processor architecture but not in others, complicating software portability.

[0005] The growing demand for improved performance in areas such as artificial intelligence, cybersecurity, autonomous systems, data analysis, and real-time decision making has also highlighted limitations in how conventional processor systems handle certain types of computations. Vector and matrix operations, for instance, are common in many of these domains but may not be natively supported or optimally implemented in all processor architectures.

[0006] Additionally, there is a need to facilitate edge processing (applications on the periphery of a data processing ecosystem) within reduced performance / power consumption / cost limits, especially when the latency and / or security risk of exporting data to a data centre is unacceptable.

[0007] It has been appreciated that a processor system is needed that overcomes one or more of these problems.SUMMARY

[0008] The following disclosure provides various processor systems which comprise one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.

[0009] In other words, embodiments of this disclosure provide hardware which can act as a functional wrapper around traditional processor cores and / or processor elements, and which can enhance the functionality of the traditional processor cores and / or processor elements by, for example, modifying, translating or otherwise converting instructions before the instructions are passed to the core(s) and / or elements. This has many useful applications, including:

[0010] - Abbreviated instructions for common behaviours, where the processor interface unit is configured to decompress an abbreviated instruction into a corresponding set of standard instructions to cause the processor cores to perform a required operation. This can, for example, decrease the number of instructions which need to be retrieved from a remote memory, and reduce the risk that instruction retrieval times can act as a bottleneck to execution speed.

[0011] - Optimisation of instructions for a given processor core configuration. In contrast to softwarebased compilation, the processor interface unit can perform optimisation immediately prior to execution and can take into account the most detailed available picture of the available execution resources. This can include re-ordering of standard instructions in order to optimize usage of available resources. In some embodiments, the processor interface unit even monitors available resources in real-time and / or directly interacts with internal resources of processor cores.

[0012] - Performing calculations to handle a range of common engineering problems such as neural nets and / or Al, filtering and signal processing, cryptography, Position, Navigation and Timing (PNT),motor control, etc. For example, embodiments may be configured to efficiently perform one or more of:

[0013] — Matrix multiplication (such as matrix dot product), element-wise multiplication (such as matrix Hadamard product), or fast fourier transform,

[0014] — Calculations for various machine learning, statistical analysis and / or searching techniques including LSTM, CNN, Transformer, Probabilistic Graphical Models, Reinforcement Learning, Linear / Logistic Regression, Decision Tree, K means clustering, Gaussian Mixture Model, K nearest neighbours, Monte Carlo, Hidden Markov Model, Random Forest, Naive Bayes Classifier, Linear Discriminant Analysis, Breadth-First Search (BFS), Depth-First Search (DFS)

[0015] — Sorting and / or filtering algorithms

[0016] — Image processing algorithms

[0017] — C / C++ data structures and algorithms (DSA).

[0018] This may provide measurable benefits such as: fewer clock cycles taken to execute an algorithm or operation; fewer clock cycles taken to execute a repetitive stream or sequence of algorithms or operations; reduced power consumption by the one or more processor core(s) for the execution of algorithms or operations; reduced power consumption by the overall processor system for the execution of algorithms or operations; and / or reduced additional functionality (such as additional silicon area) required to increase the performance in a processor system.

[0019] Furthermore, customised embodiments of this disclosure can provide some of the benefits of ASIC design, while still using existing, mass-produced processor core components.

[0020] In particular, according to a first aspect, the following disclosure provides a processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.

[0021] For example, the configuration of the one or more processor cores may comprise a plurality of resources including processing resources, memory resources (e.g. registers, caches and any otherprocessor core memory) and / or communication pathway resources (e.g. input / output interfaces, or inter-core connections) in the one or more processor cores.

[0022] In more detail, in some embodiments of the first aspect, generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more processor cores; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more processor cores to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more processor cores.

[0023] Furthermore, in some embodiments of the first aspect, the processor interface unit is configured to: iteratively repeat the steps of determining a flow of processing sub-operations and allocating resources, and evaluate an expected performance characteristic for the one or more processor cores performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources. Such iterative improvement may be used on test examples of operations, in order to train a model which can be used to more immediately determine an improved flow of processing sub-operations and allocation of resources, for future operations. In other words, when a suitably trained model is available, the processor may additionally or alternatively be configured to use a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.

[0024] Furthermore, in some embodiments of the first aspect, the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more processor cores. Such access may, for example, be exposed by an existing commercially available processor core package. Alternatively, such access may be provided by a custom package in which the processor interface unit is at least partly integrated with the one or more processor cores.

[0025] According to a second aspect, the following disclosure provides a processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions; and control the one or more processor cores to perform theprocessing operation based on the one or more second processing instructions, wherein: the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, and the first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.

[0026] According to a third aspect, the following disclosure provides a processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to: receive a first processing instruction associated with a processing operation, the first processing instruction comprising an instruction prefix and an instruction body; determine an algorithm for generating one or more second instructions based on the instruction prefix; generate one or more second processing instructions based on the algorithm and the instruction body; and control the one or more processor cores to perform the processing operation based on the one or more second processing instructions.

[0027] In some embodiments, the processor interface unit is configured to maintain an internal state, and may receive state management instructions which manage this internal state. In embodiments of the third aspect, this internal state may be used when determining the algorithm based on the instruction prefix and the internal state. As one example, a prefix may indicate that the processor interface unit should use a vector processing algorithm (i.e. the processor interface unit should generate second processing instructions for vector processing). As further examples, a prefix may indicate that the interface unit should use a matrix processing algorithm, a tensor processing algorithm, or more generally any non-scalar format for data processing. This may, for example, be useful in order to select different algorithms based on a contextual state of the processor interface unit and / or of the one or more processor cores. This may also be useful in order to, for example, provide different externally-controlled general modes of the processor interface unit, wherein different modes may be associated with different algorithms for the same instruction prefix.

[0028] Each of the first, second and third aspects achieves benefits on their own, but any of these aspects can also be combined to provide even greater combined benefits.

[0029] Embodiments of the processor interface unit are preferably implemented in hardware which is as close as possible to the processor core(s), such as in a same chip or chiplet, or on a same circuit board. Nevertheless, in many embodiments, the processor interface unit is not integrated, or not fully integrated, with the processor core(s), and may be initially manufactured as a separate self-contained component, separate from the one or more processor core(s). For example, in some embodiments, the processor interface unit may be a coprocessor associated with the processor core(s).

[0030] Any of the methods performed by the processor interface unit may themselves be defined using executable instructions, such as a computer program, which may be stored in a memory.

[0031] In particular, according to a fourth aspect, the following disclosure provides a method performed by a processor interface unit, in a processor system comprising one or more processor cores and the processor interface unit, the method comprising: receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; and controlling the one or more processor cores to perform the processing operation based on the one or more second processing instructions.

[0032] According to a fifth aspect, the following disclosure provides a method performed by a processor interface unit, in a processor system comprising one or more processor cores and the processor interface unit, the method comprising: receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions; and controlling the one or more processor cores to perform the processing operation based on the one or more second processing instructions, wherein: the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, and the first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.

[0033] According to a sixth aspect, the following disclosure provides a method performed by a processor interface unit, in a processor system comprising one or more processor cores and theprocessor interface unit, the method comprising: receiving a first processing instruction associated with a processing operation, the first processing instruction comprising an instruction prefix and an instruction body; determining an algorithm for generating one or more second instructions based on the instruction prefix; generating one or more second processing instructions based on the algorithm and the instruction body; and controlling the one or more processor cores to perform the processing operation based on the one or more second processing instructions.

[0034] According to a seventh aspect, the following disclosure provides a computer program comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any one or more of the fourth, fifth and sixth aspects.

[0035] According to an eighth aspect, the following disclosure provides a non-transitory storage medium storing instructions which, when executed by one or more processors, cause the processors to perform a method according to any one or more of the fourth, fifth and sixth aspects.

[0036] According to a ninth aspect, the following disclosure provides a data signal comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any one or more of the fourth, fifth and sixth aspects.BRIEF DESCRIPTION OF THE FIGURES

[0037] Embodiments of the invention will be described, by way of example, with reference to the following drawings.

[0038] Fig. 1 is a block diagram schematically illustrating a processor system according to an embodiment;

[0039] Fig. 2 is a flow chart schematically illustrating a method performed by a processor interface unit, according to an embodiment;

[0040] Fig. 3 is a flow chart schematically illustrating a method performed by the processor interface unit, according to an embodiment;

[0041] Fig. 4 is a block diagram schematically illustrating various optional modules of the processor system, according to embodiments;

[0042] Figs. 5A to 5C are flow charts schematically illustrating further details of methods performed by the processor interface according to embodiments;

[0043] Fig. 6 is a block diagram schematically illustrating a processor system according to a partially integrated embodiment;

[0044] Fig. 7 is a block diagram schematically illustrating a processor system according to a partially integrated embodiment;

[0045] Fig. 8 is a block diagram schematically illustrating an example application of embodiments; and

[0046] Fig. 9 is a block diagram schematically illustrating an example configuration for a re-ordering module.

[0047] Common reference numerals are used throughout the figures to indicate similar features. DETAILED DESCRIPTION

[0048] The applicant has developed a system concept called VISC (Versatile Intrinsic Structured Computing), aspects of which are schematically explained in the following detailed disclosure. VISC uses a combination of new instructions and new hardware to improve the performance of microprocessor architectures. The new instructions are called "Abstracted VISC Instructions" (AVIs). These instructions can invoke functional elements of VISC's hardware in a manner optimised to a combination of one or more of: an algorithm or operation being executed; the data being processed; and the resource utilisation and availability in an underlying microprocessor architecture. The new hardware includes a number of functional elements which may be used individually or in combinations in order to accelerate certain functions defined by special or standard instructions. The elements may be logically positioned as a front-end pre-processor to one or more microprocessor core(s) and / or integrated within a microprocessor's internal pipeline. The aspects of VISC can be used together with various microprocessor configurations, to deliver performance improvements with minimal disruption to established software and hardware design techniques and methodology.

[0049] FIG. 1 illustrates a block diagram of a processor system. The processor system includes a processor interface unit 100 and one or more processor core(s) 200.

[0050] The processor core(s) may be commercially available standard processor cores. For example, each of the processor cores may be configured to use an instruction set architecture, ISA, such as open-source architectures like RISC-V or Power, or traditional proprietary architectures like x86 or Arm. Alternatively, each of the processor cores may be configured to use an applicationspecific architecture like MIPS or XMOS. Furthermore, each processor core may be configured to use a generic architecture like a neural processor architecture or a digital signal processor architecture. In embodiments where there is more than one processor core 200, the processor cores may behomogeneous or heterogeneous - in other words, the processor interface unit may be configured to interact with multiple cores using a same ISA or a mixture of cores each using a respective ISA. For example, in an embodiment, the processor cores 200 may include a RISC-V compatible core and an Arm core.

[0051] Some embodiments of a processor system may comprise managed processor elements, in addition or alternative to the processor core(s), or including the processor core(s). The managed processor elements may, for example, comprise compute units. The compute units may be specialised or may be general-purpose. The managed processor elements may be controlled according to the techniques described herein for controlling one or more processor core(s).

[0052] Each managed processor element may be configured to perform mathematical operations and / or may comprise one or more arithmetic logic units.

[0053] Additionally or alternatively, each managed processor element may comprise one or more interface elements suitable for communicating with other compute unit(s). For example, compute units may comprise one or more logic gates multiplexers or demultiplexers for sending or receiving data to or from other compute units.

[0054] The managed processor elements may be homogeneous. Alternatively, different managed processor elements may have different features and / or different arrangements.

[0055] In some examples, the managed processor units may be interconnected such that the plurality of managed processor elements performs operations as an interconnected fabric. In some examples, the managed processor units may be configured to perform operations as a systolic array. In some examples, the managed processor elements may be configured to operate as a field-programmable processing array.

[0056] The processor interface unit 100 can be integrated into a system design in various ways. In some examples, the processor interface unit 100 can be integrated as a front-end accelerator / translator / interpreter to one or more standard microprocessor cores. In other examples, the processor interface unit 100 can be fully integrated within the microprocessor architecture.

[0057] In some examples, the processor interface unit 100 may be configured to control one or more interconnections between pairs of processor cores and / or between pairs of managed processor elements. For example, the processor interface unit 100 may be configured to control interconnections between processor cores and / or between managed processor elements such thatthe processor cores and / or the managed processor elements perform operations as an interconnected fabric, or even as a systolic array or as a field-programmable processing array.

[0058] In an example, the processor interface unit 100 and the processor cores (and / or managed processor elements) 200 may be configured to operate together to provide a block capable of reconfigurable mathematics, to provide hardware-level efficiency improvements in applications such as performing calculations to handle a range of common engineering problems such as neural nets and / or Al, filtering and signal processing, cryptography, Position, Navigation and Timing (PNT), motor control, etc.

[0059] In some implementations, the processor interface unit 100 can be provided on a separate functional silicon dice, known as a chiplet, connected on a substrate to one or more microprocessor chiplets comprising the one or more processor core(s). Alternatively, the processor interface unit 100 can be provided on a separate functional chip connected on a printed circuit board to one or more microprocessor chips comprising the one or more processor core(s).

[0060] The processor interface unit 100 is configured to receive first processing instructions 110. The first processing instructions 110 may include a mixture of abstracted instructions 31 and / or core instructions 32. The first processing instructions 110 also include data such as operands within abstracted instructions 31 or core instructions 32.

[0061] Herein, it should be understood that core instructions 32 are instructions which can be received and executed by one or more of the processor cores 200 (and / or one or more of the managed processor elements).

[0062] On the other hand, abstracted instructions 31 are instructions intended for the processor interface unit 100. Abstracted instructions 31 may have different possible values, or a different instruction structure, from core instructions 32. For example, abstracted instructions 31 may have a shorter or longer bit length than core instructions 32. Additionally or alternatively, abstracted instructions may include a type header at the start of the instruction indicating that a payload body of the instruction is intended for the processor interface unit 100.

[0063] In some examples, the abstracted instructions 31 may be regarded as conventional instructions according to a known instruction set, such as RISC-V instructions. On the other hand, the core instructions 32 sent to one or more of the processor cores 200 (and / or one or more of the managed processor elements) may be instructions specially designated for controlling processingoperations and / or state memory at the processor cores 200 (and / or one or more of the managed processor elements), and may be instructions of an application-specific instruction set or general computation instructions.

[0064] Various types of abstracted instruction 31 are discussed below and these may, for example, include state management instructions for managing a state of the processor interface unit 100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions or tensor instructions), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.

[0065] The processor interface unit 100 may be configured to receive instructions passively, at whatever rate they arrive at an input interface of the processor interface unit 100. Alternatively, the processor interface unit 100 may be configured to fetch instructions from a configured static or dynamic location. For example, on first powering up the processor system, the processor interface unit 100 may fetch BIOS instructions, or fetch instructions from a predetermined memory.Subsequently, the processor interface unit 100 may fetch instructions from various locations according to a dynamic state of the processor interface unit 100.

[0066] The processor interface unit 100 processes the first processing instructions 110 and generates second processing instructions 120. The second processing instructions typically comprise core instructions 32 which can be executed by one or more of the processor core(s) 200. In some embodiments, the second processing instructions may also comprise operand data 33, which is data to be provided to the one or more processor core(s) 200 outside of core instructions 32.

[0067] For example, when the first processing instructions comprise a core instruction 32, generating second processing instructions may simply comprise passing the core instruction 32 as a second processing instruction for the processor core(s) to execute. Additionally, the processor interface unit may perform some management operations, such as selecting which of a plurality of processor cores 200 should execute the core instruction 32.

[0068] On the other hand, when the first processing instructions comprise an abstracted instruction 31 , the process of generating second processing instructions 120 may be more complex. For example, generating the second processing instructions 120 may comprise any one or more of: identifying an algorithm for generating the second processing instructions 120; allocating resources atthe processor core(s); organising vector, matrix, tensor and / or any type of non-scalar or SIMD processing operations; or re-ordering a sequence of generated instructions in order to more-optimally use available resources at the one or more processor cores.

[0069] The processor interface unit 100 provides the second processing instructions 120 to the processor core(s) 200 for execution.

[0070] For example, the processor interface unit 100 may be connected to an instruction input interface for each of the one or more processor core(s) 200, and may be configured to provide each of the second processing instructions 120 to one or more of the processor core(s) 200. For example, the processor core(s) 200 may be standard cores which are not adapted for, or in any way aware of, the processor interface unit 100.

[0071] The connection(s) between the processor interface unit 100 and the processor core(s) 200 may be exclusive connections, such that only the processor interface unit 100 can provide instructions to the one or more processor core(s) 200.

[0072] Alternatively, the processor interface unit 100 may share control of the one or more processor cores, and the processor system may additionally provide a direct interface for externally receiving core instructions 32 and providing the externally-received core instructions 32 to the one or more processor core(s) 200. However, this shared control is not preferred, as this reduces the extent to which the processor interface unit 100 can predict and / or manage available resources in the one or more processor core(s) 200.

[0073] As a compromise, the processor system may be configured to support dynamically switching between exclusive control and shared control of the one or more processor core(s). For example, the processor system may comprise a direct interface for externally receiving core instructions 32 as mentioned above, but the processor interface unit 100 may be configured with the capability to disconnect the direct interface in an exclusive mode, and to connect the direct interface in a nonexclusive mode.

[0074] Outputs from the processor core(s) 200 may be handled in a conventional manner. For example, outputs from the processor core(s) 200 may be stored in a volatile or non-volatile memory.

[0075] In some embodiments, the processor interface unit 100 may be configured to monitor or intercept outputs from the processor core(s) 200. For example, the processor interface unit 100 may be configured to generate further second processing instructions 120 based on outputs from theprocessor core(s) 200. In other words, in some embodiments, the processor system is configured to perform a sequence of inter-related steps without any communication outside of the processor system. This may, for example, be used to accelerate iterative calculations. This may be especially applicable when the processor interface unit 100 receives a first processing instruction 110 which is an algorithmic instruction. Such contained processing may also improve security, by reducing communication on an external data bus outside of the processor system, and enabling algorithms such as encryption to be performed with fewer opportunities for interception of sensitive data such as encryption keys.

[0076] In many embodiments, the processor interface unit 100 works in a real-time streaming configuration, simultaneously performing all of the functions of receiving first processing instructions, generating second processing instructions and controlling the processor core(s) 200. The processor interface unit 100 may also buffer selected first processing instructions 110 and / or second processing instructions 120, in order change the order of instructions or change the grouping of instructions, in order to more optimally use available resources at the one or more processor cores.

[0077] To summarise, the processor interface unit 100 is configured to act as an interface between the incoming first processing instructions 110 and the processor core(s) 200, translating and optimizing the instruction flow. This can enable various effects including: reducing the data quantity of first processing instructions supplied to the processor system (for example by using algorithmic instructions) and optimizing utilization of resources at the one or more processor core(s) 200.

[0078] In some embodiments, a processor interface unit 100 and one or more processor core(s) 200 may be a subsystem of a greater processor system.

[0079] For example, a processor system may comprise multiple processor interface units 100, each configured to control a respective one or more processor core(s) or managed processor elements. Additionally or alternatively a processor system may additionally comprise one or more conventionally-accessible processor(s) which are configured to receive instructions directly, without a processor interface unit 100.

[0080] In some processor systems, a subsystem comprising the processor interface unit 100 and the one or more processor core(s) 200 may be configured to act as a coprocessor together with one or more further subsystems. For example, the processor system may be configured to use a “main”processor subsystem for some common operations and to use the processor interface unit 100 and the one or more processor core(s) 200 to perform some specialised operations.

[0081] Preferably, the processor interface unit 100 is also supported through a development tool chain including compiler and assembler support for producing functions or programs comprising first processing instructions 110, suitable for the processor interface unit 100. This tool chain may allow developers to efficiently create and optimize code which takes advantage of the capabilities provided by the processor interface unit 100 and the processor core(s) 200. For example, this may enable mathematicians to take algorithms to code even if they are unfamiliar with a target microprocessor architecture, because the processor interface unit 100 can manage in hardware some details which the developer would have to consider in software for traditional microprocessor code.

[0082] Embodiments of the invention may also enable the creation of computer programs with reduced size, using abstracted instructions 31 and optionally also including core instructions 32, by comparison to a corresponding computer program consisting solely of core instructions 32.Furthermore, the use of abstracted instructions may reduce a total number of instructions to be read into the processor system, thereby potentially reducing latency and / or power consumption.

[0083] FIG. 2 is a flowchart schematically illustrating an outline method performed by the processor interface unit 100.

[0084] At step S201 , the processor interface unit 100 receives the first processing instructions 110 associated with a processing operation. The first processing instructions 110 may include a mixture of abstracted instructions 31 and / or core instructions 32. The first processing instructions 110 also include data such as operands within abstracted instructions 31 or core instructions 32.

[0085] Herein, it should be understood that core instructions 32 are instructions which can be received and executed by one or more of the processor cores 200.

[0086] On the other hand, abstracted instructions 31 are instructions intended for the processor interface unit 100. Abstracted instructions 31 may have different possible values, or a different instruction structure, from core instructions 32. For example, abstracted instructions 31 may have a shorter or longer bit length than core instructions 32. Additionally or alternatively, abstracted instructions 31 may include a type header at the start of the instruction indicating that a payload body of the instruction is intended for the processor interface unit 100.

[0087] Various types of abstracted instruction 31 are discussed below and these may, for example, include state management instructions for managing a state of the processor interface unit 100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions, tensor instructions or any other non-scalar or SIMD instruction format), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.

[0088] The processor interface unit 100 may be configured to receive instructions passively, at whatever rate they arrive at an input interface of the processor interface unit 100. Alternatively, the processor interface unit 100 may be configured to fetch instructions from a configured static or dynamic location. For example, on first powering up the processor system, the processor interface unit 100 may fetch BIOS instructions, or fetch instructions from a predetermined memory.Subsequently, the processor interface unit 100 may fetch instructions from various locations according to a dynamic state of the processor interface unit 100.

[0089] At step S202, the processor interface unit 100 generates the second processing instructions 120 based on the received first processing instructions 110. The second processing instructions typically comprise core instructions 32 which can be executed by one or more of the processor core(s) 200. In some embodiments, the second processing instructions may also comprise operand data 33, which is data to be provided to the one or more processor core(s) 200 outside of core instructions 32.

[0090] For example, when the first processing instructions comprise a core instruction 32, generating second processing instructions may simply comprise passing the core instruction 32 as a second processing instruction for the processor core(s) to execute. Additionally, the processor interface unit may perform some management operations, such as selecting which of a plurality of processor cores 200 should execute the core instruction 32.

[0091] On the other hand, when the first processing instructions comprise an abstracted instruction 31 , the process of generating second processing instructions 120 may be more complex. For example, generating the second processing instructions 120 may comprise any one or more of: identifying an algorithm for generating the second processing instructions 120; allocating resources at the processor core(s); organising vector, matrix, tensor and / or any type of non-scalar or SIMDprocessing operations; or re-ordering a sequence of generated instructions in order to more-optimally use available resources at the one or more processor cores.

[0092] At step S203, the processor interface unit 100 controls the processor core(s) 200 to perform the processing operation based on the second processing instructions 120.

[0093] FIG. 3 is a flowchart schematically illustrating a method for processing instructions performed by the processor interface unit 100, according to an embodiment.

[0094] The method begins at a step S301 , where the processor interface unit 100 receives a first processing instruction 110. The first processing instruction 110 is an abstracted instruction 31 and / or a core instruction 32. The first processing instruction 110 also include data such as an operand within the abstracted instruction 31 or core instruction 32.

[0095] Herein, it should be understood that core instructions 32 are instructions which can be received and executed by one or more of the processor cores 200.

[0096] On the other hand, abstracted instructions 31 are instructions intended for the processor interface unit 100. Abstracted instructions 31 may have different possible values, or a different instruction structure, from core instructions 32. For example, abstracted instructions 31 may have a shorter or longer bit length than core instructions 32. Furthermore, different sub-types of abstracted instruction 31 may have different bit lengths. Additionally or alternatively, abstracted instructions 31 may include a type header at the start of the instruction indicating that a payload body of the instruction is intended for the processor interface unit 100.

[0097] Various sub-types of abstracted instruction 31 are discussed below and these may, for example, include state management instructions for managing a state of the processor interface unit 100, prefix instructions indicating that subsequent instructions should be processed in a certain way (e.g. as vector instructions, matrix instructions, tensor instructions or any other non-scalar or SIMD instruction format), and / or algorithm instructions indicating that the processor interface unit should control the processor core(s) to perform a multi-instruction algorithm.

[0098] At a step S302, the processor interface unit 100 determines an instruction type of the first processing instruction 110.

[0099] For example, the processor interface unit may determine whether the first processing instruction 110 is a core instruction 32 or an abstracted instruction 31. This may, for example, be determined based on a value or a structure of the first processing instruction 110. For example, anabstracted instruction 31 may be detected based on a value, a bit length or a type header of the first processing instruction 110. Similarly, a core instruction 32 may be detected based on a value, a bit length or other structure of a core instruction 32.

[0100] The processor interface unit 100 may, for example, store a set of properties associated with different instruction types. The processor interface unit 100 may determine whether the received first processing instruction 110 matches the stored properties of any of the instruction types. For each instruction type, these stored properties may comprise a partial or complete instruction set architecture (ISA) specification.

[0101] If the first processing instruction is a core instruction 32 then, at step S303, the processor interface unit 100 passes the instruction unchanged as a second processing instruction 120 to the processor core(s) 200. Additionally, the processor interface unit 100 may perform some management operations for the core instruction 32, such as selecting which of a plurality of processor cores 200 should execute the core instruction 32.

[0102] On the other hand, if the first processing instruction 110 is an abstracted instruction 31 , the method proceeds to step S304, where the processor interface unit 100 generates one or more second processing instructions 120. The second processing instructions typically comprise core instructions 32 which can be executed by one or more of the processor core(s) 200. In some embodiments, the second processing instructions may also comprise operand data 33, which is data to be provided to the one or more processor core(s) 200 outside of core instructions 32. Some example implementations of step S304 are illustrated using Figs. 4 and 5A to 5C as discussed below.

[0103] Following the generation of the second processing instructions 120, the method moves to a step S305, where the processor interface unit 100 passes the second processing instruction(s) 120 to the processor core(s) 200.

[0104] According to the method of Fig. 3, the processor interface unit 100 can pass through core instructions 32, as well as handling abstracted instructions 31. This enables the processor system to be used as a standard processor system having standard cores, in a case where no specific abstracted instructions 31 have been generated. In other words, the processor system can run compiled code which was compiled without being specifically intended for a processor system according to the invention, increasing flexibility. At the same time, the processor system can providethe described advantages - including reduced program size, and increased execution efficiency - for programs which include abstracted instructions 31.

[0105] FIG. 4 is a block diagram schematically illustrating a processor interface unit 400 having modules according to some more detailed embodiments. The processor interface unit 400 may be used in processor systems, similar to the processor interface unit 100 of Fig. 1. According to various embodiments, different combinations of these modules may be included or omitted, depending on the requirements of a given application of the technology. The functions of the modules are independent, so each module can be included or omitted independently. The modules are only shown together in a single figure for brevity.

[0106] In some embodiments, the processor interface unit 400 includes a first instruction buffer module 410 for storing received first processing instructions 110. Including a buffer enables the processor interface unit 400 to receive multiple first processing instructions 110, and to generate second processing instructions 120 based on the multiple first processing instructions. For example, a first abstracted instruction 31 of the first processing instructions 110 may be a vector processing instruction, indicating that the processor interface unit 400 should control the one or more processor cores to perform vector processing, and a second abstracted instruction 31 of the first processing instructions may contain data for the vector processing.

[0107] In some embodiments, the processor interface unit 400 includes a fetch module 415.

[0108] The fetch module 415 may be configured to fetch first processing instructions 110. For example, the processor interface unit 100 may be configured to fetch instructions from a configured static or dynamic location. For example, on first powering up the processor system, the processor interface unit 100 may fetch BIOS instructions, or fetch instructions from a predetermined memory. Subsequently, the processor interface unit 100 may fetch instructions from various locations according to a dynamic state of the processor interface unit 100.

[0109] Additionally or alternatively, the fetch module 415 may be configured to retrieve further instructions or data based on previously processed first processing instructions 110. For example, where first processing instruction 110 is an algorithm instruction or a vector processing instruction, the fetch module 415 may retrieve any additional data required for the algorithm or for the vector processing, and buffer the additional data in a memory at the processor interface unit 400. Thus, when the one or more processor core(s) 200 are subsequently controlled to perform steps of thealgorithm or steps of the vector processing, any required additional data is ready to be swiftly provided to the one or more processor core(s), and the processing can be performed efficiently by the one or more processor core(s), without data fetch delays in the middle of the processing at the one or more processor core(s). The above example is equally applicable to matrix processing, tensor processor, and other non-scalar and / or SIMD processing types, as alternatives to vector processing.

[0110] In some embodiments, the processor interface unit 400 includes a second instruction buffer module 420 for storing generated second processing instructions 120 before they are sent to the one or more processor core(s). For example, the second instruction buffer module 420 may be used to control instruction timing, so that instructions are sent to a suitable processor core 200 which has available resources. As another example, storing second processing instructions in the second instruction buffer module 420 may enable re-ordering of the second processing instructions after further second processing instructions have been generated based on subsequent first processing instruction(s). This may, for example, help to make more efficient usage of available processor core resources.

[0111] In some embodiments, the processor interface unit 400 comprises an interface state module 430 configured to maintain state information for the processor interface unit 400.

[0112] This interface state may, for example, include any of: information about ongoing instructions which are being executed in the one or more processor cores 200; information about planned instructions which are due to be executed in the one or more processor cores 200; information about iterative processes in which an output from the one or more processor cores 200 may be used when generating a further second processing instruction 120; information about processes which require the fetching of additional data; information about an algorithm which is currently being executed or which is due to be executed.

[0113] The processor interface unit 400 may, for example, use any of this interface state information when controlling the fetch module 415, the vector execution module 440, the SIMD module 450, the algorithm module 460, or the re-ordering module 490.

[0114] In some embodiments, the processor interface unit 400 comprises a vector execution module 440 configured to handle vector processing in the processor system. Herein, it should be understood that vector processing comprises performing a same instruction yt= f xj multiple times, for each of multiple operand values (a "vector" of operand values x), in order to produce multiple result values (a"vector" of output values y. The vector processing unit is configured to manage execution resources for vector processing such that, when sufficient resources are available at the one or more processor core(s) 200, multiple copies of the instruction for different operands are executed simultaneously using different processing resources. This managed vector processing operation can be contrasted with a for-loop methodology found in known vector processing systems, in which a loop is executed for successive executions of the instruction using a same processing resource, meaning that the processing time scales with the number of operand values in the "vector", and the vector elements of the output from the instruction become available sequentially. Furthermore, even when parallel execution is not possible for all elements of the "vector" of operand values, the vector execution module 440 may be configured to start a downstream operation on elements of the vector result z = even before the vector result has been fully calculated, if the processor interface unit 400 has received first processing instructions corresponding to such a sequence of operations (for example in the case of an algorithmic instruction). Additionally or alternatively, the processor interface unit 400 may comprise a matrix execution module or a tensor execution module configured for performing a same instruction for each of a multi-dimensional collection of operand values.

[0115] In some embodiments, the processor interface unit 400 comprises a SIMD module 450 configured to control SIMD (single-instruction multiple-data) execution resources in the one or more processor core(s) 200. The SIMD module has similar uses to the vector execution module, but makes use of SIMD execution resources in the one or more processor core(s) 200 which support parallel execution for multiple operands, on a hardware level. For example, the SIMD module may rearrange a data sequence according to a configuration of the one or more processor core(s). More specifically, the SIMD module may rearrange a data sequence to fit available SIMD resources, according to their respective capacities for multiple data in each execution step. Additionally, in a scenario where a data length (such as a data word length) is shorter than a corresponding register length in an available SIMD resource, the SIMD module 450 may adjust alignment of data in registers to improve execution order and efficiency for execution across a broader set of operand data.

[0116] In some embodiments, the processor interface unit 400 comprises an algorithm module 460 configured to control execution of any of one or more predetermined algorithms. As examples, the predetermined algorithms may include frequently used mathematical operations such as Fast Fourier Transform, Discrete Cosine Transform, matrix operations, encoding operations or encryptionoperations, or ML model operations performed while training or operating an ML model. The predetermined algorithms need not be basic elementary algorithms, and can be complex algorithms. The predetermined algorithms may, in some cases, be tailored to individual customers, depending on which algorithms are frequently needed for their use case.

[0117] Each predetermined algorithm may simply comprise reading a predetermined sequence of second processing instructions from a memory, and adding any required operands. Alternatively, the pre-defined algorithm may comprise rules for generating the second processing instructions, without the sequence of second processing instructions being fully defined in advance. This may, for example, be necessary for more complex branching algorithms, or for feedback algorithms (such as iterative algorithms) where the total length of the algorithm may not be known in advance.

[0118] In addition to predetermined algorithms, the algorithm module 460 may support dynamically updated algorithms. For example, the algorithm module may be configured to optimise a predetermined algorithm over time based on monitored performance when the algorithm is used in the processor system. Such optimisation may comprise evaluation of candidate modified versions of the predetermined algorithm, in order to determine whether any of the modified versions provide improved performance. Candidate modified versions of the predetermined algorithm may be generated randomly, or may be generated using a trained machine learning model. Using a machine learning model may have the benefit of attempting modifications in parts of the algorithm which are likely to be improvable (such as loops which increase memory usage over time), while reducing the risk of modifications which are likely to break the algorithm.

[0119] For example, the algorithm module 460 may be implemented together with an algorithm library module 470 which stores predetermined algorithms for generating second processing instructions 120.

[0120] When controlling the one or more processor core(s) to execute second processing instructions 120 corresponding to a predetermined algorithm, the processor interface unit 400 may use a local memory, such as the interface state module 430, in order to perform the complete algorithm with minimal instances of storing data in an external memory or retrieving data from an external memory, thus substantially eliminating a slow operation type, and increasing the execution speed and efficiency for the predetermined algorithm. Furthermore, the processor interface unit 400 may use the fetch module 415 in order to anticipate any fetch operations so that data is retrievedbefore it is required, and the processor core(s) 200 do not need to wait for fetch operations any more than necessary.

[0121] In some embodiments, the processor interface unit 400 is configured to process specific subtypes of abstracted instruction 31 , which may feed into the control of the previously described modules.

[0122] As a first example, prefix instructions may be processed as a sub-type of abstracted instruction 31. Prefix instructions may be used to set up certain repetitive operations such as vector processing. Prefix instructions may for example include a vector processing instruction indicating that a subsequent instruction should be repeated for each of a plurality of operands. This can be contrasted with known systems which define vector instructions as separate instructions in the ISA-in other words, in this application, the combination of a prefix instruction with a subsequent instruction can enable vector processing for any subsequent instruction, without requiring the processor interface unit 400 to be configured to distinctly process a defined vector instruction (multiple-data instruction) for each single-data instruction. Prefix instructions may be shorter than other abstracted instructions 31 (in terms of bit length). The processor interface unit 400 may use a vector processing instruction as a trigger to perform operations using the vector execution module 440, and / or to allocate space in a memory (for example in the interface state module) for vector processing. This can be contrasted with known systems which have a bank of registers dedicated to vector operations, where those registers are not used for any other purpose, leading to inefficient usage of resources.

[0123] As a second example, algorithmic instructions may be processed as a sub-type of abstracted instruction 31. An algorithmic instruction may comprise a reference to a predetermined algorithm in the algorithm library module 470. For example, the algorithmic instruction may comprise a pointer to a memory location in the algorithm library module 470, in a case where the contents of the algorithm library module 470 are fixed and known to external entities producing first processing instructions 110 - such as compilers configured to produce algorithmic instructions. For example, a manufacturer of a processor interface unit 100 or a processor system according to this application may supply, to customers, an SDK (Software Development Kit) for the processor interface unit. This SDK may include the available algorithms which are stored in the algorithm library module 470, and which can be efficiently incorporated into programs executed on the processor system. The algorithm library module 470 (and the corresponding SDK) may, in some cases, be tailored to individual customers,depending on which algorithms are frequently needed for their use case. In some embodiments, the processor interface module 400 may additionally support adding algorithms to the algorithm library module 470 and / or removing algorithms from the algorithm library module 470. For example, after a processor interface module 400 has been manufactured (such as when the processor interface module 400 is installed in a specific computer system) a user may wish to add, modify or remove algorithms, in order to customise the processor interface module 400 to be maximally beneficial for their specific use case. In other words, while the application refers to "predetermined" algorithms, these algorithms may be "predetermined" at any stage prior to an instance of executing first processing instructions 110. For example, the processor system may be used to perform operations corresponding to a first set of first processing instructions 110 with a first set of predetermined algorithms in the algorithm library module 470, then modified to replace the first set of predetermined algorithms with a second set of predetermined algorithms in the algorithm library module 470, and then used again to perform operations corresponding to a second set of first processing instructions 110 with the second set of predetermined algorithms in the algorithm library module 470. The system may be configured to support changes to the algorithms stored in the algorithm library module 470 as a regular, generally available operation. Alternatively, system may be configured to support changes to the algorithms stored in the algorithm library module 470 in the form of a firmware update. As mentioned above, the predetermined algorithms need not be elementary algorithms, and may be complex - nevertheless, each predetermined algorithm may advantageously be identified using a single algorithmic instruction.

[0124] In some examples, algorithmic instructions may comprise a prefix and a body, where the prefix comprises a reference to a predetermined algorithm, and the body comprises data to be used as one or more parameters in the algorithm. In such cases, an algorithmic instruction may be implemented as two or more of the first processing instructions 110 - an algorithm prefix instruction, and one or more algorithm body instructions. When the algorithm prefix instruction is received, the processor interface unit 400 may store information in the interface state module 430 indicating that one or more algorithm body instructions is expected subsequently in the first processing instructions 110. Alternatively, the algorithmic instruction may be implemented as a single instruction including both of a prefix and a body.

[0125] As a third example, management instructions may be processed as a sub-type of abstracted instruction 31. Management instructions may be used to configure various settings and behaviours of the processor interface unit 400, and data indicating settings may be stored in the interface state module 430. Examples of management instructions include:

[0126] - Step

[0127] A Step instruction is a management instruction which may be used to increment or decrement a memory offset. For example, the memory offset may indicate a location in a memory resource (such as a register) of a processing core 200. More specifically, the memory offset may indicate a register file offset. The processor interface unit 400 may additionally or alternatively be configured to automatically increment a memory offset. However, manually controlling a memory offset may be useful in some scenarios, such as processing multiple first processing instructions 110 without incrementing a register file offset.

[0128] - Set Vector Length

[0129] A Set Vector Length instruction is a management instruction which may be used to set a vector length for vector processing handled by the vector execution module 440. In other words, the vector execution module 440 may be configured to read a vector length setting and, after receiving a vector processing instruction and a subsequent instruction in the first processing instructions 110, generate a number of second processing instructions 120 based on the vector length setting and based on the subsequent instruction. For example, if the subsequent instruction is a core instruction 32, the vector execution module 440 may generate (vector length) copies of the core instruction 32.

[0130] The Set Sequence Length instruction may also be used to set a vector operating mode, such as setting an "ignore vector processing instructions" mode, a mode in which a register file offset is automatically incremented for each element of a vector of operations, or a mode in which a register file offset is not incremented for each of (vector length) operations. That last mode may be used to "copy" a core instruction 32 for each operation, by using the same register location multiple times, without actually generating copies of the core instruction 32.

[0131] - Store Algorithm

[0132] A Store Algorithm instruction is a management instruction which may be used to indicate that subsequent first processing instructions 110 should be stored in the algorithm library module 470. Forexample, this may be used to add, remove, or update the predetermined algorithms stored in the algorithm library module 470.

[0133] - Configure Re-order

[0134] A Configure Re-order instruction is a management instruction which may be used to configure the re-ordering module 490. For example, this may be used to set sequence generator rules for rearranging instructions and operations.

[0135] In some embodiments, the processor interface unit 400 comprises a processor core state module 480 that maintains information about the state of the processor core(s) 200. This information may, for example, be used by the reordering module 490 to make informed decisions about instruction re-ordering or may be used by the second instruction buffer module 420 in order to make decisions about when to send second processing instructions 120 from the processor interface unit 400 to the one or more processor cores 200. Additionally, the processor core state module 480 may store a history of the state of the processor core(s) 200. For example, the history may comprise, for each of a plurality of time steps, associated data of a state of the processor core(s) 200 and data a state of the processor interface unit 400. By using such processor core state information (and optionally using historical state information), the processor system can help to reduce inefficiencies such as waiting for busy processing resources, or errors or delays due to simultaneous attempts to access a same memory resource, which can in turn lead to stall events, wasted clock cycles or wasted power.

[0136] As discussed below, in embodiments where the processor interface unit has read-access and / or write-access to resources within the one or more processor core(s), the processor interface unit may directly monitor availability of resources in the one or more processor core(s).

[0137] Alternatively, the processor core state module 480 may predict usage of resources within the one or more processor core(s), based on the configuration of the one or more processor core(s). For example, commercially-available processor cores may include a known combination of processing resources (such as decoders, arithmetic logic units, floating point units, SIMD units or adders), memory resources (such as registers, caches and any other processor core memory) and / or communication pathway resources (such as input / output interfaces, or inter-core connections). The processor interface unit may use a known configuration for each of the one or more processor cores, together with a knowledge of the instructions (and data) which have been input to the processor cores, in order to predict a current state of the resources in the one or more processor cores.

[0138] The processor core(s) state may, for example, include any of: a configuration of the one or more processor cores 200; a monitored current state of the one or more processor cores 200; a predicted current state of the one or more processor cores 200; data or statistics regarding historical performance of the one or more processor cores 200.

[0139] The processor interface unit 400 may, for example, use any of this processor core(s) state information when controlling the fetch module 415, the vector execution module 440, the SIMD module 450, the algorithm module 460, or the re-ordering module 490.

[0140] In some embodiments, the processor interface unit comprises a re-ordering module 490 configured to re-arrange instructions and operations. For example, the re-ordering module may be configured to apply a set of sequence generator rules for re-arranging instructions and operations. Additionally or alternatively, the re-ordering module may comprise one or more sequence generators, an example of which is described below with reference to Fig. 9.

[0141] Such re-ordering of instructions is useful in a wide variety of cases, to improve execution speed and / or improve efficient usage of processing resources. In particular, in many cases, it is useful for the re-ordering module 490 to re-order a sequence of second processing instructions 120 based on a configuration of the one or more processor core(s) 200. Additionally, such re-ordering may ensure that operands can be accessed using a linear order of data access operations, reducing or avoiding the need for branched data access when second processing instructions 120 are executed in the core(s) 200.

[0142] As a first case, the re-ordering module 490 may be configured to re-order execution of a sequence of second processing instructions 120 depending on a predicted availability of resources in the one or more processor core(s).

[0143] Based on actual and / or predicted knowledge of resource usage in the one or more processor cores, the processor interface unit may try to optimise an order of execution for the second processing instructions 120 at the one or more processor cores 200.

[0144] As a second case, the re-ordering module 490 may be configured to determine a linear or non-linear sequence for loading operand data into registers. As discussed further below, in embodiments where the processor interface unit has write-access to resources within the processor cores, the re-ordering module may determine a sequence for loading operand data into registers within the one or more processor cores. Additionally or alternatively, the re-ordering module 490 maybe configured to determine a sequence for loading operand data into registers of the processor interface unit, such as for inclusion in second processing instructions 120.

[0145] As a third case, the re-ordering module 490 may be configured to calculate register offsets or other memory locations for storing and / or retrieving operand data, to enable a computed execution on the operands in a linear or non-linear order.

[0146] As a fourth case, the re-ordering module 490 may be configured to calculate register offsets or other memory locations for storing and / or retrieving output data following a calculation at the one or more processor cores. Such output destinations may be configured using additional second processing instructions 120 or operands within the second processing instructions.

[0147] As a fifth case, the re-ordering module may be configured to re-order operands in registers for efficient execution of SIMD operations, whether as part of vector processing or in other contexts.

[0148] In some embodiments, the processor interface unit 400 also includes a memory management module and / or a garbage collection module (not shown), in order to manage any data that it stores. For example, a memory management module may allocate memory space for a matrix calculation, until a result is ready to be output from the processor system. As another example, data from intermediate steps of an algorithm, which may be stored while the processor interface unit 400 is controlling execution of an algorithm, is preferably cleared up as soon as it is no longer required for the algorithm.

[0149] In some embodiments, the processor interface unit 400 also includes a pointer alignment module and / or a programme counter module. Misalignment due to a branch instruction or a jump instruction can sometimes require re-alignment of pointers and memory address offsets, or flushing of data pipelines. Such misalignments may be detected or predicted based on a state of the processor core(s) or based on the second processing instructions 120. By quickly detecting, or even predicting, misalignments, some interruptions and latency can be avoided, improving overall performance.

[0150] The modules within the processor interface unit 400 interact to optimize instruction execution. For example, the first instruction buffer module 410 may distribute first processing instructions 110 to the appropriate processing modules (e.g., vector execution module 440, SIMD module 450, algorithm module 460) based on instruction type. The re-ordering module 490 may then coordinate with these processing modules to optimize instruction execution order based on information from the processor core state module 480 and the interface state module 430.

[0151] Through these interactions, both in terms of the individual modules, and the combined effects of the multiple modules, the processor interface unit 400 can control the processing core(s) 200 to more efficiently perform tasks indicated by the first processing instructions 110.

[0152] Referring again to Fig. 3, the processor interface unit 100, 400 (which at least has functionality as described for Fig. 1 , and may further have functionality as described for Fig. 4) can perform various functions within the step of "generating one or more second processing instructions" (step S304). Examples of these detailed functions will now be described with reference to Figs. 5Ato 5C.

[0153] FIG. 5A is a flow chart schematically illustrating a first embodiment S304-A of step S304 of Fig. 3.

[0154] At a step S501 , the processor interface unit 100, 400 determines a configuration of the processor core(s) 200. This configuration may include information about available processing resources, memory resources, and communication pathways within the processor core(s) 200.

[0155] This configuration may be provided in advance to the processor interface unit 100, 400. For example, the processor interface unit 100, 400 this configuration may be stored in a processor core(s) state module 480, and step S501 may simply comprise accessing the predetermined configuration. For example, commercially-available processor cores may include a known combination of processing resources (such as decoders, arithmetic logic units, floating point units, SIMD units or adders), memory resources (such as registers, caches and any other processor core memory) and / or communication pathway resources (such as input / output interfaces, or inter-core connections). A known configuration for each of the one or more processor cores may be stored in the processor core(s) state module 480.

[0156] Additionally or alternatively, the processor interface unit 100, 400 may have access to monitor the one or more processor core(s) and determine the configuration of the processor core(s) based on observations.

[0157] Then, at an optional step S502, the processor interface unit 100 determines a current state of the processor core(s) 200. This state information may, for example, include details about the current utilization of resources and the status of ongoing operations.

[0158] As discussed below, in embodiments where the processor interface unit has read-access and / or write-access to resources within the one or more processor core(s), the processor interface unit may directly monitor availability of resources in the one or more processor core(s).

[0159] Alternatively, the processor core state module 480 may predict usage of resources within the one or more processor core(s), based on the configuration of the one or more processor core(s). The processor interface unit may use a known configuration for each of the one or more processor cores, together with a knowledge of the instructions (and data) which have been input to the processor cores, in order to predict a current state of the resources in the one or more processor cores.

[0160] The process then moves to a step S503, where the processor interface unit 100, 400 determines a flow of processing sub-operations. This flow represents the sequence of operations required to execute the first processing instructions 110. In other words, before generating second processing instructions - which are real-world executable instructions designed for the specific execution context in the one or more processor core(s) - the processor interface unit 100, 400 may perform a functional or conceptual assessment of the required operations, and break this functional or conceptual assessment into a series of steps.

[0161] At a step S504, the processor interface unit 100, 400 allocates resources in the processor core(s) 200 to perform each sub-operation. This allocation is based on the determined configuration and current state of the processor core(s) 200, as well as the required processing for each suboperation.

[0162] Finally, at step S505, the processor interface unit 100 generates the second processing instructions 120 based on the determined flow of sub-operations, and based on the resources allocated for each sub-operation. This generation process may, for example, involve the re-ordering module 490 to optimize the instruction sequence based on the available resources and current processor state.

[0163] Further to the steps illustrated in Fig. 5A, the determining of sub-operations (step S503) and the allocation of resources (step S504) may be an iterative process. For example, the processor interface unit 100, 400 may be configured to predict a performance characteristic (e.g. total processing resources occupied over time, or total time to reach a final result) of a given flow of processing sub-operations and a given allocation of resources. The processor interface unit may use such a predicted performance characteristic in order to compare multiple candidate implementations,each comprising a flow of processing sub-operations and an allocation of resources. The processor interface unit 100, 400 may use such comparisons in order to iteratively locate an improved implementation which may be the best available implementation, or may be an implementation which exceeds a threshold performance characteristic. Alternatively, rather than iteration, the processor interface unit 100, 400 may simply compare a set of candidate implementations in parallel, and select the best of the candidate implementations.

[0164] Such optimisation strategies may comprise re-ordering processing sub-operations, or even re-ordering operations at the level of the received first processing instructions 110. For example, when the first processing instructions 110 comprise an abstracted instruction 31 followed by a core instruction 32, the steps may in some circumstances be re-ordered so that functionality corresponding to the core instruction 32 is executed by the one or more processor core(s) 200 before functionality corresponding to the abstracted instruction 31 by the one or more processor core(s) 200.

[0165] Of course, such optimisations as described in the preceding paragraphs may be difficult to perform in real-time. As an alternative, a model may be trained using machine learning for determining the flow of processing sub-operations (S503) and allocating resources (S504) based on the first processing instructions 110 and the processor core configuration (S501) - and optionally based on the processor cores current state (S502). Such a pre-trained model may be used to implement steps S503 and S504. A training phase for training the model may be performed for a range of training scenarios with different first processing instructions, different processor core configurations and different processor core current states. This training phase may involve iterative and / or parallel evaluation of candidate implementations, as discussed in the preceding paragraphs.

[0166] As a further extension to the method of Fig. 5A, the processor interface unit 100, 400 may be configured to store determined flows of processing sub-operations (S503) and / or store resource allocations (S504) for future reference. For example, if a sequence of operations is performed repeatedly, the flow of sub-operations the resources may be determined and allocated a first time the sequence is performed. On the other hand, in subsequent times when the sequence is performed, steps S503 and S504 may be efficiently implemented by reading the determined sub-operations and the allocated resources from memory. In other words, the efficiency gains of using a processor interface unit 100 may increase over time, as the processor system learns how to efficiently handle regular sequences of operations.

[0167] FIG. 5B is a flow chart schematically illustrating a second embodiment S304-B of step S304 of Fig. 3. The embodiments S304-A and S304-B provide different but compatible functions, and they may be implemented separately or together.

[0168] At step S511 , the processor interface unit 100, 400 determines an architecture of the processor core(s) 200.

[0169] The processor core(s) 200 may be commercially available standard processor cores. For example, each of the processor cores may be configured to use an instruction set architecture, ISA, such as open-source architectures like RISC-V or Power, or traditional proprietary architectures like x86 or Arm. Alternatively, each of the processor cores may be configures to use an applicationspecific architecture like MIPS or XMOS. Furthermore, each processor core may be configured to use a generic architecture like a neural processor architecture or a digital signal processor architecture.

[0170] The determination at step S511 may involve identifying, for each of the one or more processor core(s), an instruction set architecture (ISA) for instructions which can be externally received by the processor core(s). In other words, the processor interface unit 100, 400 may determine a set of available instruction types which can be used as second processing instructions 120 for communicating with each of the one or more processor core(s). In embodiments where there is more than one processor core 200, the processor cores may be homogeneous or heterogeneous - in other words, the processor interface unit may be configured to interact with multiple cores using a same ISA or a mixture of cores each using a respective ISA.

[0171] Additionally or alternatively, this determination at step S511 may involve determining, for each of the one or more processor core(s), an architecture for instructions and data which can be used internally within the processor core. For example, the architecture may be a microarchitecture of the processor core. This may be particularly applicable in embodiments where the processor interface unit 100, 400 has read-access and / or write-access to resources within the one or more processor core(s) - for example in embodiments where the processor interface unit 100, 400 is partially or fully integrated with the processor core(s) in a single package (such as a chip or a chiplet).

[0172] Following the architecture determination, at step S512, the processor interface unit 100 generates the second processing instructions 120 according to the architecture of the processor core(s) 200.

[0173] This step S512 may involve generating some ISA-specific instructions that can be received and executed by the processor core(s) 200.

[0174] As previously mentioned, some of the first processing instructions may be core instructions 32 suitable for execution in a processor core. Nevertheless, the processor interface unit 100, 40 may control instructions to be executed in a core which uses a different ISA from the received core instruction 32. Different ISAs may have some instruction types which are the same in both architectures, but in general a first ISA will include at least one instruction type that is not included in a second, different ISA, or the first ISA will lack at least one instruction type that is included in the second, different ISA. In such cases, step S512 may comprise processing a first processing instruction 110 which is a core instruction 32 conforming to a first ISA, and generating a second processing instruction 120 according to a second ISA different from the first ISA, where the second processing instruction 120 is sent to a processor core 200 instead of the received first processing instruction 110. By this mechanism, the processor system may control "execution" of received core instructions 32 even when the received core instructions do not match an ISA of an available processor core.

[0175] This step S512 may also involve generating some second processing instructions which are designed to be inserted using access to the internal structure of a processor core. For example, some of the second processing instructions may conform to a microarchitecture used within the processor core, or may be configured to be read and / or executed by a specific processing resource within the processor core. The second processing instructions 120 may, for example, include opcodes which do not need to be decoded or translated in a processor core prior to execution.

[0176] FIG. 5C is a flow chart schematically illustrating a third embodiment S304-C of step S304 of Fig. 3. The embodiments S304-A and S304-C provide different but compatible functions, and they may be implemented separately or together. The embodiments S304-B and S304-C provide different but compatible functions, and they may be implemented separately or together. Furthermore, in some embodiments, all three of S304-A, S304-B and S304-C are implemented together.

[0177] At an optional initial step S521 , the processor interface unit 100, 400 receives a state management instruction and updates an interface state of the processor interface unit 100, 400. This interface state may be maintained by an interface state module 430 and used to inform subsequent instruction processing decisions.

[0178] At step S522 (which may follow step S521 , or which may be performed as a first part of step S304 in Fig. 3), the processor interface unit 100 determines an algorithm based on an algorithm prefix of a first processing instruction 110 and optionally based on the interface state. This step may involve the algorithm module 460 working in conjunction with the algorithm library 470 to retrieve an appropriate algorithm for generating second processing instructions 120.

[0179] At S523, the processor interface unit 100 generates the second processing instructions 120 based on the determined algorithm and an algorithm body of the first processing instructions 110. As previously described, an algorithmic instruction may comprise an algorithm prefix and the algorithm body in a single first processing instruction 110, or the algorithm prefix and algorithm body may be split between multiple first processing instructions 110. This generation process may utilize various components of the processor interface unit 100, such as the vector execution module 440 or the SIMD module 450, depending on the nature of the algorithm and the instruction.

[0180] FIG. 6 is a block diagram schematically illustrating a processor system with a partially-integrated configuration. The processor system includes the processor interface unit 100 and the processor core(s) 200.

[0181] Fig. 6 is similar to Fig. 1 , with some additional features, and like features are indicated with like reference numerals. The existing features of Fig. 1 will not be repeated unnecessarily.Additionally, the processor interface unit 100 may comprise some or all of the modules described with reference to Fig. 4.

[0182] As illustrated in Fig. 6, in some embodiments, some or all of the one or more processor core(s) 200 may include an interface connection 210.

[0183] The interface connection 210 may, for example, be a hardware modification which enables the processor interface unit 100 to read information from, and / or write information to, at least one of the memory resources and / or processing resources within the processor core. Such an interface connection may be used to monitor a current processor core state for the one or more processor cores 200, which the processor interface unit 100 can use in its various methods as discussed above. For example, a monitored current processor core state may be stored in a processor core(s) state module 480 as discussed above.

[0184] Additionally or alternatively, some types of processor core 200 may include an interface for monitoring a processor core state, as a standard feature of their outputs, without requiring anymodification - this may be used as the interface connection 210. In other words, some degree of monitoring a processor core state may be possible, even without any hardware integration between the processor interface unit 100, 400 and the one or more processor core(s) 200. For example, some Arm processors include a performance monitoring unit PMU which includes counters for certain event types which occur within the Arm core, and which may be used to at least partially understand a state of the Arm core.

[0185] The interface connection 210 may facilitate taking into account an internal state of the processor core(s) 200 when operating the processor interface unit 100. For example, the processor interface unit 100 may receive information about the current state of the processor core(s) 200 through the interface connection 210. This information may include details about resource utilization, execution status, or other relevant parameters of the processor core(s) 200.

[0186] By receiving this state information, the processor interface unit 100 can make informed decisions when generating the second processing instructions 120. For instance, the processor interface unit 100 may adjust the order or timing of instructions based on the current resource availability in the processor core(s) 200. This can lead to more efficient execution of instructions and improved overall system performance.

[0187] The interface connection 210 may support bidirectional communication. In other words, the interface connection 210 may also support write access to one or more internal resources (e.g. processing resources or memory resources) of a processor core. This may enable finer control of the processor core(s), and may enable more efficient execution of processing operations.

[0188] The configuration of Fig. 6 may correspond to a partial or full integration of the processor interface unit 100 and the processor core(s) 200, where at least one of the processor core(s) 200 is not a standard commercially available core and is modified for use together with a processor interface unit 100. For example, in embodiments where the processor interface unit 100 is partially or fully integrated with the processor core(s) 200, the processor interface unit 100 may comprise an integrated element integrated within an integrated processor core of the one or more processor cores. For example, the integrated element of the processor interface unit may be arranged on a same chip, or a same chiplet, as the integrated processor core. Alternatively, the configuration of Fig. 6 may be a non-integrated configuration, where the processor core(s) 200 are not modified specifically for use with a processor interface unit 100.

[0189] FIG. 7 is a block diagram schematically illustrating a partially or fully integrated processor system, in an embodiment where the processor interface unit 100 has write access to individual processing resources within a processor core 201 of the one or more processor cores 200.

[0190] As shown in Fig. 7, in an example scenario, the processor interface unit 100 controls the processor core 201 to perform a processor operation using a pipeline comprising multiple resources 220-1 to 220-n. Passed information (which may be unchanged, or may be computed or otherwise modified by each resource) is passed from resource to resource along the pipeline ultimately leading to an output 34 from the processor core 201. Each resource may, for example, comprise one or more processing resources (such as decoders, arithmetic logic units, floating point units, SIMD units or adders) and / or one or more memory resources (such as registers, caches and any other processor core memory)

[0191] At each stage of the pipeline, the processor interface unit 100 can provide an instruction 32 and / or data 33 to the respective resource 220-1 to 220-n. For example, if the resource 220-2 were a memory resource such as a register, the processor interface unit 100 could store relevant data 33 in the register.

[0192] Although the pipeline 220-1 to 220-n is linear in this example, processor cores often have branching and / or looping arrangements of resources, and the processor interface unit 100 may accordingly be configured to control any or all of the resources in the processor core 201.

[0193] Such an integrated scenario, where the processor interface unit 100 has access to resources within a processor core, enables further improvement of processing of instructions according to criteria such as execution time and resource usage.

[0194] For example, referring again to Figs. 4 and 5A, the processor interface unit 100 may use the re-ordering module 490 to optimize the pipeline of resources 220-1 to 220-n and a corresponding sequence of instructions and data provided to the resources. As another example, the direct access to individual resources may be used together with the vector execution module 440 and the SIMD module 450 to process vector and SIMD operations efficiently within the resources of the one or more processor cores.

[0195] As mentioned above, one example application of the techniques described herein is to facilitate edge processing, using constrained processing resources, where efficiency is very important.

[0196] FIG. 8 is a block diagram schematically illustrating an example data processing sequence that may be performed in an edge processing scenario, using the processor system described herein. The sequence includes multiple stages for processing and analysing data from input to output.

[0197] The sequence begins with a stage 810, comprising digital signal processing applied to input received from sensors and data sources. The digital signal processing may involve manipulating sensor signals to generate meaningful data streams. Such digital signal processing may benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may utilize the vector execution module 440 and SIMD module 450 to efficiently process multiple data points in parallel during this stage.

[0198] Following the digital signal processing, the sequence moves to a stage 820, comprising data fusion from multiple inputs. Data fusion may involve combining data from different types and viewpoints to create a constant 'picture' of situational awareness for the edge system. Such fusion processes are commonly mathematically intense and often involve 3D transformations. Such data fusion may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may receive algorithmic instructions and employ the algorithm module 460 and algorithm library 470 to perform complex mathematical operations required for data fusion.

[0199] After data fusion, the sequence proceeds to a stage 830, comprising data analysis. The data analysis stage may involve looking for changes or anomalies in the created 'picture' and using mathematical algorithms to detect objects or features. Such data analysis may again benefit from various of the above-described techniques. For example, in an implementation, the reordering module 480 of the processor interface unit 100 may optimize the execution of analysis algorithms across the processor core(s) 200.

[0200] The sequence then moves to a stage 840, where decision making is performed. Such decision making may, for example, comprise using Al and machine learning techniques. This stage may involve using algorithms to react appropriately and immediately to changing data patterns. Such decision-making may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may receive abstracted instructions 31 which efficiently represent complex Al and machine learning operations, thus reducing the total size of aprogram comprising instructions for the Al and machine learning operations, and reducing the data rate at which instructions need to be read into the processor system.

[0201] From stage 840, the sequence branches into two paths. The first path comprises outputting of system control signals. These control signals may, for example, be used to modify physical aspects of the system's behavior based on the decision-making results.

[0202] The second path flows through three additional processing stages for data storage and communication functions.

[0203] A stage 850 comprises data encoding to minimize size for storage or transmission. In many scenarios, a paper-trail of situations and responses must be stored and / or communicated with other elements of the system. The transmitted or stored data often needs to be compact, so encoding or compression algorithms are executed. Such data encoding may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may use specialized instructions and algorithms, re-ordering of instructions, and or knowledge of available resources at the one or more processor core(s), to efficiently compress data.

[0204] A stage 860 comprises data encryption for privacy and security. It is often desirable to prevent theft, interception or manipulation of data in the ecosystem by an adversary. Such threats can be mitigated using encryption to protect privacy and security, which often requires significant processing resources. Algorithms for creating keys, authenticating users and encrypting data are increasingly complex, and the data stream may need to be continuous and real-time. Such encryption may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may be configured to receive abstracted instructions corresponding to dedicated encryption algorithms stored in the algorithm library 470, and may be configured to generate second processing instructions according to those algorithms to secure the data.

[0205] A stage 870 comprises handling data communication with the ecosystem. This stage may involve preparing data for transmission using various communication protocols, which may for example involve error checking and correction. Such communication may again benefit from various of the above-described techniques. For example, in an implementation, the processor interface unit 100 may be specifically configured with algorithms for generating signals according to the signal structure requirements of a communication protocol, especially where such structures are used repeatedly over the course of regular communication. Thus, a program for communication accordingto the communication protocol may contain fewer, smaller abstract instructions referring to the algorithms stored at the processor interface unit 100, and an overall data size of the program may be reduced.

[0206] The output from stage 870 provides data for storage or communication within the larger system ecosystem.

[0207] This example data processing sequence demonstrates how the described processor system may be used to efficiently handle complex, multi-stage data processing tasks. The processor interface unit 100 enables optimized execution of various algorithms and operations across the different stages, potentially improving overall system performance and efficiency. Furthermore, where multiple stages of a data processing pipeline are known at an early stage, a suitable processor interface unit 100 may be configured to efficiently perform multiple stages of the data processing pipeline by any combination of: using algorithms stored at the processor interface unit, using vector processing, using re-ordering of instructions to more efficiently utilize the associated processor core resources.

[0208] Fig. 9 is a block diagram schematically illustrating an example configuration for a re-ordering module 490.

[0209] Referring to Fig. 9, the re-ordering module may comprise a series of operational layers, wherein outputs from a preceding layer are passed as inputs to a next layer.

[0210] The operational layers may perform various functions, including: counting clock pulses, mixing, modifying or skipping values, and combining values.

[0211] In the example of Fig. 9, a first layer 910 comprises a series of counters COUNTERO, COUNTER1 and COUNTER2, each of which counts up to a respective maximum MAX0, MAX1 , MAX2 before overflowing and incrementing a next counter in the series (except for the last counter COUNTER2 which does not have an overflow destination). The first counter COUNTERO is incremented by a clock signal Clk (which may be generated in the processor interface unit, in a processor core 200, or elsewhere). Using three counters may often be useful, for example, when performing operations over a set of data points in a 3D grid.

[0212] A second layer 920 is configured to permute outputs from the first layer 910. In other words, the second layer may change an order of the counters, and change which counter is least / most significant. A third layer 930 is configured to invert outputs from the second layer 920. For example, the third layer may change a counting direction (decreasing vs increasing). A fourth layer 940 isconfigured to skip outputs from the third layer 930. For example, the fourth layer may be used to ignore a single counter, by setting it to zero. These are just examples, and such mixing, modifying and / or skipping can be used to produce a wide variety of mathematical sequences.

[0213] A fifth layer 950 is configured combine values from the fourth layer 940. More specifically, in this example, the fifth layer includes two multiply units and an adder, which combine the outputs from the fourth layer 940 to generate an output sequence Seq. In this example, the fifth layer calculates elements of the sequence based onSeqi = AS_Out0i + AS MaxO * AS_0utli+ AS MaxO * AS_Maxl * AS_0ut2i

[0214] See the below examples for settings and generated sequences:Example MaxO Maxi Max2 Permute Skip Seq Order CounterA 3 2 2 0,2,1 1 0,0, 2, 2, 4, 4, 1 ,1, 3, 3, 5, 5 B 3 2 2 0,2,1 3 0,1, 0,1, 0,1, 2, 3, 2, 3, 2, 3 C 3 2 2 0,1,2 3 0,1, 2, 3, 4, 5, 0,1, 2, 3, 4, 5

[0215] In the above, the "Permute Order" column indicates the order in which counters are arranged by the permute layer, in increasing significance order.

[0216] The example sequences may, for example, be relevant for a matrix multiplication operation, where a 3x2 matrix is multiplied by a 2x2 matrix. Using a generated sequence, the matrices may be stored in a row-first order in a memory resource of a core 200 and accessed in the required by using the generated sequence to control a memory location offset before a corresponding second processing instruction 120 is decompressed

[0217] A series of operational layers, such as the example of Fig. 9, may be denoted as a "sequence generator". The re-ordering module may comprise one sequence generator, or a plurality of sequence generators. Some or all of the sequence generators may be implemented in hardware. Additionally, the re-ordering module 490 may comprise programmable re-ordering hardware such that different sequence generators can be configured. For example, when performing certain types of mathematical operation (and potentially working together with the algorithm module 460), the processor interfaceunit 400 may be configured to configure one or more sequence generators according to predetermined settings for the mathematical operation.

[0218] The subject-matter of the application also includes the following numbered feature combinations.1. A processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to:receive one or more first processing instructions associated with a processing operation; generate one or more second processing instructions based on the received one or more first processing instructions; andcontrol the one or more processor cores to perform the processing operation based on the one or more second processing instructions.Different Instruction Sets / Different Ordering2. A processor system according to feature combination 1 , wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores. 3. A processor system according to feature combination 2, wherein the configuration of the one or more processor cores comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in the one or more processor cores.4. A processor system according to feature combination 3, wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more processor cores; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more processor cores to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing suboperations and the allocated resources in the one or more processor cores.5. A processor system according to feature combination 4, wherein the processor interface unit is configured to: iteratively repeat the steps of determining a flow of processing sub-operations and allocating resources, and evaluate an expected performance characteristic for the one or more processor cores performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.6. A processor system according to feature combination 4, wherein the processor interface unit is configured to use a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.7. A processor system according to any of feature combinations 3 to 6, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more processor cores.8. A processor system according to any of feature combinations 1 to 7, wherein the one or more second processing instructions comprises at least one of the first processing instructions.9. A processor system according to feature combination 8, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.10. A processor system according to feature combination 9, wherein the at least two processing instructions comprise a memory read instruction.11. A processor system according to any of feature combinations 1 to 10, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.Exclusive control12. A processor system according to any of feature combinations 1 to 11 , wherein the processor interface is configured to operate with exclusive control of the one or more processor cores.13. A processor system according to feature combination 12, wherein the one or more cores comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.14. A processor system according to feature combination 12, wherein the exclusive control of the one or more cores comprises disconnecting a dynamically switchable connection between the one or more processor cores and another component of the system.Hardware arrangements15. A processor system according to any of feature combinations 1 to 14, wherein the processor interface unit comprises a non-integrated element separate from the one or more processor cores.16. A processor system according to feature combination 15, comprising a plurality of chips, wherein the non-integrated element of the processor interface unit is arranged in a first chip, and the one or more processor cores are arranged in one or more other chips separate from the first chip. 17. A processor system according to feature combination 15, comprising a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the one or more processor cores are arranged in one or more chiplets separate from the first chiplet.18. A processor system according to any of feature combinations 1 to 17, wherein the processor interface unit comprises an integrated element integrated within a processor core of the one or more processor cores.19. A processor system according to feature combination 18, wherein the integrated element of the processor interface unit and the one or more processor cores are arranged in a chip.20. A processor system according to feature combination 18, wherein the integrated element of the processor interface unit and the one or more processor cores are arranged in a chiplet.Instruction transforming21. A processor system according to any of feature combinations 1 to 20, wherein the processor interface unit is configured to:receive a first instruction comprising an instruction prefix and an instruction body; determine an algorithm for generating one or more second instructions based on the instruction prefix; andgenerate one or more second processing instructions based on the algorithm and the instruction body.22. A processor system according to feature combination 21 , wherein the processor interface unit is configured to maintain an internal state, and the algorithm is determined based on the instruction prefix and the internal state.23. A processor system according to feature combination 22, wherein the processor interface unit is further configured to receive a state management instruction for managing the internal state.24. A processor system according to any of feature combinations 21 to 23, wherein the one or more second processing instructions comprise a memory read instruction.25. A processor system according to any of feature combinations 21 to 24, wherein the algorithm is a vector processing algorithm.26. A processor system according to any of feature combinations 21 to 25, wherein the instruction prefix is a vector processing prefix.27. A processor system according to any of feature combinations 21 to 26, wherein the processor interface unit comprises a memory storing one or more algorithms, and the algorithm for generating the one or more second instructions is determined from the stored one or more algorithms.28. A processor system according to any of feature combinations 21 to 27, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.29. A processor system according to feature combination 27 or feature combination 28, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.30. A processor system according to any of feature combinations 27 to 29, wherein the processor interface unit is configured to dynamically update a stored algorithm or a stored sequence of instructions.31. A processor system according to feature combination 30, wherein the processor interface unit comprises a machine learning model for dynamically updating the stored algorithm or the stored sequence of instructions.Management32. A processor system according to any of feature combinations 1 to 31 , wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.Integrated functionality33. A processor system according to any of feature combinations 1 to 32, wherein the one or more processor cores comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.34. A processor system according to feature combination 33, wherein the one or more second processing instructions comprise a machine code instruction, and the processor interface unit is configured to input the machine code instruction directly to a register or functional unit of the one or more processor cores.35. A processor system according to feature combination 34, wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the processor interface unit is configured to modify or rearrange data in a register of the one or more processor cores while the one or more cores execute the SIMD instruction.36. A processor system according to feature any of feature combinations 33 to 35, wherein the one or more registers or functional units comprise a parallel execution unit.37. A processor system according to feature combination 21 and according to feature combination 33, wherein the determined algorithm comprises:monitoring a state of the one or more system cores to obtain a current core state; and generating a second processing instruction based on the current core state, or re-ordering at least two second processing instructions based on the current core state.38. A processor system according to feature combination 37, wherein the current core state comprises a core resource utilization.39. A processor system according to feature combination 27 and according to feature combination 33, wherein the processor interface unit is configured to:monitor a state of the one or more system cores over time to obtain a core state history; and dynamically update the stored algorithm or the stored sequence of instructions based on the core state history.40. A processor system according to feature combination 39, wherein the core state history comprises a core resource utilization history.41. A method performed by a processor interface unit, in a processor system comprising one or more processor cores and the processor interface unit, the method comprising:receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions; andcontrolling the one or more processor cores to perform the processing operation based on the one or more second processing instructions.42. A method according to feature combination 41 , wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores.43. A method according to feature combination 42, wherein the configuration of the one or more processor cores comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in the one or more processor cores.44. A method according to feature combination 43, wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more processor cores; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more processor cores to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more processor cores.45. A method according to feature combination 44, comprising: iteratively repeating the steps of determining a flow of processing sub-operations and allocating resources, and evaluating an expected performance characteristic for the one or more processor cores performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.46. A method according to feature combination 44, comprising using a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.47. A method according to any of feature combinations 43 to 46, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more processor cores.48. A method according to any of feature combinations 41 to 47, wherein the one or more second processing instructions comprises at least one of the first processing instructions.49. A method according to feature combination 48, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.50. A method according to feature combination 49, wherein the at least two processing instructions comprise a memory read instruction.51. A method according to any of feature combinations 41 to 50, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.52. A method according to any of feature combinations 41 to 51 , wherein the processor interface is configured to operate with exclusive control of the one or more processor cores.53. A method according to feature combination 52, wherein the one or more cores comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.54. A method according to feature combination 52, wherein the exclusive control of the one or more cores comprises disconnecting a dynamically switchable connection between the one or more processor cores and another component of the system.55. A method according to any of feature combinations 41 to 54, wherein the processor interface unit comprises a non-integrated element separate from the one or more processor cores.56. A method according to feature combination 55, wherein the processor system comprises a plurality of chips, wherein the non-integrated element of the processor interface unit is arranged in a first chip, and the one or more processor cores are arranged in one or more other chips separate from the first chip.57. A method according to feature combination 55, wherein the processor system comprises a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit isarranged in a first chiplet, and the one or more processor cores are arranged in one or more chiplets separate from the first chiplet.58. A method according to any of feature combinations 41 to 57, wherein the processor interface unit comprises an integrated element integrated within a processor core of the one or more processor cores.59. A method according to feature combination 58, wherein the integrated element of the processor interface unit and the one or more processor cores are arranged in a chip.60. A method according to feature combination 58, wherein the integrated element of the processor interface unit and the one or more processor cores are arranged in a chiplet.61. A method according to any of feature combinations 41 to 60, comprising:receiving a first instruction comprising an instruction prefix and an instruction body; determining an algorithm for generating one or more second instructions based on the instruction prefix; andgenerating one or more second processing instructions based on the algorithm and the instruction body.62. A method according to feature combination 61 , further comprising maintaining an internal state of the processor interface unit, and determining the algorithm based on the instruction prefix and the internal state.63. A method according to feature combination 62, further comprising receiving a state management instruction for managing the internal state.64. A method according to any of feature combinations 61 to 63, wherein the one or more second processing instructions comprise a memory read instruction.65. A method according to any of feature combinations 61 to 64, wherein the algorithm is a vector processing algorithm.66. A method according to any of feature combinations 61 to 65, wherein the instruction prefix is a vector processing prefix.67. A method according to any of feature combinations 61 to 66, wherein the processor interface unit comprises a memory storing one or more algorithms, and method comprises determining the algorithm for generating the one or more second instructions from the stored one or more algorithms.68. A method according to any of feature combinations 61 to 67, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.69. A method according to feature combination 67 or feature combination 68, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.70. A method according to any of feature combinations 67 to 69, further comprising dynamically updating a stored algorithm or a stored sequence of instructions.71. A method according to feature combination 70, further comprising using a machine learning model to dynamically update the stored algorithm or the stored sequence of instructions.72. A method according to any of feature combinations 41 to 71 , wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.73. A method according to any of feature combinations 41 to 72, wherein the one or more processor cores comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.74. A method according to feature combination 73, wherein the one or more second processing instructions comprise a machine code instruction, and the method comprises inputting the machine code instruction directly to a register or functional unit of the one or more processor cores.75. A method according to feature combination 74, wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the method comprises modifying or rearranging data in a register of the one or more processor cores while the one or more cores execute the SIMD instruction.76. A method according to feature any of feature combinations 73 to 75, wherein the one or more registers or functional units comprise a parallel execution unit.77. A method according to feature combination 61 and according to feature combination 73, wherein the determined algorithm comprises:monitoring a state of the one or more system cores to obtain a current core state; and generating a second processing instruction based on the current core state, or re-ordering at least two second processing instructions based on the current core state.78. A method according to feature combination 77, wherein the current core state comprises a core resource utilization.79. A method according to feature combination 67 and according to feature combination 73, further comprising:monitoring a state of the one or more system cores overtime to obtain a core state history; anddynamically updating the stored algorithm or the stored sequence of instructions based on the core state history.80. A method according to feature combination 79, wherein the core state history comprises a core resource utilization history.81 A computer program comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of feature combinations 41 to 80.82. A non-transitory storage medium storing instructions which, when executed by one or more processors, cause the processors to perform a method according to any of feature combinations 41 to 80.83. A data signal comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of feature combinations 41 to 80.

[0219] Features of any of the examples or embodiments outlined above may be combined to create additional examples or embodiments without losing the intended effect. It should be understood that the description of an embodiment or example provided above is by way of example only, and various modifications could be made by one skilled in the art. Furthermore, one skilled in the art will recognise that numerous further modifications and combinations of various aspects are possible. Accordingly, the described aspects are intended to encompass all such alterations, modifications, and variations that fall within the scope of the appended claims.

Claims

CLAIMS1. A processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to:receive one or more first processing instructions associated with a processing operation; generate a plurality of second processing instructions based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; andcontrol the one or more processor cores to perform the processing operation based on the plurality of second processing instructions.

2. A system according to claim 1 , wherein the configuration of the one or more processor cores comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in the one or more processor cores.

3. A system according to claim 2, wherein generating the plurality of second processing instructions based on the received one or more first processing instructions comprises:determining the configuration of the one or more processor cores;determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more processor cores to perform each sub-operation of the flow of processing sub-operations;generating the plurality of second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more processor cores.

4. A system according to claim 3, wherein the processor interface unit is configured to:iteratively repeat the steps of determining a flow of processing sub-operations and allocating resources, and evaluate an expected performance characteristic for the one or more processor cores performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.

5. A system according to claim 3, wherein the processor interface unit is configured to use a pretrained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.

6. The system according to any of claims 2 to 5, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more processor cores.

7. The system according to any of claims 1 to 6, wherein the one or more second processing instructions comprises at least one of the first processing instructions.

8. The system according to claim 7, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.

9. The system according to claim 8, wherein the at least two processing instructions comprise a memory read instruction.

10. A processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to:receive one or more first processing instructions associated with a processing operation; generate a plurality of second processing instructions based on the received one or more first processing instructions; andcontrol the one or more processor cores to perform the processing operation based on the plurality of second processing instructions,wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.

11. The system according to any of claims 1 to 10, wherein the processor interface is configured to operate with exclusive control of the one or more processor cores.

12. The system according to claim 11 , wherein the one or more cores comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.

13. The system according to claim 12, wherein the exclusive control of the one or more cores comprises disconnecting a dynamically switchable connection between the one or more processor cores and another component of the system.

14. A processor system comprising one or more processor cores and a processor interface unit, wherein the processor interface unit is configured to:receive a first processing instruction associated with a processing operation, the first processing instruction comprising an instruction prefix and an instruction body;determine an algorithm for generating a plurality of second instructions based on the instruction prefix;generate the plurality of second processing instructions based on the algorithm and the instruction body; andcontrol the one or more processor cores to perform the processing operation based on the plurality of second processing instructions.

15. A system according to claim 14, wherein the processor interface unit is configured to maintain an internal state, and the algorithm is determined based on the instruction prefix and the internal state.

16. A system according to claim 15, wherein the processor interface unit is further configured to receive a state management instruction for managing the internal state.

17. A system according to any of claims 14 to 16, wherein the one or more second processing instructions comprise a memory read instruction.

18. A system according to any of claims 14 to 17, wherein the algorithm is a vector processing algorithm.

19. A system according to any of claims 14 to 18, wherein the instruction prefix is a vector processing prefix.

20. A system according to any of claims 14 to 19, wherein the processor interface unit comprises a memory storing one or more algorithms, and the algorithm for generating the one or more second instructions is determined from the stored one or more algorithms.

21. A system according to any of claims 14 to 20, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.

22. A system according to claim 20 or claim 21 , wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.

23. A system according to any of claims 20 to 22, wherein the processor interface unit is configured to dynamically update a stored algorithm or a stored sequence of instructions.

24. A system according to feature combination 23, wherein the processor interface unit comprises a machine learning model for dynamically updating the stored algorithm or the stored sequence of instructions.

25. A processor system according to any preceding claim, wherein the processor interface unit comprises a non-integrated element separate from the one or more processor cores.

26. A processor system according to claim 25, comprising a plurality of chips, wherein the nonintegrated element of the processor interface unit is arranged in a first chip, and the one or more processor cores are arranged in one or more other chips separate from the first chip.

27. A processor system according to claim 25, comprising a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the one or more processor cores are arranged in one or more chiplets separate from the first chiplet.

28. A processor system according to any preceding claim, wherein the processor interface unit comprises an integrated element integrated within an integrated processor core of the one or more processor cores.

29. A processor system according to claim 28, wherein the integrated element of the processor interface unit and the integrated processor core are arranged in a chip or in a chiplet.

30. A processor system according to any of claims 1 to 29, wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.

31. A processor system according to any of claims 1 to 30, wherein the one or more processor cores comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.

32. A processor system according to claim 31 , wherein the one or more second processing instructions comprise a machine code instruction, and the processor interface unit is configured to input the machine code instruction directly to a register or functional unit of the one or more processor cores.

33. A processor system according to claim 32, wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the processor interface unit is configured to modify or rearrange data in a register of the one or more processor cores while the one or more cores execute the SIMD instruction.

34. A processor system according to feature any of claims 31 to 33, wherein the one or more registers or functional units comprise a parallel execution unit.

35. A processor system according to claim 14 and according to claim 31 , wherein the determined algorithm comprises:monitoring a state of the one or more system cores to obtain a current core state; and generating a second processing instruction based on the current core state, or re-ordering at least two second processing instructions based on the current core state.

36. A processor system according to claim 35, wherein the current core state comprises a core resource utilization.

37. A processor system according to claim 23 and according to claim 31 , wherein the processor interface unit is configured to:monitor a state of the one or more system cores overtime to obtain a core state history; and dynamically update the stored algorithm or the stored sequence of instructions based on the core state history.

38. A processor system according to claims 37, wherein the core state history comprises a core resource utilization history.

39. A method performed by a processor interface unit, in a processor system comprising one or more processor cores and the processor interface unit, the method comprising:receiving one or more first processing instructions associated with a processing operation; generating one or more second processing instructions based on the received one or more first processing instructions, wherein an order of the plurality of second processing instructions is based on a configuration of the one or more processor cores; andcontrolling the one or more processor cores to perform the processing operation based on the one or more second processing instructions.

40. A method according to claim 39, wherein the configuration of the one or more processor cores comprises a plurality of resources including processing resources, memory resources and / or communication pathway resources in the one or more processor cores.

41. A method according to claim 40, wherein generating the one or more second processing instructions based on the received one or more first processing instructions comprises: determining the configuration of the one or more processor cores; determining a flow of processing sub-operations associated with the processing operation; allocating resources in the one or more processor cores to perform each sub-operation of the flow of processing sub-operations; generating the one or more second processing instructions based on the flow of processing sub-operations and the allocated resources in the one or more processor cores.

42. A method according to claim 41 , comprising: iteratively repeating the steps of determining a flow of processing sub-operations and allocating resources, and evaluating an expected performance characteristic for the one or more processor cores performing the processing operation using the flow of processing sub-operations and the allocated resources, in order to determine an improved flow of processing sub-operations and allocation of resources.

43. A method according to claim 41 , comprising using a pre-trained model for determining the flow of processing sub-operations and allocating the resources based on the received one or more first processing instructions.

44. A method according to any of claims 39 to 43, wherein the processor interface unit is configured with read and / or write access to one or more of the plurality of resources in the one or more processor cores.

45. A method according to any of claims 39 to 44, wherein the one or more second processing instructions comprises at least one of the first processing instructions.

46. A method according to claim 45, wherein the one or more second processing instructions comprise at least two processing instructions of the first processing instructions, and an order of the at least two processing instructions in the second processing instructions is different from an order of the at least two processing instructions in the received first processing instructions.

47. A method according to claim 46, wherein the at least two processing instructions comprise a memory read instruction.

48. A method according to any of claims 39 to 47, wherein:the first processing instructions are instructions according to a first architecture and the second processing instructions are instructions according to a second architecture, wherein the first architecture is an instruction set architecture, ISA, and the second architecture is an ISA of the one or more processor cores or a microarchitecture of the one or more processor cores, andthe first architecture comprises an instruction type that is not included in the second architecture; or the second architecture comprises an instruction type that is not included in the first architecture.

49. A method according to any of claims 39 to 48, wherein the processor interface is configured to operate with exclusive control of the one or more processor cores.

50. A method according to claim 49, wherein the one or more cores comprise a pathway for receiving instructions, and the processor system is statically configured with an exclusive connection between the processor interface unit and the pathway for receiving instructions.

51. A method according to claim 49, wherein the exclusive control of the one or more cores comprises disconnecting a dynamically switchable connection between the one or more processor cores and another component of the system.

52. A method according to any of claims 39 to 51 , wherein the processor interface unit comprises a non-integrated element separate from the one or more processor cores.

53. A method according to claim 52, wherein the processor system comprises a plurality of chips, wherein the non-integrated element of the processor interface unit is arranged in a first chip, and the one or more processor cores are arranged in one or more other chips separate from the first chip.

54. A method according to claim 52, wherein the processor system comprises a plurality of chiplets in a package, wherein the non-integrated element of the processor interface unit is arranged in a first chiplet, and the one or more processor cores are arranged in one or more chiplets separate from the first chiplet.

55. A method according to any of claims 39 to 54, wherein the processor interface unit comprises an integrated element integrated within a processor core of the one or more processor cores.

56. A method according to claim 55, wherein the integrated element of the processor interface unit and the one or more processor cores are arranged in a chip.

57. A method according to claim 55, wherein the integrated element of the processor interface unit and the one or more processor cores are arranged in a chiplet.

58. A method according to any of claims 39 to 57, comprising:receiving a first instruction comprising an instruction prefix and an instruction body; determining an algorithm for generating one or more second instructions based on the instruction prefix; andgenerating one or more second processing instructions based on the algorithm and the instruction body.

59. A method according to claim 58, further comprising maintaining an internal state of the processor interface unit, and determining the algorithm based on the instruction prefix and the internal state.

60. A method according to claim 59, further comprising receiving a state management instruction for managing the internal state.

61. A method according to any of claims 58 to 60, wherein the one or more second processing instructions comprise a memory read instruction.

62. A method according to any of claims 58 to 61 , wherein the algorithm is a vector processing algorithm.

63. A method according to any of claims 58 to 62, wherein the instruction prefix is a vector processing prefix.

64. A method according to any of claims 58 to 63, wherein the processor interface unit comprises a memory storing one or more algorithms, and method comprises determining the algorithm for generating the one or more second instructions from the stored one or more algorithms.

65. A method according to any of claims 58 to 64, wherein the processor interface unit comprises a memory storing one or more sequences of instructions, and the determined algorithm comprises reading a stored sequence of instructions.

66. A method according to claim 64 or claim 65, wherein the stored algorithms comprise an algorithm for a mathematical operation, or wherein the stored sequence of instructions comprise a sequence of instructions for a mathematical operation.

67. A method according to any of claims 64 to 66, further comprising dynamically updating a stored algorithm or a stored sequence of instructions.

68. A method according to claim 67, further comprising using a machine learning model to dynamically update the stored algorithm or the stored sequence of instructions.

69. A method according to any of claims 39 to 68, wherein the first processing instructions comprise a state management instruction for managing a state of the processor interface unit.

70. A method according to any of claims 39 to 69, wherein the one or more processor cores comprise one or more registers or functional units, and the processor interface unit is configured with read and / or write access to the one or more registers or functional units.

71. A method according to claim 70, wherein the one or more second processing instructions comprise a machine code instruction, and the method comprises inputting the machine code instruction directly to a register or functional unit of the one or more processor cores.

72. A method according to claim 71 , wherein the one or more second processing instructions comprise single-instruction multiple-data, SIMD, instruction, and the method comprises modifying or rearranging data in a register of the one or more processor cores while the one or more cores execute the SIMD instruction.

73. A method according to any of claims 70 to 72, wherein the one or more registers or functional units comprise a parallel execution unit.

74. A method according to claim 58 and according to claim 70, wherein the determined algorithm comprises:monitoring a state of the one or more system cores to obtain a current core state; and generating a second processing instruction based on the current core state, or re-ordering at least two second processing instructions based on the current core state.

75. A method according to claim 74, wherein the current core state comprises a core resource utilization.

76. A method according to claim 64 and according to claim 70, further comprising:monitoring a state of the one or more system cores overtime to obtain a core state history; anddynamically updating the stored algorithm or the stored sequence of instructions based on the core state history.

77. A method according to claim 76, wherein the core state history comprises a core resource utilization history.

78. A computer program comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of claims 39 to 77.

79. A non-transitory storage medium storing instructions which, when executed by one or more processors, cause the processors to perform a method according to any of claims 39 to 77.

80. A data signal comprising instructions which, when executed by one or more processors, cause the processors to perform a method according to any of claims 39 to 77.60