Memory, Network, Processor
The multiprocessor system with a dynamically configurable communication network addresses inefficiencies in existing processor architectures by enabling efficient parallel processing and reduced power consumption, improving performance and simplifying software development.
Patent Information
- Application Number
- JP2023094618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-11-03
- Filing Date
- 2023-06-08
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2038-11-02
AI Technical Summary
Existing hardware architectures for processors, such as GPP, GPU, DSP, FPGA, and multi-core architectures, face inefficiencies in power consumption, flexibility, and complexity in software programming, leading to high energy costs and complex software development processes, especially when transitioning between hardware platforms.
A multiprocessor system with a dynamically configurable communication network and message nodes that facilitate efficient data routing and processing, enabling flexible communication paths and reduced power consumption through conditional execution and static scheduling of instructions.
The system achieves improved performance per watt of power dissipation and reduces software complexity by allowing efficient parallel processing with lower power consumption and simplified programming models.
Smart Images

Figure 0007698327000008 
Figure 0007698327000009 
Figure 0007698327000010
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) The inventors are Michael B. Doerr, Carl S. Dobbs, Mich ael B. Solka, Michael R. Trocino, Kenneth R. Faulkner, Keith M. Bindloss, Sumeer Ary a, John Mark Beardslee, and David A. Gibson of "Memory - Network Processor with Progr ammable Optimizations (Memory - Network Processor with Programmable Optimizations)", U.S. Patent No. 9,430,369, is hereby incorporated by reference in its entirety as shown in all respects in this specification.
[0002] The inventors are Michael B. Doerr, Carl S. Dobbs, Mich ael B. Solka, Michael R. Trocino, and David A. Gibson of "Multiprocessor fabric hav ing configurable communication that is s electively disabled for secure processin g (Multiprocessor configuration selectively disabled for secure processing)", U.S. Patent No. 9,424,441, is hereby incorporated by reference in its entirety as shown in all respects in this specification.
[0003] The present invention relates to a multi-processor system, and more particularly to improving the operation and execution of processors.
Background Art
[0004] The main purpose of a general hardware system is to achieve application-specific (non-programmable) hardware performance while maintaining full programmability. Historically, these two concepts have been at opposite extremes. Application-specific hardware is a fixed hardware solution that performs a specific function in the most efficient way possible. This solution is typically measured in terms of energy per function, or energy per one or more operations, and also in terms of functionality per (circuit) area, which may be related to the cost of the product. The cost of a chip product consists of many factors including die area and the final package. The cost should also take into account the entire ecosystem for developing the product. This ecosystem cost consists of the time to convert a specific application to a specific hardware solution, the number of specific hardware solutions required to configure the entire system, and the time required to integrate all of the specific hardware solutions with customized communication and memory structures. Thus, a fully integrated solution requires porting all of the many specific hardware solutions with custom interconnects, resulting in very large area requirements on a single chip die. Historically, this process has resulted in inefficient solutions in terms of area, energy, and time to market.
[0005] When you think about the world of programmability and the concept of target hardware, Market or situation in terms of hardware architecture and software development style is a general-purpose processor offered by Intel, AMD, and ARM. Purpose Processor (GPP), available from NVIDIA and AMD Graphical Processing Unit (G PU), from Texas Instruments and Analog Devices The resulting digital signal processor ssor, DSP), FPGA (Field Programmable Gate Array (FPGA) from Xilinx, Altera, etc. d Programmable Gate Array), Cavium and Tile Multi-core architectures / many-core processors derived from ra, as well as Application Specific Integrated Circuit ASIC (Advanced Semiconductor Device) or System On Chip (System On These are represented by 3G chips and SoCs.
[0006] GPP is for general purpose processing, i.e. old but proven, over 40 years old. Based on a hardware architecture that is considered to be a jack of all trades. The general purpose of GPP is to provide a set of supported operating systems (e.g. Windows ws and Linux (registered trademark) to terface, UI), and high-performance applications such as MSWord, Excel, email, etc. The main challenge is running applications that use interactive UI intensively. The hardware characteristics that have an impact are multi-level caches, complex hardware memory management units, large buses, and large clock control structures. In short, GPP dissipates a large amount of power to perform these tasks. From the perspective of software development, the goal is considered to be the simplest software programming model. This model is obtained from the perspective that the user is developing a single thread that runs continuously or serially. When introducing parallel processing, or multiple hardware threads (more than about four threads), the ability to program the threads efficiently becomes much more difficult. This is basically due to the fact that an architecture for supporting parallel thread operation has not been developed, and as a result, the hardware architecture requires complexity that results in an enormous amount of overhead to manage. The software programming model needs to introduce API or language extensions to support the specification of multiple software threads. This extension does not need to be complex, but unfortunately the current GPP hardware architecture requires such complexity. At a high level, the API that has been widely used for many years on all supercomputers in the world along with C, C++, Fortran, etc. is the MPI (message passing interface) API, which has been the
[0007] industry standard since the early 1990s. This MPI is a very simple and well-understood API that does not limit hardware implementation. The MPI API is independent of the hardware and is a widely used API in the industry for many years on all supercomputers in the world along with C, C++, Fortran, etc. It is the MPI (message passing interface) API, which has been the industry standard since the early 1990s. This MPI is a very simple and well-understood API Enables the specification of software threads and communication in an unrelated manner. This MPI AP I is different from other languages / APIs that natively specify the assumed underlying hardware model, such as OpenMP, Coarray Fortran, OpenCL, etc., and thus limits the flexibility of interpretation and causes forward compatibility issues. In other words, when using these other languages / APIs, the programmer has to rewrite the program for every new target hardware platform. GPU has historically been developed to process and target data display. GPU is hardware constrained by the off-core (external) memory model requirements and internal core memory model requirements of the GPU. Off-core memory requires data to be placed within the GPU memory space relative to the GPP. The GPU then fetches the data, operates on the data in a pipelined fashion, and then places the data back into the external memory space of the GPU. From here, the data can be sent to the display device, or the GPP has to move the data outside of the GPU memory space for further use / storage in an operation that receives general processing. The hardware is inefficient because (1) the support required to move data around to support off-core hardware constraints, and (2) the limited internal core memory structure similar to a deeply pipelined SIMD machine that is constrained to process data in an optimized pipeline. As a result, power consumption is high due to the inefficiency of the hardware for processing data. The software used
[0008] GPU has historically been developed to process and target data display. GPU is hardware constrained by the off-core (external) memory model requirements and internal core memory model requirements of the GPU. Off-core memory requires data to be placed within the GPU memory space relative to the GPP. The GPU then fetches the data, operates on the data in a pipelined fashion, and then places the data back into the external memory space of the GPU. From here, the data can be sent to the display device, or the GPP has to move the data outside of the GPU memory space for further use / storage in an operation that receives general processing. The hardware is inefficient because (1) the support required to move data around to support off-core hardware constraints, and (2) the limited internal core memory structure similar to a deeply pipelined SIMD machine that is constrained to process data in an optimized pipeline. As a result, power consumption is high due to the inefficiency of the hardware for processing data. The software used by the GPU is also affected by these hardware constraints. Off-core memory requires data to be placed within the GPU memory space relative to the GPP. The GPU then fetches the data, operates on the data in a pipelined fashion, and then places the data back into the external memory space of the GPU. From here, the data can be sent to the display device, or the GPP has to move the data outside of the GPU memory space for further use / storage in an operation that receives general processing. The hardware is inefficient because (1) the support required to move data around to support off-core hardware constraints, and (2) the limited internal core memory structure similar to a deeply pipelined SIMD machine that is constrained to process data in an optimized pipeline. As a result, power consumption is high due to the inefficiency of the hardware for processing data. The software used by the GPU is also affected by these hardware constraints. The hardware is inefficient because (1) the support required to move data around to support off-core hardware constraints, and (2) the limited internal core memory structure similar to a deeply pipelined SIMD machine that is constrained to process data in an optimized pipeline. As a result, power consumption is high due to the inefficiency of the hardware for processing data. The software used by the GPU is also affected by these hardware constraints. As a result, power consumption is high due to the inefficiency of the hardware for processing data. ·The programming model is extremely hardware - centric, such as OpenCL and CUDA, etc. Therefore, it is complex to achieve efficiency and not very portable. When trying to migrate to a new hardware - targeted platform, the code has to be rewritten and restructured.
[0009] DSP can be considered as a GPP (General - Purpose Processor) that is scaled down for general signal processing and uses an instruction set targeted at that processing. DSP has the same cache, MMU, and bus annoyances as its big brother / sister, the GPP. Additionally, any high - throughput processing functions, such as Viterbi / Turbo decoding or motion estimation, usually become ASIC accelerators with limited capabilities that support only a limited set of specific standards in the commercial market. The programming model is similar to GPP when targeting a single hardware thread, but since the hardware of the execution unit is the way of dealing with signal - processing instructions, to achieve any high efficiency, it is necessary to require hand - assembly of functions or use proprietary software libraries. When creating a multiple - parallel DSP architecture similar to the parallel GPP discussed above, the problem gets even worse.
[0010] FPGA has a completely different hardware approach where the function specification can be done at the bit - level and communication between logic functions is done by a programmable wired structure. This hardware approach introduces an enormous overhead and complexity. Due to this, hardware programming languages such as Verilog or VHDL perform efficient programming. Programmable wiring and programmable logic have , a timing convergence obstacle is introduced, which is similar to that required by ASIC / SOC but with a structured wired configuration, making the compilation process much more complex. Power dissipation and performance throughput regarding specific functions are much better than GPP or GPU when comparing one function at a time because the FPGA only accurately performs what is programmed and nothing else. However, if all the capabilities of GPP are to be implemented within the FPGA, it is obvious that the FPGA is much inferior to GPP. The difficulty of programming at the hardware level is obvious (e.g., timing convergence). The programming of FPGA is actually not "programming" but rather logic / hardware design. VHDL / Verilog are logic / hardware design languages, not programming languages.
[0011] Almost all multi-core architectures / many-core architectures adopt core processors, caches, MMUs, buses, and all related logic from a hardware perspective, and fabricate these together with their surrounding communication buses / configurations on the die. Examples of multi-core architectures are IBM's Cell, Intel and AMD's quad-core and N-multi-core, products of Cavium and Tilera, several custom SoCs, etc. Additionally, the power reduction achieved by multi-core architectures is mostly negligible. This result is due to the way multi-core It is derived from the fact that it merely replaces the way the GPU works. Mar In a multi-core architecture, the only actual power savings are in reducing some of the I / O drivers, which are no longer needed because the cores are now connected by a communication bus on the chip, whereas they used to be on separate chips. Therefore, the multi-core approach does not save much energy. Second, the software programming model has not improved from the GPP discussed above. In a multi-core architecture, the only actual power savings are in reducing some of the I / O drivers, which are no longer needed because the cores are now connected by a communication bus on the chip, whereas they used to be on separate chips. Therefore, the multi-core approach does not save much energy. Second, the software programming model has not improved from the GPP discussed above. In a multi-core architecture, the only actual power savings are in reducing some of the I / O drivers, which are no longer needed because the cores are now connected by a communication bus on the chip, whereas they used to be on separate chips. Therefore, the multi-core approach does not save much energy. Second, the software programming model has not improved from the GPP discussed above. In a multi-core architecture, the only actual power savings are in reducing some of the I / O drivers, which are no longer needed because the cores are now connected by a communication bus on the chip, whereas they used to be on separate chips. Therefore, the multi-core approach does not save much energy. Second, the software programming model has not improved from the GPP discussed above. In a multi-core architecture, the only actual power savings are in reducing some of the I / O drivers, which are no longer needed because the cores are now connected by a communication bus on the chip, whereas they used to be on separate chips. Therefore, the multi-core approach does not save much energy. Second, the software programming model has not improved from the GPP discussed above.
[0012] The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions. The list of problems identified with other approaches arises in specific markets because system developers entrust custom chips with specific GPPs, DSPs, and ASIC accelerators to form a system-on-chip (SoC). The SoC provides programmability when necessary and ASIC performance for specific functions to balance power dissipation and cost. However, the software programming model is now even more complex than that discussed under the programmable hardware solutions above. Additionally, the SoC may result in a loss of the flexibility associated with fully programmable solutions.
[0013] What is common among all these programmable solutions is that the software programming model representative of today's market often focuses on extrapolating to support more applications more efficiently rather than making the execution model and underlying hardware architecture hardware-independent. What is common among all these programmable solutions is that the software programming model representative of today's market often focuses on extrapolating to support more applications more efficiently rather than making the execution model and underlying hardware architecture hardware-independent. What is common among all these programmable solutions is that the software programming model representative of today's market often focuses on extrapolating to support more applications more efficiently rather than making the execution model and underlying hardware architecture hardware-independent. What is common among all these programmable solutions is that the software programming model representative of today's market often focuses on extrapolating to support more applications more efficiently rather than making the execution model and underlying hardware architecture hardware-independent. What is common among all these programmable solutions is that the software programming model representative of today's market often focuses on extrapolating to support more applications more efficiently rather than making the execution model and underlying hardware architecture hardware-independent.
[0014] OpenCL supports writing kernels using the ANSI C programming language with some restrictions and additional considerations. OpenCL does not permit the use of function pointers, recursion, bit fields, variable length arrays, and standard header fields. The language is extended to support parallel processing with vector types and vector operations, synchronization, and functions for working with work items / groups. An application programming interface (API) is used to define the platform and then control the platform. OpenCL supports parallel computing using course-level task-based parallel processing and data-based parallel processing. ing interface, API) to define the platform and then control the platform. OpenCL supports parallel computing using course-level task-based parallel processing and data-based parallel processing. and data-based parallel processing to support parallel computing.
[0015] Traditional approaches to developing software applications for parallel execution on multi-processor systems generally require a trade-off between ease of development and efficiency of parallel execution. In other words, generally, the easier the development process is for a programmer, the less efficient the resulting executable program will be when run on the hardware, and conversely, to execute more efficiently, generally the programmer has to make a considerably greater effort, that is, to avoid inefficient processing, the program has to be designed in more detail and use the efficiency-enhancing features of the target hardware, which was the case. effort, that is, to avoid inefficient processing, the program has to be designed in more detail and use the efficiency-enhancing features of the target hardware, which was the case. effort, that is, to avoid inefficient processing, the program has to be designed in more detail and use the efficiency-enhancing features of the target hardware, which was the case. That is the fact.
Prior Art Documents
Patent Documents
[0016]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0017] Therefore, improved systems and methods are desired for facilitating software descriptions from the perspectives of application and system levels, targeting execution models and the underlying hardware architectures, and promoting the use of software programming models and subsequent software programming models. Through this process, improvements are also desired to provide a mechanism that enables an efficient and programmable implementation of applications. MPI (Message Passing Interface) is a standardized, language-independent, extensible, and portable message passing communication protocol API. The MPI API is intended to provide essential virtual topology, synchronization, and communication functionality among a set of processes (mapped to nodes / servers / computer instances) in a language-independent manner using language-specific syntax (bindings). The MPI API standard defines the core syntax and semantics of library routines that include, but are not limited to, support for point-to-point communication, collective communication / broadcast communication send and receive operations, and synchronization of processes, which can define various behaviors. MPI is currently the dominant model used in high-performance computing. At the system level, further significant progress in obtaining higher performance per watt of power dissipation is in a dense communication state. Many processing elements, distributed high-speed memory, and more sophisticated software development tools that divide the system into a hierarchy of modules are possible. At the bottom of the hierarchy, there are tasks assigned to the processing elements, supporting memory, and flexible communication paths across a dynamically configurable interconnect network. If used, it is possible. At the bottom of the hierarchy, there are tasks assigned to the processing elements, supporting memory, and flexible communication paths across a dynamically configurable interconnect network. There are tasks assigned to the processing elements, supporting memory, and flexible communication paths across a dynamically configurable interconnect network. There are tasks assigned to the processing elements, supporting memory, and flexible communication paths across a dynamically configurable interconnect network.
Means for Solving the Problem
[0018] Disclosed are various embodiments for a multiprocessor integrated circuit including a plurality of message nodes. Generally speaking, the plurality of message nodes are connected in an array scattered among a plurality of processors included in the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. The specific message node is also configured to modify the routing information of the first message based on the different message node, generate a second message, and send the second message to the different message node. Disclosed are various embodiments for a multiprocessor integrated circuit including a plurality of message nodes. Generally speaking, the plurality of message nodes are connected in an array scattered among a plurality of processors included in the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor. A specific message node among the plurality of message nodes is configured to receive a first message including a payload and routing information, and select a different message node among the plurality of message nodes based on the routing information and the operation information of the multiprocessor.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
DETAILED DESCRIPTION OF THE INVENTION
[0020] While the present disclosure is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the drawings and detailed description thereto of the embodiments are not intended to limit the present disclosure to the particular forms shown, but on the contrary, the intention is to cover all modifications, equivalents, and alternative forms falling within the spirit and scope of the present disclosure as defined by the appended claims. Headings used herein are for the purpose of organization only and are not intended to limit the scope of the specification. Throughout this application, the word "may" is used in a permissive sense (i.e., having the potential to) rather than the mandatory sense (i.e., must). Similarly, the words "include", "including", and "includes" mean including but not limited to. Flowcharts are provided to illustrate representative embodiments and are not intended to limit the present disclosure to the particular steps shown. In various embodiments, some of the elements of the methods shown may be performed simultaneously, performed in a different order than shown, or omitted. Additional elements of the methods may also be performed as desired. With respect to various units, circuits, or other components, one or more tasks are to be performed. It should be understood that the present disclosure is intended to cover all modifications, equivalents, and alternative forms falling within the spirit and scope of the present disclosure as defined by the appended claims. The headings used herein are for the purpose of organization only and are not intended to limit the scope of the specification. Throughout this application, when the word "may" is used, it is used in a permissive sense (i.e., having the potential to) rather than in a mandatory sense (i.e., must). Similarly, the words "include", "including", and "includes" mean including but not limited to. That is, it does not mean an obligatory meaning (i.e., must), but a permissive meaning (i.e., having the possibility to do). Similarly, the words "include", "including", and "includes" mean including but not limited to. That is, it means including but not necessarily limited to.
[0021] Flowcharts are provided to illustrate representative embodiments and are not intended to limit the present disclosure to the particular steps shown. In various embodiments, some of the elements of the methods shown may be performed simultaneously, performed in a different order than shown, or omitted. Additional elements of the methods may also be performed as desired. That is, some of the elements of the illustrated method may be performed simultaneously, performed in a different order than shown, or omitted.
[0022] With respect to various units, circuits, or other components, one or more tasks It may be described as being “configured to” perform. In such a context, “configured as” generally refers to a comprehensive detailed description of the structure that “has a circuit” for performing one or more tasks during operation. Thus, a unit / circuit / component can be configured to perform a task even when it is not currently operating. Generally, a circuit forming a structure corresponding to “configured as” may include a hardware circuit. Similarly, for various units / circuits / components, for the sake of convenience of description, it may be described as performing one or more tasks. Such a description should be interpreted as including the phrase “ configured as”. Describing a unit / circuit / component configured to perform one or more tasks is not explicitly intended to cite the interpretation of Paragraph 6 of 35 U.S.C. § 112 with respect to that unit / circuit / component. More generally, the description of any element is not explicitly intended to cite the interpretation of Paragraph 6 of 35 U.S.C. § 112 with respect to that element unless the terms “means for” or “step for” are specifically described.
[0023] Detailed Description of Embodiments Referring to FIG. 1, a block diagram illustrating one embodiment of a multi-processor system (MPS) is depicted. In the illustrated embodiment, MPS10 includes a plurality of processor elements ( PEs) and a dynamically configurable communicator or dynamically configurable communication element that is interconnected to communicate data and instructions with each other and is referred to as such. A plurality of data memory routers (DMR) that may be exposed are included. As used herein, a PE may also be referred to as a PE node and, in some cases, a DMR may also be referred to as a DMR node.
[0024] FIG. 2 shows a data path diagram for an embodiment with dual / quad processing elements (PEs) and their local auxiliary memory (SM) for data. The upper left corner of FIG. 2 shows an address generator for the data SM (data RAM), and the upper right shows a register file and several control registers for the data memory router (DMR). The DMR is a node within the primary interconnection network (PIN) between PEs. Large-scale MUXes are used to switch different data sources into the main input operand registers A, B, and C. Another large-scale MUX switches the operand data to the X and Y inputs of the arithmetic pipelines DP1 and DP0. A third large-scale MUX switches the arithmetic pipeline outputs Z1 and Z0 to data path D and returns them to the register file or data RAM. The data RAM is shared with neighboring PEs, and access conflicts are arbitrated by hardware. The control of the pipeline is obtained from the instruction decode block of FIG. 3. Referring to the address generator of FIG. 4, the programmable arithmetic unit includes a sophisticated address calculation that may start several operations before actually using the address. The address generator shown in FIG. 4 includes three integer addition units (AGU0, AGU1, and AGU2) and a multiplier unit (MULT) to support efficient address calculations. The address generator also includes a register file (AGU RF) to store intermediate results during address calculation. The address generator is connected to the data SM (data RAM) through a multiplexer (MUX). The control of the address generator is obtained from the instruction decode block of FIG. 3.
[0025] Referring to the address generator of FIG. 4, the programmable arithmetic unit includes a sophisticated address calculation that may start several operations before actually using the address to support efficient address calculations. The address generator shown in FIG. 4 includes three integer addition units (AGU0, AGU1 and AGU2), a general-purpose integer ALU (GALU), and blocks for "iterative H / W (hardware) and Auto Incrementer". ware) and Auto Incrementer The registers support nested loops using up to eight indexes with independent base and stride values. Additional general-purpose registers support the GALU and iterative units. The output multiplexer supports routing any computed address to the address ports of A, B, or D of the data RAM. Support is provided for routing any computed address to the address ports of A, B, or D of the data RAM.
[0026] Conditional execution of instructions is supported by the execution registers and predicate flags shown in the center of Figure 2. The execution of instructions depends on the execution state, and the predicates are of the prior art. If every instruction were to end after its own number of clock cycles every time, conditional execution would likely not be beneficial. However, in many cases, important instructions can end within fewer clock cycles and provide the results that the next instruction or multiple instructions require. If these next instructions are conditioned to wait for a status bit, they can be started more promptly in the above cases. The predicates are not just for waiting but can be several bits more generally used by conditional instructions for selection and branching as well. Also, some instructions are used to set / clear the predicate value. In various embodiments, the PE may support a program that mixes two types of instructions, namely 64-bit and 128-bit. The shorter instructions support the assembly language programming model as shown on the left side of Figure 5 below. This is the traditional way The predicates can be several bits more generally used by conditional instructions for selection and branching as well. Also, some instructions are used to set / clear the predicate value. Use some instructions to set / clear the predicate value.
[0027] In various embodiments, the PE may support a program that mixes two types of instructions, namely 64-bit and 128-bit. The shorter instructions support the assembly language programming model as shown on the left side of Figure 5 below. This is the traditional way is useful for supporting the code and a simpler compiler. Longer 128-bit instructions support the "HyperOp" programming model shown on the right side of FIG. 5 . Longer instructions are necessary to more precisely control the dual data path hardware and make the dual data path hardware more efficient for signal processing, resulting in improved performance with respect to a given power dissipation (Pdiss). However, the programming needs to be more sophisticated.
[0028] In FIG. 2 of the PE architecture, (in contrast to hardware-assisted scheduling in other architectures) the detailed scheduling of the operations within each pipeline is defined by the program compiler. However, the PE instructions are designed for conditional execution and the conditions for execution depend on the execution state and the register values of the predicates. In some cases , six predicates appear in pairs, i.e., two pairs in the data path and one pair in the address generator . A single conditional instruction can access a single pair of predicates, but over more instructions , all predicates can be accessed. In some embodiments, conditional execution can be used to optimize the performance of the PE in the context of a dual pipeline or a multi-processor IC . Conditional execution may improve the average speed / power ratio (also referred to as the "speed / Pdiss ratio") based on the detailed flow structure of the algorithm of each application in various embodiments .
[0029] In various embodiments, the PEs included in the embodiment of FIG. 1 may include the following features : ● Two data paths, each capable of (per cycle): ○One / two 16×16 multiplications, or one 32×16 multiplication ○One / two 16-bit additions / subtractions, or one 32-bit addition / subtraction ○40-bit barrel shift ○32-bit logical operations ●40-bit accumulation, two 40-bit accumulators ○(Per cycle) Data paths can be executed together: ○One 32×32 multiplication or multiply-accumulate operation ●One 32-bit floating-point addition / subtraction / multiplication ○Three address generation units (Address Generation Unit , AGU) ○Three loads: srcA (source A), srcB (source B), srcC (source C) ○Two loads and one store: srcA, srcB, dstD (destination D) ○Eight base registers and eight index registers ●GP register file ○Three 16×32-bit registers or eight 64-bit registers accessible as 2×16-bit registers ●Instruction decoding: ○64-bit legacy assembly instructions ○128-bit HyperOp instructions ○IM provides 128 bits / cycle at any 64-bit alignment ●Loop hardware ○Zero-overhead looping ○Supports three levels of nesting using three primary index registers to ○Auto-increment of four secondary base / index registers ●Loop buffer ○Reduces instruction fetch power during one or more inner loops
[0030] An iteration loop incorporated into the design to provide repetition of small sections of code Hardware exists. This hardware may include a more efficient branching function than the execution of index counters, increment / decrement logic, completion tests, and software instructions that perform these "overhead" functions. When properly implemented, this hardware eliminates the instruction cycles for performing overhead functions. Directly program a hardware state machine that loops without software instructions for overhead functions using a REPEAT instruction to provide zero-overhead looping with up to 3 levels of nesting. Automatically manage indexing so that additional instructions are not normally required within the loop to manage operand address calculations. This allows multiple arrays to be accessed and managed within the loop without the overhead of additional instructions, saving power and providing better performance. In various embodiments, the iteration loop hardware may include: ● Eight base registers B0 - B7 ○ B0 yields a value of zero in the addressing mode ○ B0 is used as a stack pointer (SP-relative addressing mode) ● Eight index registers I0 - I7 ○ I0 yields a value of zero in the addressing mode ○ I0 can be used as another temporary register for AGU arithmetic (this register is called GR0 instead of I0 in the register map) ● Seven stride registers S1 - S7 ○ Sn is used with In or Bn ● Hardware support for 3 levels of iteration loops ○ The primary loop indices are I1, I2, and I3 ● Four additional increments for the secondary index or base register ○ Index registers I4 through I7 ○ Base registers B4 through B7 ○ Increments by stride registers S4 through S7 ○ Starting address / temporary registers T4 through T7
[0031] The repeat loop is controlled by the REPEAT instruction: ● REPEAT is similar to the previous HyperX generations and has the following improvements: ● Primary loop indices I1, I2, and I3 ● Option to select up to four base registers / index registers, i.e., I4 / B4, I5 / B5, I6 / B6, I7 / B7, which are incremented at the end of the loop ● Repeat loop information is loaded into the loop registers before the label that defines the loop instruction
[0032] The repeat buffer is an instruction FIFO for holding instructions with repeat loops. Its purpose is to reduce instruction fetch power consumption during the most time-consuming section of the code . The distribution of instructions to the buffer is determined by the HyperX tool at compile time and is not left to the user. This document only describes for the purpose of providing the user with a basic understanding. The main features of the repeat buffer may include the following: ● Determine a group of instructions by the REPEAT instruction and its label ● The usage of the repeat buffer is determined at compile time and is indicated by a flag within the REPEAT instruction ● The first instruction of every repeat loop is always loaded into the repeat buffer for performance and power reasons ● The buffer can hold 64-bit or 128-bit instructions. ● Up to 12 64-bit entries are available. For 128-bit instructions, two entries are used. ● The buffer is used for anything other than putting the first instruction of a loop into the buffer. Therefore, the entire loop must fit into the buffer.
[0033] The iteration hardware uses the primary indexes (I1~I3) and other related control registers to control the loop operation. In addition to the primary hardware, another set of registers that can be automatically managed by the iteration hardware for additional address calculation by the AGU exists. These special registers are as follows: ● B4~B7 - Four additional base registers. ● I4~I7 - Four additional index registers. ● S4~S7 - Four additional stride registers. ● T4~T7 - Four additional registers used to initialize base registers or index registers.
[0034] There are four additional adders available for performing addition on these registers. These adders can be controlled by instructions (INIT and INCR), or by the automatic increment feature of the hardware. Using the AUTOINC register described elsewhere in this specification, each primary REPEAT operation is tied together, and furthermore, address addition can be performed on one or more index registers or base registers.
[0035] Using each adder, for any primary index (I1~I3), for each loop for each iteration, a given stride (S4~D7) can be added to the same numbered base (B4~B7) or to the same numbered index (I4~I7). Additionally, whenever a starting value is to be loaded into the primary index at the beginning of a loop instruction, the same numbered T register (T4~T7) is loaded into the indicated AUTOINC BASE or INDEX. This allows multiple arrays to be accessed and managed within a loop without the overhead of additional instructions, saving power and providing better performance. In various embodiments, conditional execution may be based on a predicate flag. Such flags may include:
[0036] ● P0~P3: ○ Set by the DP test instruction ○ Set according to the timing of DP ● GP0 and GP1 ○ Set by the AGU test instruction (an example is shown in Figure 6) ○ Set according to the timing of AGU
[0037] Predicate flags are set using instructions of the TEST class that perform the following: ● Execute the TEST operation ● Verify the resulting condition ● Set the selected predicate flag
[0038]
Table 1
[0039] Conditional instructions specify a test for a pair of predicate flags. For example: ● GP0, GP1 - Used by the AGU instruction P0, P1 - Used by the DP instruction, typically DP0 P2, P3 - by DP instruction, typically used in DP1
[0040] An example of testing the predicate flags is illustrated in Figure 6. In addition, the conditional instructions of the DP, Conditional instructions and program flow instructions are illustrated in FIG.
[0041] The conditional block command is shown in FIG. 8. The command shown in FIG. 8 is a simplified version of the actual operation. The STARTIF, ELSE, and ENDIF commands can be nested, so There is a condition stack that holds the stored condition state. STARTIF pushes a new condition onto the stack, and ELSE pushes the current condition state (top of the stack) Toggles and ENDIF pops the condition stack. The current condition state is STARTIF. The actions of F, ELSE, and ENDIF may be prohibited.
[0042] You may execute a Hyper-Op in a variety of ways. Here is an example of a Hyper-Op execution: Examples are shown in Table 3.
[0043] [Table 2] TIFF0007698327000003.tif121170
[0044] The GPn will be ready in the next cycle, so the GTEST instruction can be used to start the GP If the n bit is set, no branch prediction is required at all. However, if the GPn bit is If a MOV is written from a general register, the branch prediction is delayed and the correct Branch prediction is performed. Pn is ready after 5 cycles and therefore does not require branch prediction. be desired. If n is the number of instruction cycles between a test instruction and a branch, the penalty due to a prediction miss is 5 - n cycles. If the test instruction can be moved forward in the code stream, n can be increased and the penalty due to a prediction miss can be reduced to perhaps zero (0) cycles or less.
[0045] Predicates are calculated using explicit instructions to set the predicate and are not modified by other instructions, so often it is possible to schedule the code to greatly reduce any penalty associated with a mispredicted branch. Branch prediction may be done statically, determined at compile time, based on industry standard heuristics regarding branch probability.
[0046] Hyper - Op mode may enable encoding of instructions, in which case each separate part of the data path is controlled by part of the instruction encoding. This allows more direct control of the hardware's parallel processing. A 128 - bit Hyper - Op format enables the parallel processing depicted in Table 4.
[0047]
Table 3
[0048] There are constraints under which HyperOp instructions can be executed in parallel on DP0 and DP1. Two HyperOp instructions can be executed in parallel if they have the same latency. By definition, the slots of DP0 and DP1 can always execute the same instruction in parallel (equivalent to SIMD). There are a few exceptions. Only a single FP instruction... When using hardware from both data paths in the calculation of both DP slots, it can be run in both DP slots. It supports the SIMD form of executing the same instruction. However, on the other hand, note that the usage model is much more flexible in that it allows any two instructions with the same latency to be executed in parallel.
[0049] Address instructions occur during the FD pipeline stage, take 1 cycle, and the result is available in the next cycle for use by all load / store instructions. In various embodiments, auto-increment and iteration include reloads to reduce overhead.
[0050] Each DMR may have a direct memory access (DMA) engine to support multi-point sources and multi-point destinations simultaneously. Further, the complete state of each DMA engine can be captured and stored in memory, and this state information can be retrieved later to resume DMA operation if the DMA operation is interrupted. The ability to save the state of the DMA configuration requires reading up to 11 DM registers for the PE to obtain the overall state of the DMA. Many of these registers are internal DMA registers that are exposed to external access for the purpose of capturing the state.
[0051] To save register space, the DMA can save its state in a compact form called a descriptor in memory. The PE specifies the starting location of this descriptor, and the DMA and modified push engine read from memory starting at the specified memory address. Disaster data can be written. The push engine is part of the DMA engine used to extend the routed message from one destination DMR to a second destination DMR.
[0052] The push engine already has a state machine that steps through each register in the DMA to program the DMA. This same machine is also used to read the register. The read data then needs to be directed to the in-port module of the adjacent port. The important part is to tie any DMA write function stop to the push engine. This can be done by gating the DMA function stop on the busy input signal of the push engine.
[0053] The DMA wake-up can be used to send a signal to the PE that saved the descriptor. At that point, the PE can freely swap tasks. When the new task is complete, the PE can indicate the saved descriptor and the DMA processing resumes. It is noted that the in-port or out-of-port router needs to be properly configured during the task swap.
[0054] The accumulator memory has an optional right post-shift by a predetermined shift amount. In various embodiments, the following shift amounts exist: ● 1: For taking an average ● 8: ● 16: For storing ACCH = ACC[31:16] ● 15:
[0055] These values are stored in three shift fields used as indirect shift amounts. 3 One field is designated as SHR1, SHR2, and SHR3 in HyperOp syntax and refers to the shift value fields SHIFT_CTL1~3 in the PE_CONFIG register .
[0056] There are two types of accumulator pairs, namely, one accumulator (ACC C2_ACC0, ACC2_ACC1, ACC3_ACC0, ACC3_ACC1) from each DP, and a split of the accumulator treated as 16-bit data of SIMD (ACC0 H_ACC0L, ACC1H_ACC1L, ACC2H_ACC2L, ACC3H_AC C3L). Memory with post-shift of the accumulator pair performs independent shifts for each part of the accumulator pair while having the same bit position numbers. In the following description, the "tmp (temporary)" designation is used to attempt to clarify the semantics and is not an actual hardware register .
[0057]
Table 4
[0058] Shifting may also be done with split memory and split load. More shift operations generally increase the hardware power dissipation (Pdiss). In various embodiments, the shift hardware design may be advanced by selecting the most required shift option for a given Pdiss budget. By analyzing the application code and the PE architecture , the most required shift option may be determined. In some In some embodiments, for example, instead of byte-addressing the memory, word-addressing the memory, all that is necessary may be for byte alignment / shifting from / to word boundaries. In some embodiments, to increase performance, additional auxiliary calculation units may be employed. A list of possible auxiliary calculation units is depicted in Table 6.
[0059] In some embodiments, to increase performance, additional auxiliary calculation units may be employed. A list of possible auxiliary calculation units is depicted in Table 6.
[0060]
Table 5
[0061] HyperOp instructions - Use static scheduling in program compilation to enable individual control of the dual data paths. Run the execution threads through the compiler to statically schedule all operations. In contrast, modern GPP architecture compilers place related instructions together in the machine code, but the detailed operation scheduling is performed by the (power-consuming) hardware. Static scheduling can significantly save Pdiss during execution. Run the execution threads through the compiler to statically schedule all operations. In contrast, modern GPP architecture compilers place related instructions together in the machine code, but the detailed operation scheduling is performed by the (power-consuming) hardware. Static scheduling can significantly save Pdiss during execution. In contrast, modern GPP architecture compilers place related instructions together in the machine code, but the detailed operation scheduling is performed by the (power-consuming) hardware. Static scheduling can significantly save Pdiss during execution. Static scheduling can significantly save Pdiss during execution.
[0062] During data transmission, defects within the system can cause distortion or degradation of the transmission signal. Such distortion or degradation of the transmission signal may result in incorrect data bit values at the receiving circuit. To remove such effects, in some embodiments, instructions for supporting forward error correction (FEC) encoding and decoding are included. FEC is applicable to all types of digital During data transmission, defects within the system can cause distortion or degradation of the transmission signal. Such distortion or degradation of the transmission signal may result in incorrect data bit values at the receiving circuit. To remove such effects, in some embodiments, instructions for supporting forward error correction (FEC) encoding and decoding are included. To remove such effects, in some embodiments, instructions for supporting forward error correction (FEC) encoding and decoding are included. To remove such effects, in some embodiments, instructions for supporting forward error correction (FEC) encoding and decoding are included. FEC is applicable to all types of digital It is applied to other fields such as optical communication and digital recording and playback from a storage medium. The basic idea is to take out a block of arbitrary input data and encode the block using additional parity bits in such a way that bit error correction is possible in the receiving or playback electronic circuit. An encoded block consisting of data and parity bits is called an FEC frame. The FEC frame is further processed by a modulator and then transmitted into a medium (wired, wireless, or storage medium). At the receiving side, a signal is picked up by an antenna or transducer, amplified, demodulated, and sampled by an analog-to-digital converter (ADC).
[0063] The signal in the medium may have been subject to interference, fading, and echoes, and noise may have been added by the receiving side. The output of the ADC is a sequence of digital samples. There are various methods for picking up a sequence of samples, obtaining synchronization, and formatting them into an FEC frame, but these methods are mostly irrelevant to FEC calculations and are not described here. Each bit position in the formatted FEC frame has a digital value, sometimes called a soft bit, represented by the actual bits of an integer in a digital system. FEC decoding is the process of taking out the soft bits in the formatted FEC frame, calculating bit error correction, applying the bit error correction, and outputting the corrected FEC frame. The purpose of the FEC decoding algorithm is to know the method of generating the parity bits in advance.
[0064] As a premise, it is to output the most likely correct data frame. For the FEC to operate correctly, a specific FEC decoding method (using parity bits for error correction) must be coordinated with the FEC encoding method (generating parity bits). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo).
[0065] Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo). Initial success of FEC was achieved using Hamming codes, BCH codes, and Reed - Solomon codes. Further success was obtained using convolutional codes and serial concatenation of convolutional codes with other codes. On the decoding side, the goal is to find the block of data that is most likely to be correct, assuming received soft bits with errors induced by noise. This can be achieved with a single - pass algorithm (such as Viterbi) or an iterative algorithm (such as Turbo).
[0066] The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and The calculation of FEC involves calculating the predicted correctness of two choices for the values of binary bits according to a sequence of observed values obtained from a sampler. Since a sequence of constantly changing values can be treated as a random variable, mathematical processing of probability can be applied. The main concern is whether a specific transmitted data bit was 1 or - 1, assuming the values of the soft bits within the FEC frame. Before making a hard decision regarding the transmitted data, multiple soft decisions can be calculated. These soft decisions can be calculated by comparing probabilities, and a method of comparing probabilities that includes parity bits is to calculate the ratio of conditional probabilities called the likelihood ratio (LR). The logarithm of the LR (logarithm of LR, LLR) can be integrated and Division is converted to addition and subtraction, which is of particular interest because addition and subtraction are computed more quickly and are less likely to cause overflow and underflow in the PE. As a result, FEC decoding can be calculated using the values of the LLRs stored as integers.
[0067] The sum-product of log probabilities is also called the MAX* operator. In various embodiments, the offset instruction may be used to perform the MAX* operator in a manner similar to add-compare-select (ACS). The MAX* operator provides a sum-product type of operation for mathematical processing in the logarithmic domain on exponential probabilities. In many cases, the symbolic form is written as Max* (x0+y0,x1+y1). (x0+y0,x1+y1).
[0068] In various embodiments, the PE implements a function such as: Z[15:0]=MAX((X[15:0]+Y[15:0]),(X[31:16] +Y[31:16]))+TLU offset
[0069] By slightly modifying the operand usage to provide a higher throughput form useful for turbo operation, a double log probability sum-product instruction may be achieved. This one produces two results in a single data path in the form of Max*(x1 +y0, x0+y1):Max*(x0+y0,x1+y1).
[0070] In various embodiments, the PE implements a function such as: Z[31:16]=MAX((X[31:16]+Y[15:0]),(X[15:0 +Y[31:16]))+TLU offset Z[15:0]=MAX((X[15:0]+Y[15:0]),(X[31:16] +Y[31:16]))+TLU offset
[0071] Another form of MAX* operation produces 32 bits per datapath in the following form: Max*(0,x1+y1)-Max*(x1,y1):Max*(0,x0+y0) -Max*(x0,y0)
[0072] In various embodiments, the PE implements functions such as: Z[31:16]=MAX(0,(X[31:16]+Y[31:16]))+TLU offset-(MAX(X[31:16],Y[31:16]))+TLU offset) Z[15:0]=MAX(0,(X[15:0]+Y[15:0]))+TLU offset -(MAX(X[15:0],Y[15:0])+TLU offset)
[0073] Another instruction uses two MAX* registers in each data path for a dual MAX* operation on two operands. Another instruction can provide the value in the accumulator of the By using an accumulator, you can use MAX* for large groups of numbers. Provides a fast way to obtain the result. The two 16-bit results are the sum of the two accumulators. If both data paths are used, when all input data is available, The accumulators run a MAX* operation on their input data and calculate the final result as Symbolically, the equation looks like this: ACC n+1 =Max*(ACC n+1 ,Max*(x1,y1)):ACC n =Max*( ACC n,Max*(x0,y0))
[0074] The following formula may be used to achieve a double cumulative form for the logarithmic probability sum: ACC n+1 =Max*(ACC n+1 ,x0+y0):ACC n =Max*(ACC n ,x1 +y)
[0075] It should be noted that the natural result of the special hardware for the LP instruction is to swap the "0" index data to the upper ACC CC and the "1" index data to the lower ACC. This is preferable if this data is easily swapped within the data path. In various embodiments the PE implements a function as follows: ACC n+1 =MAX(ACC n+1 ,(X[15:0]+Y[15:+1]))+TLU offset set ACC n =MAX(ACC n ,(X[31:16]+Y[31:+16]))+TLU offset set
[0076] It is also possible to generate a double sum of quotients in the logarithmic domain. In a manner similar to what has already been shown subtraction is used instead of addition to provide the following: ACC n+1 =Max*(ACC n+1 ,x0-y0):ACC n =Max*(ACC n ,x1 -y1)
[0077] In various embodiments, the PE implements a function as follows: ACC n+1 =MAX(ACC n+1 ,(X[15:0]-Y[15:0]))+TLU off Set ACC n =MAX(ACC n ,(X[31:16]-Y[31:16]))+TLU offset Set
[0078] To implement the instructions referenced above, a dedicated logic circuit as depicted in FIGS. 9 to 14 may be employed In some cases, the logic circuit may be selectively enabled based on the type of instruction, thereby enabling multiple instructions to be executed using a minimum amount of logic circuitry.
[0079] Within an application that communicates with the chip I / O ports and runs on the PE, or between such applications, a primary interconnect network (PIN) of the MPS is designed for wideband, low-latency data transfer. The primary interconnect network (PIN) is generally described as a set of nodes with links connecting pairs of nodes. The interconnect network (IN) may generally be described as a set of nodes with links connecting pairs of nodes. Most PINs require far more wiring and thus are not fully point-to-point capable in one step. Rather, most PINs are multi-stage with routers at each node of the network, and the nodes are connected to each other via the links. Messages can be routed through the PIN, and the PIN enforces rules for initiating, halting, and delivering messages from a source node to a destination node. Messages may be used indefinitely as data pipes if left open. However, due to being multi-stage, existing messages may block the setup of new messages by occupying links or destinations specified by new messages. set by the new message. It may block the uplink, and as a result, message delivery is not guaranteed. To mitigate this Some things that have emerged in the literature, such as dynamic cut-through and long-distance routes to "jump over" congestion However, the approach of the present inventors is to add interconnected layers, each layer having another set of links. Each additional layer is used to expand the PIN node router so that messages can cross from one layer to another to make it possible.
[0080] In large-scale MPS, it is necessary to manage the system without affecting the operating efficiency of the PIN This has led to the development of a secondary interconnection network (SIN) that may have a lower bandwidth than the PIN in some embodiments but can guarantee message delivery but can guarantee message delivery. Such an interconnection network is shown in FIG. 15. As illustrated, the message bus 1500 includes a plurality of message-bus nodes connected to each other in a fixed connectivity manner. The message bus provides not only a moderate bandwidth, variable latency, and guaranteed delivery method to reach all accessible registers and memory locations within the chip, including both PE and DMR locations, but also I / O controllers such as the I / O controller 1502 Such an interconnection network is shown in FIG. 15. As illustrated, the message bus 1500 includes a plurality of message-bus nodes connected to each other in a fixed connectivity manner. The message bus provides not only a moderate bandwidth, variable latency, and guaranteed delivery method to reach all accessible registers and memory locations within the chip, including both PE and DMR locations, but also I / O controllers such as the I / O controller 1502 bus provides not only a moderate bandwidth, variable latency, and guaranteed delivery method to reach all accessible registers and memory locations within the chip, including both PE and DMR locations, but also I / O controllers such as the I / O controller 1502 bus provides not only a moderate bandwidth, variable latency, and guaranteed delivery method to reach all accessible registers and memory locations within the chip, including both PE and DMR locations, but also I / O controllers such as the I / O controller 1502 bus provides not only a moderate bandwidth, variable latency, and guaranteed delivery method to reach all accessible registers and memory locations within the chip, including both PE and DMR locations, but also I / O controllers such as the I / O controller 1502
[0081] The message bus may be used not only to boot, debug, and load data outside the core array structure but also to access substantially all addressable locations through the MPS device For example, the message bus may be used to access all PE / D For example, the message bus may be used to access all PE / D The memory locations of MR data and instructions, internal PE registers, DMR registers (including register locations of the bus), and I / O peripheral devices attached to the I / O bus can be accessed to.
[0082] In some embodiments, the message bus can provide support for multiple simultaneous masters, such as PEs, development access ports (DAPs) 1503, boot control 1504, and also I / O processors. Messages are routed on the message bus using automatic routing based on the relative positioning of the source and destination. Responses are automatically routed in a manner similar to the request using the relative location. The error response route utilizes the continued location to the source maintained within the message.
[0083] In some embodiments, the message bus may comprise two or more independent message configurations. Situations may arise where independent configuration messages attempt to access the same destination where automatic arbitration is useful. The result of the arbitration may be determined in a simple manner such as a priority structure. The message bus priority structure may consider two distinct priorities established for accessing DMR addresses, namely the lowest or highest, and all accesses to PE addresses have the lowest priority at the PE.
[0084] The message bus supports multiple endpoint message groups to enable a subset of actors to respond to a single message. Multiple group memberships can be set on a single node. In various embodiments, broadcast capabilities are used And, to reach and distribute to all nodes, many independent groups may be available. It may be capable.
[0085] In various embodiments, the message bus may enable multiple chip operations. When implementing multiple die structures, the relative address of the destination may be bridged between chips. In some cases, the message format may permit up to 256 MBN nodes in the X and Y directions. In other embodiments, the message can be extended to support additional nodes. By adopting a table (e.g., Table 1807) inheritance technique, any suitable number of message nodes may be supported. When implementing multiple die structures, the relative address of the destination may be bridged between chips. In some cases, the message format may permit up to 256 MBN nodes in the X and Y directions. In other embodiments, the message can be extended to support additional nodes. Using a table (e.g., Table 1807) inheritance technique, any suitable number of message nodes may be supported. By adopting a table (e.g., Table 1807) inheritance technique, any suitable number of message nodes may be supported.
[0086] The message bus has the ability to enable a processor within the device to reach any addressable location within the device. This ability enables passing messages between processors, updating a table of values as the algorithm progresses, managing the behavior of remote I / O controllers, collecting runtime statistical information, managing cell security, and enabling various realizations including general communication of information where time is not critical between processors. This ability enables passing messages between processors, updating a table of values as the algorithm progresses, managing the behavior of remote I / O controllers, collecting runtime statistical information, managing cell security, and enabling various realizations including general communication of information where time is not critical between processors. updating a table of values as the algorithm progresses, managing the behavior of remote I / O controllers, collecting runtime statistical information, managing cell security, and enabling various realizations including general communication of information where time is not critical between processors. managing the behavior of remote I / O controllers, collecting runtime statistical information, managing cell security, and enabling various realizations including general communication of information where time is not critical between processors. managing cell security, and enabling various realizations including general communication of information where time is not critical between processors. including general communication of information where time is not critical between processors.
[0087] It is noted that the message bus may lack certain features that make it undesirable to use as a special layer of the PIN routing configuration. First, the bandwidth is much lower. For example, in some implementations, the message bus may be up to 10 times slower than the PIN, while in other implementations, it may be only 2 times slower than the PIN. It is noted that the message bus may lack certain features that make it undesirable to use as a special layer of the PIN routing configuration. First, the bandwidth is much lower. For example, in some implementations, the message bus may be up to 10 times slower than the PIN, while in other implementations, it may be only 2 times slower than the PIN. while in other implementations, it may be only 2 times slower than the PIN. There is. Second, the waiting time for data delivery varies greatly even for messages between the same source and destination pairs. There is no concept of programmed route setup and teardown. In this case, within the configuration, a route of known length is set up for the message, and each time the route is used, the same wiring is traversed to connect the two endpoints, resulting in a predictable waiting time for data delivery. When using a message bus, a relatively short message is routed from the source to the destination using a route determined by the configuration hardware. If the message is blocked at some point along the route, it waits for other messages blocking it to complete and then continues. When using only one message at a time in the configuration (with no priority delay at the destination), data delivery by messages may show a predictable waiting time. However, additional message traffic in the configuration may disrupt data delivery and may change the route taken by each subsequent message. As a result, the arrival time is not guaranteed, so the message bus may not be suitable for distributing synchronous messages. The message bus is useful for guaranteeing the delivery of shorter, lower-bandwidth messages to any destination within the MPS with high power efficiency. These messages consume a significant amount of resources within the PIN configuration, potentially tying up the link for an extended period, and data hardly passes through or the link cannot block the system, so a non-stop setup and teardown of the link is required. The message bus also simplifies the management of remote processors for chip I / O. In the MPS, in this case, these are near the I / O ports.
[0088] only the processor controls the port and any peripheral devices attached to the port well.
[0089] Although not suitable for applications where timing is critical, the message bus still makes considerable performance available. The bus itself moves one word per clock, and the clock may be the same as the functional core clock which may have a target frequency of 1 500 MHz. This clock effectively results in a transfer of 1500 MHz words / second between nodes. The message bus is designed to push data and address across the entire bus for each word delivered to a register and then release the bus to other messages as quickly as possible, so there is a unique overhead necessary to define the route to identify where to read or write a word and how to return data or status to the node that is making the request. These non-data words reduce the bus throughput in a single transaction. To limit the impact of the message structure overhead, any number of words can be transferred in one message, with the only restriction being that those words must be contiguous from a single starting address.
[0090] Under normal conditions, access to any normally accessible location within the MPS device is available. This means that any register, memory location, or peripheral device with a normal mode address can be written to or read from within the parameters accessing a particular address. In some PE internal registers, P The contents of the PE can be read while E is operating, however, the values included represent a snapshot at the time the read was taken and are updated only when a value is requested. In addition, there is a time interval between when a request is generated and when a value is read from the PE or DMR by the message bus and when the result is returned to and delivered by the requester, and this time interval represents that the time to wait for the result can be quite long depending on the load on the system and the message bus. For almost all clocks, there is a possibility of waiting excessively to access certain internal PE registers required for operation, effectively halting the request until the PE stops at a breakpoint. Gaps may appear in the code and it may become possible to read these registers , but the PE needs to operate some registers and the messages on the message bus generally result in delaying the message halt by attempting to read these registers at the default low priority. Inside the DMR, since the access priority is programmable, it is possible to make the request wait until no other requests for that address region are pending at all, or to access that address immediately and block all other requests attempting an access to the same region as the request. The normal mode address locations may include the following: ● Read / write access to any location in the DMR data memory ● Read / write access to any DMR memory-mapped register ● Read / write access to any location in the PE instruction memory
[0091] ● Read / write access to any location in the DMR data memory ● Read / write access to any DMR memory-mapped register ● Read / write access to any location in the PE instruction memory Read / write access to PE state and control registers Read / write access to clock control registers - Breakpoint control except for hardware break insertion ●PE wake-up control Parity Control ●PE message delivery PE programmer register access Memory and peripherals on the IO bus
[0092] While running a program, great care must be taken when operating on instruction memory. A block of instructions can be written to memory and As the program runs, new pieces of code are added to blocks of code. may be executed without first fully substituting, resulting in unpredictable behavior. The MPS includes a parity bit for each write to memory and The read operation may be configurable to check parity and if an error is detected, However, parity checking is much more efficient than not having it. Parity checking in MPS is done by dividing memory with and without parity into separate parts. This is an operating mode that should be used in extreme environments. You can switch between these modes while running the applications that use that memory. It is not wise to do so.
[0093] Access to the clock control is possible under all conditions, however It is not always wise to change the state of the lock register. While it is working, especially when working on a data set that is shared among multiple processors, When present, clock control changes are made locally without regard to other nodes, and other nodes may also be further accessed to update clock control at those locations. If an attempt is made to change the clock configuration while running an algorithm, the timing of data access is likely to be lost.
[0094] When a PE stops at a breakpoint, additional access to the message bus is available. When a PE is interrupted, the program counter stops and the hardware breakpoint can be updated. All of the normal mode capabilities are available, and the hardware breakpoint insertion ability is additionally available.
[0095] Due to the implementation of breakpoints, changes made by the system while it is operating can lead to unpredictable results, including missed interrupts and unintended interrupts. Therefore, breakpoint changes are most effectively made while the program execution is stopped.
[0096] When a PE stops at a breakpoint, the internal register access time is improved and the return value to the stopped processor remains accurate. Arbitration for PE register access has no higher priority requesters while active and allows the debug system to access the internal state registers of the PE more quickly. Similarly, on the DMR, after completion of a DMA operation, there are no competing accesses for the address and even the lowest priority requests are immediately satisfied.
[0097] During booting, the processor is loaded for the first time using the message bus, and the clock and security are configured, and the PE can be released from reset and start operating. During the boot operation, most of the transactions on the message bus are expected to originate from the boot processor using a destination that spans the entire device. Long burst operations are dominant, and it is expected to reduce the program load overhead caused by addressing.
[0098] After that, one use of the boot controller is to implement dynamic cells, in which case it is possible to load new cells into an already operating system. When used and defined in this specification, a cell is a part of an application assigned to one or more PEs and one or more DMRs. It should be noted that at least one DMR is included in the cell to serve as the associated PE instruction memory. In this case, there is a high possibility of more activity on the message bus, but in this case as well, since it is already around the device, arbitration is simplified and the new cell is transmitted into the array. By utilizing larger block transfers, the time to load dynamic cells can be minimized. Different from the initial load, there is a high possibility of contention at some point during the replacement cell load. Since the dynamic cell load potentially consumes a long path and leads to delays in delivering other messages, the full length of the burst should be considered when implementing the dynamic cell load. One problem common to in-system debuggers is the conflict between the debug facility and the functional
[0099] operation of the system. There may be an interaction. In certain cases, the possibility of interaction may change the behavior of the functional system when engaged in debugging, or may even result in the correction or change of errors when debugging is in operation. Any access that must be mediated cannot completely eliminate the interaction between the functional system and the parallel debugging facility, but if the debugging operation is mapped into a separate message bus instance, this mapping can eliminate all interactions except for the final data access mediation. By carefully selecting the debugging with the lowest priority, the debugging interacts only with the system while not being used in other ways and does not disrupt the functional accesses generated from the functional system. In various embodiments, the priority may be changed between high and low. When engaged in debugging, it may change the behavior of the functional system, or may even result in the correction or change of errors when debugging is in operation. Any access that must be mediated cannot completely eliminate the interaction between the functional system and the parallel debugging facility, but if the debugging operation is mapped into a separate message bus instance, this mapping can eliminate all interactions except for the final data access mediation. Any access that must be mediated cannot completely eliminate the interaction between the functional system and the parallel debugging facility, but if the debugging operation is mapped into a separate message bus instance, this mapping can eliminate all interactions except for the final data access mediation. By carefully selecting the debugging with the lowest priority, the debugging interacts only with the system while not being used in other ways and does not disrupt the functional accesses generated from the functional system. When engaged in debugging, it may change the behavior of the functional system, or may even result in the correction or change of errors when debugging is in operation. By carefully selecting the debugging with the lowest priority, the debugging interacts only with the system while not being used in other ways and does not disrupt the functional accesses generated from the functional system. By carefully selecting the debugging with the lowest priority, the debugging interacts only with the system while not being used in other ways and does not disrupt the functional accesses generated from the functional system. In various embodiments, the priority may be changed between high and low. In various embodiments, the priority may be changed between high and low.
[0100] When the processor is at a breakpoint, there is no PE that generates a request to be delivered to the DMR. This does not mean that there are no requests in the DMR, as the DMA requests can be continuously processed while the PE is stopped. This results in a split state where the PE's requests are immediately satisfied and a DMR state where the DMA transaction continues to be before the debug request. Logically, this supports the idea that debugging should not interfere with the operation in a processor without breakpoints. When the processor is at a breakpoint, there is no PE that generates a request to be delivered to the DMR. This does not mean that there are no requests in the DMR, as the DMA requests can be continuously processed while the PE is stopped. When the processor is at a breakpoint, there is no PE that generates a request to be delivered to the DMR. This does not mean that there are no requests in the DMR, as the DMA requests can be continuously processed while the PE is stopped. When the processor is at a breakpoint, there is no PE that generates a request to be delivered to the DMR. This does not mean that there are no requests in the DMR, as the DMA requests can be continuously processed while the PE is stopped. When the processor is at a breakpoint, there is no PE that generates a request to be delivered to the DMR. This does not mean that there are no requests in the DMR, as the DMA requests can be continuously processed while the PE is stopped. When the processor is at a breakpoint, there is no PE that generates a request to be delivered to the DMR. This does not mean that there are no requests in the DMR, as the DMA requests can be continuously processed while the PE is stopped.
[0101] Before exploring the details of the bus itself, in relation to the message bus, what does the message mean? It is helpful to understand this first. In the most general sense, a message requires a means to deliver the message to the intended destination, the data to be delivered, and a means to return a response to the starting point. There are several different messages that the message bus passes through its network, which will be addressed next. The message is constructed and sent by programming the configuration registers inside the message bus node. There are two sets of these registers for the two channels (A and B) of the message bus. The programming of these registers will be discussed below. There are several available message formats. These can be categorized as follows:
[0102] ● Point-to-Point Message - Allows information to be read from, or written to, any other single node. ● Multi-Point Message - Allows a single message to be read from, or written to, a group of end-point nodes. ● Response Message - Not directly generated by the user. Used by the message bus to provide a positive response to other messages. ● Secure Configuration Message - A form of message used to configure the security for the chip. To send a message, the program must configure the basic components of the message into the configuration registers and then program the signals for the message to be sent. The programming components are listed in Figure 16.
[0103]
[0104] Use the STATUS register to observe the status of the sent message. Before sending the next message, in addition to these registers that directly control the sent message, there are several other configuration registers, described later, used to control other aspects of the message bus system. Note that only the registers that need to be modified to define a new message need to be updated before sending the next message. For example, to send the same message to five locations, simply update the routing information in DELTA_OFFSET and resend the message using GEN_MSG. The message format is described more fully below.
[0105] The most basic message that can be used by any master to reach any location within the chip is the point-to-point message. As its name implies, this message targets a single location and is issued from a single location. None of the intermediate locations have any means to snoop the passing data, so the information passed between two nodes is only visible outside the message bus by the two nodes, and in this regard, all point-to-point messages are secure. A variety of options are available to build the message to balance the capabilities and overhead for this message type.
[0106] Since the slave can receive and process only one message at a time, it does not need to know which node is requesting data access, and only the route back It is necessary, and as a result, a response can be delivered. Part of the request message is the return path for the response, which is necessary to complete the round-trip of the point-to-point message. It includes.
[0107] The point-to-point message can be a read request or a write request. The read request generates a response message that includes the requested read data, and the write request generates a response indicating the success or failure of the write operation. The message of the read request or write request is very similar to balancing the capabilities and performance. The response message also sacrifices some flexibility at the master to minimize the overhead. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. The write request generates a response indicating the success or failure of the write operation. The read request or write request message is very similar to balancing the capabilities and performance. The response message also sacrifices some flexibility at the master to minimize the overhead. The read request or write request message is very similar to balancing the capabilities and performance. The response message also sacrifices some flexibility at the master to minimize the overhead. The read request or write request message is very similar to balancing the capabilities and performance. The response message also sacrifices some flexibility at the master to minimize the overhead. The read request or write request message is very similar to balancing the capabilities and performance. The response message also sacrifices some flexibility at the master to minimize the overhead.
[0108] For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed. For each point-to-point read request or write request, there is one response message. If multiple data words are included, multiple response words are included. As a result, any address sent will return a response indicating the write status or data for reading. The data within the message body is returned in the same order as when the request was sent. To ensure that the data is quickly removed from the bus when it arrives back at the requested node, the address to be stored separately for the response should be programmed into the MBN. Since only one response location can be stored at a time, if more than two words are expected on the return and the automatic memory loading mechanism is used, each MBN may have one unprocessed transaction at a time. If the processor extracts all the data from the node, there may be as many requests as the processor desires left unprocessed.
[0109] Use the same response message format for all of the multiple endpoint responses However, in these multiple endpoint responses, a single response word is inserted into the payload For example, in a read message, the response may include the value of the requested address or a security error code if an attempt is made to read an address that is not valid for the safe area. Alternatively, in a write message, the response may include a success or failure value indicating whether the write attempt made by the request was carried out or a security error code if an attempt is made to read an address that is not valid for the safe area For example, in a read message, the response may include the value of the requested address or a security error code if an attempt is made to read an address that is not valid for the safe area. Alternatively, in a write message, the response may include a success or failure value indicating whether the write attempt made by the request was carried out or a security error code if an attempt is made to read an address that is not valid for the safe area
[0110] Two or more active nodes receive multiple endpoint write messages. These messages are useful for communication within the cell, in which case the cell can be instructed which message to respond to through the message bus component address write. Security may prevent writes from occurring, and since individual word states can potentially cause interference, a single write state is returned for the entire message rather than for individual word states. Multicast messages are distributed through the array, so the response address is recorded as the calculated delta offset from the requesting node. This results in using many paths to return the response message to the master, and many do not follow the same path as expected for the request. An example of a multiple endpoint message is a broadcast message that addresses all nodes at once These messages are useful for communication within the cell, in which case the cell can be instructed which message to respond to through the message bus component address write These messages are useful for communication within the cell, in which case the cell can be instructed which message to respond to through the message bus component address write Security may prevent writes from occurring, and since individual word states can potentially cause interference, a single write state is returned for the entire message rather than for individual word states Security may prevent writes from occurring, and since individual word states can potentially cause interference, a single write state is returned for the entire message rather than for individual word states Multicast messages are distributed through the array, so the response address is recorded as the calculated delta offset from the requesting node. This results in using many paths to return the response message to the master, and many do not follow the same path as expected for the request Multicast messages are distributed through the array, so the response address is recorded as the calculated delta offset from the requesting node. This results in using many paths to return the response message to the master, and many do not follow the same path as expected for the request This results in using many paths to return the response message to the master, and many do not follow the same path as expected for the request An example of a multiple endpoint message is a broadcast message that addresses all nodes at once An example of a multiple endpoint message is a broadcast message that addresses all nodes at once
[0111] Also, the same address can be read from a collection of message bus nodes This may be useful in some cases. In these cases, multiple endpoint reads are available. The operation functions such that only nodes that match the multi-node address respond. Similar to other multiple endpoint messages, the response path is determined by the delta offset calculated along the way to the responding node. The response returns to the requesting node following several routes, many of which are different from the path taken from source to destination. Also, all nodes may
[0112] respond and return one word. At each node, maintain a security configuration that describes the operations permitted at the node. The setting of this configuration must be a secure activity and is implemented through the I / O processor selected as part of the boot operation. Using this message, the security configuration can be updated and the security configuration can be generated by a selected processor within the system. The format of this message is unique and can be constructed by data writing, and thus only the identified security processor can generate this message. Since the only consideration is the delivery of the message, the underlying security determination that leads to the generation of the security configuration message is outside the scope of the message bus. Similar to being able to disable the master to enforce separation of debugging from the
[0113] functional network, nodes not selected as the security master are unable to send security messages, but in this case, only a certain type of message is constrained rather than all messages as in the case of network isolation.The message bus may be implemented as a two-dimensional mesh network as shown in FIG. 15. Additionally, two identical networks operate in parallel, and the recombination points are located inside each node. Each link shown has both an input port and an output port replicated for both networks, and both networks can simultaneously transmit and receive on the same side of the node for up to a total of 4 messages valid on either side of the node. In one of the maximum usage cases of the network, all 4 input ports and 4 output ports can be utilized to transfer messages across the entire node. When operating in the maximum usage case, the only constraint on routing is that U-turn routes are not permitted at all, but any other combination of routes to the other 3 outputs is acceptable. Although there are two networks present at every node, the two networks operate completely independently of each other, and routing between the two networks is not possible. A blockage in one network may occur, while the other network is idle. There are several advantages to implementing as a mesh network compared to other topologies, and the greatest advantage is the ability to route around obstacles. Since the message bus is a network that dynamically routes between nodes, there is always a possibility of encountering an obstacle on the direct path between two nodes, from a node that is already in use by other messages to a node that has been powered off to reduce power consumption across the chip. The mesh structure provides options for reaching the destination, and in most cases, a closer to the endpoint.
[0114] There are two logical directions in which the country message can be moved, and one of the directions that gets closer to the end is Even when blocked, it means that there is generally another direction. There may still be messages that cannot be routed, but this is not necessarily a system failure that prevents routing. It is located in an area where the power is turned off, such as an end point where there is no path at all between the required endpoints, such as .
[0115] Another advantage of a mesh network is that the message travel distance is reduced. For multiple nodes, there are several possible connection methods, such as a serial chain, multiple loops, row or column-oriented buses, and a mesh . In a serial chain, the main drawback is that the distance the message may have to travel between two points can be long. Additionally , typically there is only one path available through the chip, so the number of messages that can coexist in the network is generally reduced. The access timing of a serial chain can be variable and difficult to design for an appropriate timing margin . .
[0116] Another problem associated with large-scale serial chains is that power, and any of the nodes, cannot be turned off in any area if it is part of the path that needs to access an unrelated node. An improvement over a single serial bus is several smaller loops, but this leads to a concentration problem of having to move between loops and the possibility of significant delays if a collision occurs at the connection points between loops. Multiple loops also have loops Since the overall power supply requires a significant power step, there are ongoing issues associated with power optimization. The number of simultaneous accesses is increasing, but there are still limitations at points where data must move between independent loops. The number of simultaneous accesses is increasing, but there are still limitations at points where data must move between independent loops. The number of simultaneous accesses is increasing, but there are still limitations at points where data must move between independent loops.
[0117] Multibus-oriented arrays have problems similar to those of multiple loop configurations, i.e., the points where data must move between different bus segments ultimately become bottlenecks in the entire interconnection network. Multibus-oriented arrays have problems similar to those of multiple loop configurations, i.e., the points where data must move between different bus segments ultimately become bottlenecks in the entire interconnection network. Bus arrays actually enable an easier means of transmitting multiple messages at once. However, the ease of obtaining messages on one bus segment is reduced by the complexity of having to mediate between different bus segments. Bus arrays actually enable an easier means of transmitting multiple messages at once. However, the ease of obtaining messages on one bus segment is reduced by the complexity of having to mediate between different bus segments. Bus arrays actually enable an easier means of transmitting multiple messages at once. However, the ease of obtaining messages on one bus segment is reduced by the complexity of having to mediate between different bus segments. As a result, depending on the location of the bus interconnection, certain areas of the device may have to remain powered on just to be able to move data between bus segments. As a result, depending on the location of the bus interconnection, certain areas of the device may have to remain powered on just to be able to move data between bus segments. With IO scattered around the device, data has potential affinity for both sides of the device, so there is no ideal location to place the bus connectors. With IO scattered around the device, data has potential affinity for both sides of the device, so there is no ideal location to place the bus connectors. This results in some layouts being relatively power-efficient, but on the other hand, many otherwise unused nodes have to be powered on to be able to continue interacting with other bus segments, This results in some layouts being relatively power-efficient, but on the other hand, many otherwise unused nodes have to be powered on to be able to continue interacting with other bus segments, leaving other layouts with inferior performance.
[0118] Meshes also support many messages operating in parallel. Since there are no common bottlenecks in the routes, many messages can move through the network simultaneously. Meshes also support many messages operating in parallel. Since there are no common bottlenecks in the routes, many messages can move through the network simultaneously. Routes do not merge through significant obstacles and are not restricted to passing through a single node. As long as possible, each message can often proceed without ever encountering another message. If each processor supports one message at a time, the upper limit of simultaneous messages with a long duration is equal to the number of processors within the system. However, depending on the routes required to deliver parallel messages and return responses to the messages, congestion may reduce the actual upper limit.
[0119] All nodes within the message bus structure function as masters, slaves, or intermediate points of the route, so these basic functions of each node are generally described in detail in this section. The exact interface details may vary across embodiments, and this description provides a functional overview of the message bus node components. An example of the general interface of a message bus node leading into the system is illustrated in FIG. 17. Although the PE link is not required and there are variations in how the node is attached
[0120] to the IO bus, the underlying operations are similar. As illustrated, the message bus node 1701 receives a first message containing payload and routing information, and is configured to select different message nodes among a plurality of message nodes based on the routing information and the operation information of the multi-processor array. As used herein, the operation information is information related to the past or This may be the case. In some instances, the operation information may be current information regarding the performance of the multi-processor array, while in other instances, it may include historical information regarding the performance of the multi-processor array. It should be noted that in some embodiments, the message bus node may receive operation information from the multi-processor array during operation. This may be the case. In some instances, the operation information may be current information regarding the performance of the multi-processor array, while in other instances, it may include historical information regarding the performance of the multi-processor array. It should be noted that in some embodiments, the message bus node may receive operation information from the multi-processor array during operation. This may be the case. In some instances, the operation information may be current information regarding the performance of the multi-processor array, while in other instances, it may include historical information regarding the performance of the multi-processor array. It should be noted that in some embodiments, the message bus node may receive operation information from the multi-processor array during operation. This may be the case. In some instances, the operation information may be current information regarding the performance of the multi-processor array, while in other instances, it may include historical information regarding the performance of the multi-processor array. It should be noted that in some embodiments, the message bus node may receive operation information from the multi-processor array during operation.
[0121] Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent. Message bus node 1701 is further configured to modify the routing information of the first message based on different message nodes to generate a second message and send the second message to different message nodes. Routing information, as used in this specification, is information that specifies the absolute or relative destination of a message. When specifying a relative destination, the number of nodes and the corresponding direction from the starting node are specified to determine the destination of the message. Alternatively, when specifying an absolute destination, an identifier that specifically refers to a particular node is specified as the destination. Each message node may then determine the best possible node for sending the message in order to propagate the message to the specified absolute destination. As will be described in more detail below, the routing information can include an offset that specifies several message nodes and the direction in which the message should be sent.
[0122] As used and described in this specification, a message is a collection of data that includes a payload (i.e., the content of the message) along with routing information. Additionally, a message can include operation information or any suitable portion of the operation information. As used and described in this specification, a message is a collection of data that includes a payload (i.e., the content of the message) along with routing information. Additionally, a message can include operation information or any suitable portion of the operation information. As used and described in this specification, a message is a collection of data that includes a payload (i.e., the content of the message) along with routing information. Additionally, a message can include operation information or any suitable portion of the operation information.
[0123] A message bus node (or simply "message node") may be implemented according to various design styles. A particular embodiment is depicted in FIG. 18. As an illustration, the message bus node 1800 includes a router 1801, a router 1802, a network processor 1803, a network processor 1804, an arbiter 1805, a configuration circuit 1806, and a table 1807. The message bus node 1800 is attached to the PEs and DMRs through the arbiter 1805. In the case of the I / O bus, the arbiter 1805 is a bridge between the I / O bus and the message bus. There are three targets for accesses entering the message bus node 1800 from the local processor, namely, the configuration register (disposed within the configuration circuit 1806) and the network processors 1803 and 1804. Additionally, the network processors 1803 and 1804 may generate accesses to the local node, and there is only one access path returning from the message bus node 1800 to the DMR or PE. Based on the configuration of the node, the type of access, remote request processing, the local requests being generated, or the stored responses, the arbiter 1805 connects one of the network processors 1803 and 1804 to the PE and DMR interfaces. Since only request generation is susceptible to a functional stop from the network side, all writes to the DMR or PE can be generated immediately. To fill in the data for a write request or to request a read in response to a remote access being processed
[0124]
[0125] When combining, arbiter 1805 must wait for one request to complete before switching to the other network processor. If DMR or PE has disabled a request and the access is configured to have a higher priority, the current request can be removed and it is possible to switch to the other network processor. Since PE or DMR has already disabled the access, there is no in - transit data that would be affected by switching the access to the other processor.
[0126] Arbiter 1805 is also configured to direct register bus traffic to the appropriate network processor or configuration register based on the requested address. When remote access is using the configuration register, since the configuration register is the only contention point within message bus node 1800 between the local node and the remote access, arbiter 1805 generates a disable back on the register bus interface.
[0127] Network processors 1804 and 1805 are responsible for the interaction between the attached PE / DMR or IO bus and the rest of the message bus node. Network processors 1803 and 1804 have three responsibilities. The first responsibility is to generate request messages into the network. The second function is to process messages received from the network (including modifying the routing information of the message) and access the local address requested by the message for writing or reading. The last function is to respond to the request message by sending the received reply. It is to process the response message.
[0128] The first function of the network processor (for example, network processor 1803) is to generate a new message in the network. This is achieved in one of two ways i.e., the first way. For a single-word message, the PE accesses the node delta leading to the remote node or nodes in the endpoint group to be accessed, the address of the remote node to be accessed, and in the case of writing, can write the write data. The network processor then generates a message structure and sends the message to the router for delivery. For a longer message, meaning a word length of two or more words, the PE can find the node delta leading to the remote node, the start address of the remote node, the end address of the remote node, and the local address in the DMR where the write data can be found, or in the case of reading, the location where the return data is to be stored. Once these values are configured, the network processor generates a message structure to the router and generates a read request to the DMR to fetch the
[0129] The second function of the network processor is to process the message received from the network and provide a response. In this case, the arriving message structure is disassembled, and the first and last addresses to be accessed are separated and stored. For reading, a read request to the DMR is generated that starts at the first address and continues until the last address is reached. A check is performed to In this configuration, an error value is returned instead of data for read words that cannot be accessed. . In the case of writing, the network processor waits until the first data word arrives, and then generates a write to the DMR for each received word. The write causes an additional check to be performed to verify that the received message is also of the type of the security message if the address is the security configuration address.
[0130] The third function of the network processor is to receive a response to a request and return and store the response for the processor to read. There are two options for this step. The first option relates to single-word responses where the processor can directly read the response from the response register of message bus node 1800. To prevent multiple-word messages from malfunctioning within the network work, when two or more words are returned, the network processor returns and stores these words in the DMR memory. When a read request is generated, a response that stores the address range is also configured within message bus node 1800. The network processor uses the pre-programmed address range to return and store the response and, as a security measure, discards any additional data that may be returned in the message.
[0131] Since there are three functions competing for a single resource, the network processor must also determine which activity to perform at any given point in time. In practice, only the servicing of responses or requests becomes active on the and request generation can be active on the PE / DMR side, so only two of the three can exist simultaneously. The main problem with arbitration is to ensure that no deadlock conditions are formed, and deadlock avoidance is more important than system performance under operations that may deadlock. Since the system can plan how messages flow within the system, the arbitration method is selected from one of three options. In the first method, the first one in is the first one processed. In this mode, a node processes the first request arriving from either the network or the processor side and processes that message until it is completed. This is the simplest way to fully maintain network performance, however, it is prone to deadlocks. The second method is round-robin processing that alternates between two requests for access. Unfortunately, due to the depth of the DMR interface pipeline, the second method can reduce the access speed to 2 / 5. What actually happens is that a write regarding a return, or a remote read, or a write, occupies one cycle, and the next cycle processes the read of the local write message for the written data, and then the interface has to wait for these two accesses to complete. By waiting, the performance is sacrificed to be significantly lower, but network outages interacting with the DMR pipeline are avoided. There is a means to determine that it is not the case that both the messages entering and leaving the MBN are between the same nodes. Multiple There may be a node deadlock, but the system must actively generate scenarios that the hardware does not protect. By checking where the data comes from and comparing where the data is going, it is possible to determine whether two competing messages can generate a deadlock. In such a scenario, a round-robin operation can be selected. Otherwise, a FIFO running at full speed can be set as the default for message delivery across the system, and the messages will complete more quickly than when implementing round-robin. Each of router 1801 and router 1802 is connected to the corresponding network and is configured to receive messages from the network and send the messages generated by network processors 1803 and 1804 to the next destination corresponding to the message. Routers 1801 and 1802 may include a plurality of switches or other appropriate circuitry configured to connect network processors 1803 and 1804 to their corresponding networks. Routers 1801 and 1802 are identical and each perform two main operations on the data passed through the nodes. The first operation is to identify messages intended for the nodes. This involves looking at the two bytes of the delivered node - delta address, starting to extract the next content of the message when a set of zero values is found, and delivering that content to the slave processor.
[0132]
[0133]
[0134] When no agreement is found, the second major operation is to send a message to the next node and proceed towards the destination. The progress towards the destination potentially has two directions. If the paths along the two options that lead closer to the destination are not available, it involves the option of detouring in the third direction. Since going back is not allowed, the direction in which the data arrives is not optional. The basic requirement of the system design is that when following the routing rules, a path should be allowed between two nodes that need to communicate so that the route does not need to make a U-turn. The router is also responsible for inserting new messages into the network. To insert a message into the network, the destination delta offset is known, and the message is accepted unless one of the outputs of the two logical directions towards the destination is already in use, and it is placed in the message bus. Immediately before the first address and data pair, a response data slot is inserted into the message so that the destination node can reply with the result of the requested operation. The response delta is automatically updated based on the path the message takes through the network, and in the case of an error response, any node along the route or the destination node can have the exact destination to which it should send a response in reply to the request message. When discussing the addresses inside the message bus, it is important to distinguish between the address of the message bus node and the value placed in the message for routing to that node. The address of the node effectively includes the IO node, PE, and DMR.
[0135]
[0136] The location of the core array, and the core that includes only the DMR so as to occur at the upper right corner of the array It is the location of the X and Y coordinates of the nodes inside the entire array, including the nodes. Location (0,0 ) is found at the lower left corner of the device, connected to the boot processor, and outside the main core array is arranged. As shown above the entire array in FIG. 19, the core array is bounded by these four corners (1,1 ), (1,17), (17,17), and (17,1), and it should be noted that the format of the figure is (upper number, lower number).
[0137] The address of the location of the message bus node is used when generating the routing delta information for use in the message header. To calculate the necessary routing deltas for the message two signed differences in location are used to identify the number of nodes that need to be traversed in each direction of the mesh to move from the source node to the destination node . For example, to move from (2,2) to (4,7), the delta address (+2,+5) is used, and the return route is (-2,-5). This indicates that the destination is 2 nodes east and 5 nodes north of the current location. This allows for a flexible placement of cells since the routing information is relative and when the cell moves, the endpoints move the same distance and the delta between the two locations remains the same.
[0138] In some cases, the information stored in Table 1807 may be used to determine the routing deltas. For example, the destination information included in the message may be used as an index to Table 1807 to retrieve the data. Such data is for the message and the data may be retrieved. It may be specified which next message bus node is to be sent. Table 1807 may be implemented as a static random-access memory (SRAM), a register file, or other suitable storage circuitry. In various embodiments, the information stored in table 1807 may be loaded during boot sequences and updated during operation of a multi-processor array.
[0139] Assuming 8-bit row and column address values, the message bus may span a 256 x 256 node array. Implementing such a node array and being able to scale the message bus while remaining constant as the technology node shrinks or support multiple die array configurations may be done in later generations and an address format may be selected that does not need to be revised between several generations.
[0140] When a message reaches a destination node, a second address is needed to place a value at the destination node to be accessed. Unlike row and column addresses that have enough room to expand, the PE / DMR destination node local address component actually has no spatial room. As currently defined, there is a 16k word DMR data memory, an 8k word PE instruction memory, a DMR register bus space, PE internal registers, and message bus internal configuration registers. Local addresses do not need all 16 bits of a word and read / write instructions need only 1 bit, The location of bit 15 is used as a control bit. This also writes or reads Since the address is repeated for each burst, it is convenient, and for each burst, reading and writing are selected so that a flexible and efficient means for applying access control is provided.
[0141] In the IO bus interface node, the bus operates with a 32-bit address. Based on the message format, only 15 bits are transferred for each burst, and as a result, 17 bits are unaccounted for in the message. For these remaining bits, a page register is used, with the implicit upper bit 0, and as a result, the IO bus has a potentially 31-bit address more than sufficient to map all necessary memory space and peripheral space. As part of the message accessing the IO bus, the message should start with a write to the page register since the page holds the last value written, and if another master sets the page register to a value different from what the current master expects, it can potentially lead to an unintended access location. To further illustrate the operation of the message bus node, a flowchart depicting an embodiment of a method for operating the message bus node is illustrated in FIG. 22. The method may be applied to the message bus node 1800 or any other suitable message bus node and begins at block 2201.
[0142] To further illustrate the operation of the message bus node, a flowchart depicting an embodiment of a method for operating the message bus node is illustrated in FIG. 22. The method may be applied to the message bus node 1800 or any other suitable message bus node and begins at block 2201.
[0143] The method is a first message including a payload and routing information by a specific message node among a plurality of message nodes included in a multi-processor array including the step of receiving (block 2202). As described above, the first message may be received via one of a plurality of message buses coupled to a particular message node.
[0144] The method also includes the step of a particular message node selecting a different message node among a plurality of message nodes based on routing information and operation information of a multi-processor array (block 2203). As pointed out above, the different message nodes may be based on the relative offset included in the routing information and the congestion or other heuristics included in the operation information.
[0145] Additionally, the method includes the step of a particular message node generating a second message based on different message nodes (block 2204). In various embodiments, a network processor (e.g., network processor 1803) may generate the second message based on which message node was selected. In some cases, the second message may include modified routing information that can be used by a different message node to send a message over a subsequent message node.
[0146] The method further includes the step of a particular message node sending the second message to a different message node (block 2205). In some embodiments, a router (e.g., router 1801) may send the second message based on the relative offset included in the routing information of the first message. The router may send the message Such relative offsets can be used when determining whether to send in a particular direction. The method ends at block 2206.
[0147] HyperOp Data Path Referring to FIG. 20, an embodiment of a HyperOp data path is shown. As illustrated, the HyperOp data path includes two data paths identified as DP0 and DP1. Each of DP0 and DP1 may be identical and may include not only an accumulation circuit, an addition circuit, a shifter circuit, but also additional circuits for moving operands through the data path. It should be noted that a given PE within the multi-processor array may include the HyperOp data path depicted in FIG. 20.
[0148] Using the multi-processor architecture described above, different programming models may be adopted. An example of such a programming model is depicted in FIG. 21. As illustrated, FIG. 21 includes programming models for ASM and HyperOp. Additional details and coding examples regarding different programming models are described below. Each example includes the following: ● C - Reference code describing the functional operation / algorithm. ● ASM - One or more examples of how to implement the operation / algorithm using 64-bit instructions. ASM also includes examples of using vector intrinsic functions (pseudo-ASM instructions) to access the dual DP. The vector intrinsic functions are instructions similar to ASM that are mapped to HyperOp instructions. ● HyperOp - How to implement the operation / algorithm using 128-bit instructions One or more examples of doing.
[0149] Memory operand
[0150] ASM code add16s M1.H,M2.H,M3.H add16s M1.L,M2.L,M3.L
[0151] HyperOp code |A| ld32 M1,%A; / / Load SIMD data from 32-bit M1 Do |B| ld32 M2,%B; / / Load SIMD data from 32-bit M2 Do |DP1| add16s %AH,%BH,%ACC2; / / ACC2 = M1[0 +M2[0] |DP0| add16s %AL,%BL,%ACC0; / / ACC0 = M1[1 +M2[1] |D| dst16 %ACC2_ACC0,M3; / / Store the result of SIMD in 32-bit M3
[0152] Immediate operand
[0153] ASM code sub16 %r2,$10,%r8
[0154] HyperOp code { |A| ld16 %r2,%AL; / / Load 16-bit R2 |C| ld16 $10,%CLH; / / Load 16-bit immediate value 10 |DP1| sub16s %AL,%CLH,%D1; / / D1 = R2 - 10 |D| st16 %D1,%r8; / / Store the result in 16-bit R8 } load immed uses slot C to load the 16-bit segment of the %C register, but note that it can use slot B to load the 16-bit segment of the %B register.
[0155] Conditional execution on scalar
[0156] C code int16 a, b, c, d, e; if (a > b) e = c + d;
[0157] ASM code / / Assumption: / / a is in %R2 / / b is in %R3 / / c is in %R4 / / / / d is in %R5 / / e is in %R6 / / Use %R7 as tmp tcmp16s GT %R2, %R3, %P0 add16s %R4, %R5, %R7 cmov16 (%P0) %R7, %R6
[0158] HyperOp code (conditional memory slot) - Version 1 { |A| ld16s %R2, %AL; / / Load 16-bit R2 |B| ld16s %R3, %BL; / / Load 16-bit R3 |DP0| tcmp16s GT %AL, %BL, %P0; / / Test R2 > R3 and set predicate P0 } { |A| ld16s %R4, %AH; / / Load 16-bit R4 |B| ld16s %R5, %BH; / / Load 16-bit R5 |DP0| add16s %AH,%BH,%D0; / / D0 = R4 + R5 |D| st16 (%P0) %D0,%R6; / / If P0 is true, store the result in 16-bit R6 to the result in 16-bit R6 }
[0159] HyperOp Code (Conditional Store Slot) - Version 2 { |A| ld32 %R2.d,%A; / / Load 32-bit R2:R3 |B| ld32 %R4.d,%B; / / Load 32-bit R4:R5 |DP1| tcmp16s GT %AH,%AL,%P0; / / Test R2 > R3 and set predicate P0 to test R2 > R3 and set predicate P0 |DP0| add16s GT %BH,%BL,%D0; / / D0 = R4 + R5 } { |D| st16 (%P0) %D0,%R6; / / If P0 is true, store the result in 16-bit to store the result in 16-bit R6 }
[0160] Note: ● Conditional execution in the ASM model is only available when using CMOV ● Requires calculating the result in a temp register and then conditionally moving it to the destination to move it conditionally to the destination ● In the HyperOp model, conditional execution is made such that the condition can be applied independently to the slot to make the condition applicable independently to the slot ● Predicate execution uses the predicate flag Pn set in the previous instruction, not the same instruction ● Conditional store is done in a separate instruction slot D ● It may be possible to hide the conditional store in subsequent HyperOps
[0161] Conditional Execution on Vectors
[0162] C code int16 a[2], b[2], c[2], d[2], e[2]; if (a[0] > b[0]) e[0] = c[0] + d[0]; if (a[1] > b[1]) e[1] = c[1] + d[1];
[0163] ASM code / / Assumption: / / a[0], a[1] are in %R2, %R3 / / b[0], b[1] are in %R4, %R5 / / c[0], c[1] are in %R6, %R7 / / d[0], d[1] are in %R8, %R9 / / e[0], e[1] are in %R10, %R11 / / Use %R12, %R13 as temp tcmp16s GT %R2, %R4, %P1 tcmp16s GT %R3, %R5, %P0 add16s %R6, %R8, %R12 add16s %R7, %R9, %R13 cmov16 (%P1) %R12, %R10 cmov16 (%P0) %R13, %R11
[0164] HyperOp code (dual conditional memory) { |A| ld32 %R2.D, %A; / / Load 32-bit R2:R3 ||B| ld32 %R4.D, %B; / / Load 32-bit R4:R5 |DP1| tcmp16s GT %AH, %BH, %P1; / / Test R2 > R4 and set predicate P1 Verify and set predicate P1 |DP0| tcmp16s GT %AL, %BL, %P0; / / Test R3 > R5 and set predicate P0 Verify and set predicate P0 } { |A| ld32 %R6.D,%A; / / Load 32-bit R6:R7 |B| ld32 %R8.D,%B; / / Load 32-bit R8:R9 |DP1| add16s %AH,%BH,%D1; / / D1 = R6 + R8 |DP0| add16s %AL,%BL,%D0; / / D0 = R7 + R9 |D| dst16 (%P1 %P0) %D1_D0,%R10.D; / / If P1 is true, store D1 in 16-bit R10, if P0 is true, / / store D0 in 16-bit R11 }
[0165] Note: ● Conditional execution applies to slot D instructions ● Use SIMD predicate execution mode ● if(%P1 %P0) {...} ● %P1 controls the upper word ● %P0 controls the lower word
[0166] Detect non-zero elements of the array and save the values
[0167] C code int16 a[N],b[N]; int16 i,j; j = 0; for(i = 0,i < N;i++) { if(a[i]<>0) b[j++] = a[i]; }
[0168] ASM code using GPn / / Assumption: / / Use %I1 as i / / Use %I2 as j / / %B1 points to a[] / / %B2 points to b[] / / Use %I0 as a temporary GR gmovi $0,%I2 / / I2 = 0 repeat $0,$N - 1,$1,%I1,L_loop_start,L_lo op_end L_loop_start: mov16 0[%B1 + %I1],%I0 / / I0 = a[i] / / Load %I0 into EX and stop the function for +4 cycles when used in FD gtcmps NE %I0,$0,%GP0 / / Test a[i]<>0 and set predicate G P0 cmov16 (%GP0) 0[%B1 + %I1],0[%B2 + %I2] / / If G P0 is true, move a[i] to 16 - bit b[j] gadd (%GP0) %I2,$1,%I2 / / If GP0 is true, then j++ L_loop_end: Cycles: 2 + N(1 + 4 + 3)=2 + 8N
[0169] ASM code using Pn gmovi $0,%I2 / / I2 = 0 repeat $0,$N - 1,$1,%I1,L_loop_start,L_lo op_end L_loop_start: tcmp16s NE 0[%B1 + %I1],$0,%P0 / / Test a[i]<>0 and set predicate P0 cmov16 (%P0) 0[%B1 + %I1],0[%B2 + %I2] / / If P0 is true, move a[i] to 16 - bit b[j] / / Set %P0 in EX and stop the function for +3 cycles when used in FD gadd (%P0) %I2,$1,%I2 / / If P0 is true, then j++ L_loop_end: Cycle: 2 + N(2 + 3 + 1) = 2 + 6N
[0170] Simple HyperOp code using Pn (conditional G slot execution) / / Assumption: / / %B1 points to a[], and i is in %I1 / / %B2 points to b[], and j is in %I2 gmovi $0, %I2 / / I2 = 0 repeat $0, $N - 1, $1, %I1, L_loop_start, L_lo op_end L_loop_start: { |A| ld16 0[%B1 + %I1], %AL; / / Load 16-bit a[i] / / |DP0| mov16s %AL, %D0; |DP1| tcmp16 NE %AL, $0, %P0; / / Test a[i]<>0 and set predicate P0 / / } { |D| st16 (%P0) %D0, 0[%B2 + %I2]; / / If P0 is true, move a[i] to 16-bit b[j] / / } / / Set %P0 in EX and stop the +3 cycle function when used in FD { |G| gadd (%P0) %I2, $1, %I2 ·· If P0 is true, increment j / / } L_loop_end: Cycle: 2 + N(1 + 1 + 3 + 1) = 2 + 6N
[0171] Pipelined HyperOp code using Pn (conditional storage) / / Assumption: / / %B1 points to a[], and i is in %I1 / / %B2 points to b[], and j is in %I2 gdmovi $0,$1,%I2,%S2 / / I2 = 0, S2 = 1 repeat $0,$N - 1,$4 %I1,L_loop_start,L_lo op_end L_loop_start: {|A| ld16 0[%B1 + %I1],%AL;|DP1| mov16 %A L,%ACC0;|DP0| tcmp16 NE %AL,$0,%P0;} {|A| ld16 1[%B1 + %I1],%AL;|DP1| mov16 %A L,%ACC1;|DP0| tcmp16 NE %AL,$0,%P1;} {|A| ld16 2[%B1 + %I1],%AL;|DP1| mov16 %A L,%ACC2;|DP0| tcmp16 NE %AL,$0,%P2;} {|A| ld16 3[%B1 + %I1],%AL;|DP1| mov16 %A L,%ACC3;|DP0| tcmp16 NE %AL,$0,%P3;} / / Set %P0 in EX and stop the +1 cycle function when used in FD {|A| incr (%P0) $(__i2Mask);|D| st16 (% P0) %ACC0,0[%B2 + %I2];} {|A| incr (%P1) $(__i2Mask);|D| st16 (% P1) %ACC1,0[%B2 + %I2];} {|A| incr (%P2) $(__i2Mask);|D| st16 (% P2) %ACC2,0[%B2 + %I2];} {|A| incr (%P3) $(__i2Mask);|D| st16 (% P3) %ACC3,0[%B2+%I2];} L_loop_end: Cycle: 1 + N / 4(4 + 1 + 4) = 1 + 2.25N
[0172] HyperOp code using two PEs / / Use PE0 to perform tests on the input array a[]: for(i = 0; i < N; i++) { if(a[i]<>0) sendToPE1(a[i]); } / / Use PE1 to save the sparse output array b[]: idx = 0; while(1) { tmp = recvFromPE0(); b[idx++] = tmp; }
[0173] PE0 / / Assumptions: / / %B1 points to a[], and i is in %I1 repeat $0,$N - 1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: tcmp16 NE 0[%B1 + %I1],$0,%P0; cmov16 (%P0) 0[%B1 + %I1],PE0_PE1_QPORT; L_loop_end: PE0 Cycle: 1 + 2N
[0174] PE1 / / Assumptions: / / %B2 points to b[], and j is in %I2 gdmovi $0,$1,%I2,%S2 / / I2 = 0,S2 = 1 L_loop: jmp L_loop; / / Infinite loop on the Q port { |A| incr $(__i2Mask); / / I2 += S2; Useful update for next instruction Useful update |B| ld16 PE0_PE1_QPORT,%BL; |DP0| mov16 %BL,%D0; |D| st16 %D0,0[%B2 + %I2]; / / Use current value of I2 (not updated) for memory Use current value of I2 (not updated) for memory }
[0175] Note: ● By using two PEs, avoid the function stop when setting %GP0 in EX and using it in FD Avoid the function stop
[0176] Detect non - zero elements of the array and save the index
[0177] C code int16 a[N],b[N]; int16 i,j; j = 0; for(i = 0; i < N; i++) { if(a[i]!= 0) { b[j++] = i;} }
[0178] ASM code using GPn / / Assumption: / / %B1 points to a[], i is in %I1 / / %B2 points to b[], j is in %I2 gmov $0,%I2 repeat $0,$N - 1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: mov16 0[%B1 + %I1],%I0 / / Load a[i] into temporary I0 Do / / Load %I0 into EX and stop the function for +4 cycles when used in FD gtcmps NE %I0,$0,%GP0 / / Test a[i]<>0 and set predicate G P0 cmov16 (%GP0) %I1,0[%B2+%I2] / / If GP0 is true, move i to 16-bit b[j] gadd (%GP0) %I2,$1,%I2 / / If GP0 is true, increment j ++ L_loop_end: Cycles: 2 + N(1 + 4 + 3) = 2 + 8N
[0179] ASM code using Pn / / Assumptions: / / %B1 points to a[], i is in %I1 / / %B2 points to b[], j is in %I2 gmov16 $0,%I2 repeat $0,$N-1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: tcmp16s NE 0[%B1+%I1],$0,%P0 / / Test a[i]<>0 and set predicate P0 cmov16 (%P0) %I1,0[%B2+%I2] / / If P0 is true, move i to 16-bit b[j] / / Set %P0 in EX and stop the function for +3 cycles when used in FD gadd (%P0) %I2,$1,%I2 / / If P0 is true, increment j++ L_loop_end: Cycles: 2 + N(2 + 3 + 1) = 2 + 6N
[0180] ASM code using pipelined Pn / / Assumptions: / / %B1 points to a[], and i is in %I1 / / %B2 points to b[], and j is in %I2 gmov16 $0,%I2 repeat $0,$N - 1,$4 %I1,L_loop_start,L_lo op_end L_loop_start: tcmp16s NE 0[%B1 + %I1],$0,%P0 / / a[i + 0]<> Test for 0 and set predicate P0 tcmp16s NE 1[%B1 + %I1],$0,%P1 / / a[i + 1]<> Test for 0 and set predicate P1 tcmp16s NE 2[%B1 + %I1],$0,%P2 / / a[i + 2]<> Test for 0 and set predicate P2 tcmp16s NE 3[%B1 + %I1],$0,%P3 / / a[i + 3]<> Test for 0 and set predicate P3 add16s (%P0)%I1,$0,0[%B2 + %I2] / / If P0 is true Move i + 0 to 16 - bit b[j] gadd (%P0) %I2,$1,%I2 / / If P0 is true, increment j++ Do add16s (%P1) %I1,$1,0[%B2 + %I2] / / If P1 is true Move i + 1 to 16 - bit b[j] gadd (%P1) %I2,$1,%I2 / / If P1 is true, increment j++ Do add16s(%P2) %I1,$2,0 [%B2 + %I2] / / If P2 is true Move 16 - bit b[j] gadd (%P2) %I2,$1,%I2 / / If P2 is true, increment j++ Do add16s (%P3) %I1,$3,0[%B2+%I2] / / If P3 is true, move i+3 to 16-bit b[j] gadd (%P3) %I2,$1,%I2 / / If P3 is true, increment j++ L_loop_end: Cycles: 2+N / 4(4+8)=2+3N
[0181] Simple HyperOp code using GPn (conditional G slot and memory) / / Assumptions: / / %B1 points to a[], i is in %I1 / / %B2 points to b[], j is in %I2 gdmov $0,$1,%I2,%S2 repeat $0,$N-1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: { |A| ld16 0[%B1+%I1],%AL; / / Load a[i] into AL ||DP0| mov16 %AL,%D0; / / Move a[i] to D0 |D| st16 %D0,%I0; / / Store D0=a[i] temporarily in I0 } / / Write %I0 into EX, stalls +4 cycles when used in FD { |B| ld16 %I1,%BH; / / Load i into BH |DP0| mov16s %BH,%D0; / / Move i to D0 |G| gtcmps NE %I0,$0,&GP0; / / Test a[i]<>0 and set predicate P0 } { |A| incr (%GP0) $(__i2Mask); / / Increment j++ if GP0 is true In the case of, increment j++ |D| st16 (%GP0) %D0,0[%B2+%I2]; / / Move i to 16-bit b[j] if GP0 is true In the case of, move i to 16-bit b[j] } L_loop_end: Cycles: 2 + N(1 + 4 + 2) = 2 + 7N
[0182] Simple HyperOp code using Pn / / Assumptions: / / %B1 points to a[], i is in %I1 / / %B2 points to b[], j is in %I2 gdmovi $0,$1,%I2,%S2 repeat $0,$N - 1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: { |A| ld16 0[%B1+%I1],%AL; / / Load a[i] into AL Load |B| ld16 %I1,%BL; / / Load i into BL |DP1| tcmp16s NE %AL,$0,%P0; / / Test a[i]<>0 and set predicate P0 Test |DP0| mov %BL,%D0; / / Move i to D0, prepare for storage Prepare } / / Write %P0 into EX, stalls for +4 cycles when used in FD { |A| incr (%P0) $(__i2Mask); / / Increment j++ if P0 is true Increment j++ |D| st16 (%P0) %D0,0[%B2+%I2]; / / If P0 is true If so, shift i to 16-bit b[j]. } L_loop_end: Cycle: 2 + N(1 + 4 + 1) = 2 + 6N
[0183] HyperOp code pipeline using GPn / / Assumption: / / %B1 points to a[], and i is in %I1 / / %B2 points to b[], and j is in %I2 gdmovi $0,$1,%I2,%S2 repeat $0,$N - 1,$5 %I1,L_loop_start,L_lo op_end L_loop_start: / / Load the next 5 values from a[] into a temporary GR {|A| ld16 0[%B1+%I1],%AL;|DP0| mov16s % AL,%D0;|D| st16 %D0,%T4;} {|A| ld16 1[%B1+%I1],%AL;|DP0| mov16 %A L,%D0;|D| st16 %D0,%T5;} {|A| ld16 2[%B1+%I1],%AL;|DP0| mov16 %A L,%D0;|D| st16 %D0,%T6;} {|A| ld16 3[%B1+%I1],%AL;|DP0| mov16 %A L,%D0;|D| st16 %D0,%T7;} {|A| ld16 4[%B1+%I1],%AL;|DP0| mov16 %A L,%D0;|D| st16 %D0,%I0;} / / if(a[i]<>0) { b[j++]=i;} / / Test a[i+0] to {|A| ld16 %I1,%AH;|G| gtcmpi16 NE %T4,$ 0,%GP0;|DP0| add16s %AH,$0,%D0;} {|A| incr (%GP0) $(__i2Mask);|D| st16 ( %GP0) %D0,0[%B2+%I2];} / / if(a[i+1]<>0) { b[j++]=i+1;} / / a[i+1] to test {|G| gtcmpi16 NE %T5,$0,%GP0;|DP0| add1 6s %AH,$1,%D0;} {|A| incr (%GP0) $(__i2Mask);|D| st16 ( %GP0) %D0,0[%B2+%I2];} / / if(a[i+2]<>0) { b[j++]=i+2;} / / a[i+2] to test {|G| gtcmpi16 %T6,$0,%GP0;|DP0| add16s %AH,$2,%D0;} {|A| incr (%GP0) $(__i2Mask);|D| st16 ( %GP0) %D0,0[%B2+%I2];} / / if(a[i+3]<>0) { b[j++]=i+3;} / / a[i+3] to test {|G| gtcmpi16 NE %T7,$0,%GP0;|DP0| add1 6s %AH,$3,%D0;} {|A| incr (%GP0) $(__i2Mask);|D| st16 ( %GP0) %D0,0[%B2+%I2];} / / if(a[i+4]<>0) { b[j++]=i+4;} / / a[i+4] to test {|G| gtcmpi16 NE %I0,$0,%GP0;|DP0| add1 6s %AH,$4,%D0;} {|A| incr (%GP0) $(__i2Mask);|D| st16 ( %GP0) %D0,0[%B2+%I2];} L_loop_end: Cycle: 2 + N / 5(5 + 5(2)) = 2 + 3N Note: ● All function stops can be hidden by loading into five GRs
[0184] HyperOp code pipelined using Pn / / Assumption: / / %B1 points to a[], i is in %I1 / / %B2 points to b[], j is in %I2 gdmovi $0,$1,%I2,%S2 repeat $0,$N - 1,$4 %I1,L_loop_start,L_lo op_end L_loop_start: / / Test the next five values of a[] and put them into P0 - P3 {|A| ld32 0[%B1+%I1],%A; |C| ld16 %I1,%CLL; / / CLL = I1 |DP1| tcmp16s NE %AH,$0,%P0; |DP0| tcmp16s NE %AL,$0,%P1; } {|B| ld32 2[%B1+%I1],%B; |DP1| tcmp16s NE %BH,$0,%P2; |DP0| tcmp16s NE %BL,$0,%P3; } / / Set %P0 in EX, causing a +3 cycle function stop when used in FD / / if(a[i]<>0) { b[j++]=i;} / / Use P0 {|A| incr (%P0) $(__i2Mask); |DP0| add16s %CLL,$0,%D0; |D| st16 (%P0) %D0,0[%B2+%I2]; } / / if(a[i+1]<>0) { b[j++]=i+1;} / / Use P1 to {|A| incr (%P1) $(__i2Mask); |DP0| add16s %CLL,$1,%D0; |D| st16 (%P1) %D0,0[%B2+%I2]; } / / if(a[i+2]<>0) { b[j++]=i+2;} / / Use P2 to {|A| incr (%P2) $(__i2Mask); |DP0| add16s %CLL,$2,%D0; |D| st16 (%P2) %D0,0[%B2+%I2]; } / / if(a[i+3]<>0) { b[j++]=i+3;} / / Use P3 to {|A| incr (%P3) $(__i2Mask); |DP0| add16s %CLL,$3,%D0; |D| st16 (%P3) %D0,0[%B2+%I2]; } L_loop_end: Cycles: 2 + N / 4(2 + 3 + 4) = 2 + 2.25N Note: ● It is not possible to hide all function stops using four Pn
[0185] HyperOp code using tagged data / / Assumption: / / %B1 points to a[], i is in %I1 / / %B2 points to b[], j is in %I2 gdmovi $0,$1,%I2,%S2 repeat $0,$N - 1,$4 %I1,L_loop_start,L_lo op_end L_loop_start: / / Test the next four values of a[] and put them into P0~P3 { |A| ld32 0[%B1 + %I1],%A; / / Load, AH = a[i + 0], AL = a[i + 1] |C| ld16 $a,%CLL; / / CLL = &a[0] |DP1| tcmp16s NE %AH,$0,%P0; / / a[i + 0] <> 0, set predicate P0 |DP0| tcmp16s NE %AL,$0,%P1; / / a[i + 1] <> 0, set predicate P1 } { |B| ld32 2[%B1 + %I1],%B; / / Load, BH = a[i + 2], BL = a[i + 3] |DP1| tcmp16s NE %BH,$0,%P2; / / a[i + 2] <> 0, set predicate P2 |DP0| tcmp16s NE %BL,$0,%P3; / / a[i + 3] <> 0, set predicate P3 } / / Set %P0 in EX, and when used in FD, it will stop the function for +3 cycles ( INCR instruction) / / if(a[i] <> 0) { b[j++] = i;} { |A| incr (%P0) $(__i2Mask); / / If P0 is true , increment j++ |B| ld16t 0[%B1 + %I1],%B; / / Load tagged data B = {&a[i]:a[i]} |DP0| sub16s %BH,%CLL,%D0; / / D0=&a[i]-& a[0]=i |D| st16 (%P0) %D0,0[%B2+%I2]; / / If P0 is true store i into 16-bit b[j] } / / if(a[i+1]<>0) { b[j++]=i+1;} { |A| incr (%P1) $(__i2Mask); / / If P1 is true increment j++ |B| ld16t 1[&B1+%I1],%B; / / Load tagged data B={&a[i+1]:a[i+1]} |DP0| sub16s %BH,%CLL,%D0; / / D0=&a[i+1] -&a[0]=i+1 |D| st16 (%P1) %D0,0[%B2+%I2]; / / If P1 is true store i+1 into 16-bit b[j] } / / if(a[i+2]<>0) { b[j++]=i+2;} { |A| incr (%P2) $(__i2Mask); / / If P2 is true increment j++ |B| ld16t 2[&B1+%I1],%B; / / Load tagged data B={&a[i+2]:a[i+2]} |DP0| sub16s %BH,%CLL,%D0; / / D0=&a[i+2] -&a[0]=i+2 |D| st16 (%P2) %D0,0[%B2+%I2]; / / If P2 is true store i+2 into 16-bit b[j] } / / if(a[i+3]<>0) { b[j++]=i+3;} { |A| incr (%P3) $(__i2Mask); / / If P3 is true , increment j++ |B| ld16t 3[&B1+%I1],%B; / / Load tagged data B = {&a[i + 3]:a[i + 3]} |DP0| sub16s %BH,%CLL,%D0; / / D0 = &a[i + 3] -&a[0] = i + 3 |D| st16 (%P3) %D0,0[%B2+%I2]; / / If P3 is true , store i + 3 in 16 - bit b[j] } L_loop_end: Cycles: 2 + N / 4(2 + 3 + 4)=2 + 2.25N
[0186] Note: ● Tagged load LD16T loads 16 - bit data (in the lower 16 bits) and its address (in the upper 16 bits) as packed data ● The data index is the data address (or tag), i.e., the start of the array
[0187] Access the array using indirect addressing
[0188] C code int16 a[N],b[N],c[N]; int16 i,j; for(i = 0;i < N;i++) { j = b[i]; a[i]=c[j]; }
[0189] ASM code / / Assume %B1 points to a[] and i is in %I1 / / Assume %B2 points to b[] and i is in %I1 / / Assume that %B4 points to c[] and j is in %I2 repeat $0,$N - 1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: mov16 0[%B2 + %I1],%I2 / / Set %I2 in EX and stop the function for +4 cycles when used in FD mov16 0[%B4 + %I2],0[%B1 + %I1] L_loop_end: Cycles: 1 + N(1 + 4 + 1)=1 + 6N
[0190] Simple HyperOp code / / Assume that %B1 points to a[] and i is in %I1 / / Assume that %B2 points to b[] and i is in %I1 / / Assume that %B4 points to c[] and j is in %I2 repeat $0,$N - 1,$1 %I1,L_loop_start,L_lo op_end L_loop_start: {|A| ld16 0[%B2 + %I1],%AL;|DP0| mov %AL, %D0;|D| st16 %D0,%I2} / / Set %I2 in EX and stop the function for +4 cycles when used in FD {|B| ld16 0[%B4 + %I2],%BL;|DP0| mov %BL, %D0;|D| st16 %D0,0[%B1 + %I1]}; L_loop_end: Cycles: 1 + N(1 + 4 + 1)=1 + 6N
[0191] Pipelined HyperOp code / / Assume that %B1 points to a[] and i is in %I1 / / Assume that %B2 points to b[], and i is within %I1 / / Assume that %B4 points to c[], and j is within %I2 to %I7 / / j0 = b[0]; j1 = b[1]; {|A| ld32 0[%B2], %A; |DP0| mov32 %A, %D0; |D| st32 %D0, %I2I3;} / / j2 = b[2]; j3 = b[3]; {|A| ld32 2[%B2], %A; |DP0| mov32 %A, %D0; |D| st32 %D0, %I4I5;} / / j4 = b[4]; j5 = b[5]; {|A| ld32 4[%B2], %A; |DP0| mov32 %A, %D0; |D| st32 %D0, %I6I7;} / / Set %I2 and %I3 in EX, and when used in FD, the function stops for +1 cycle Stop repeat $0, $N - 1, $6 %I1, L_loop_start, L_lo op_end L_loop_start: / / a[i + 0] = c[j0]; a[i + 1] = c[j1]; j0 = b[i + 6]; j 1 = b[i + 7]; {|A| ld16 0[%B4 + %I2], %AL; |B| ld16 0[%B4 + %I3], %BL; |DP1| mov16 %AL, %D1; |DP0| mov16 %BL, %D0 ; |D| dst16 %D1_D0, 0[%B1 + %I1];} {|A| ld32 6[%B2 + %I1], %A; |DP0| mov32 %A, %D0; |D| st32 %D0, %I2I3;} / / a[i + 2] = c[j2]; a[i + 3] = c[j3]; j2 = b[i + 8]; j 3 = b[i + 9]; {|A| ld16 0[%B4+%I4],%AL;|B| ld16 0[%B4 +%I5],%BL; |DP1| mov16 %AL,%D1;|DP0| mov16 %BL,%D0 ;|D| dst16 %D1_D0,2[%B1+%I1];} {|A| ld32 8[%B2+%I1],%A;|DP0| mov32 %A, %D0;|D| st32 %D0,%I4I5;} / / a[i+4]=c[j4];a[i+5]=c[j5];j4=b[i+10]; j5=b[i+11]; {|A| ld16 0[%B4+%I6],%AL;|B| ld16 0[%B4 +%I7],%BL; |DP1| mov16 %AL,%D1;|DP0| mov16 %BL,%D0 ;|D| dst16 %D1_D0,4[%B1+%I1];} {|A| ld32 10[%B2+%I1],%A;|DP0| mov32 %A ,%D0;|D| st32 %ACC0,%I6I7;} / / Ignore the final values loaded into I1 to I7 L_loop_end: Cycles: 3 + 1 + 1 + N / 6(6) = 5 + N
[0192] Note: ● Index j is loaded in pairs from b[i] in one cycle ● Two c[j] are loaded as a pair in one cycle and stored in a[i] ● By using six index registers, pipeline bubbles such as index setting in EX and index usage in FD are avoided
[0193] Conditional accumulation using double DP
[0194] The following is an example of where Conditional HyperOp with two predicates can be used.
[0195] C code int16 a[N], b[N], c[N]; int16 i; int32 sum = 0; for (int i = 0; i < N; i++) { if (a[i] > b[i]) sum += a[i] * c[i]; }
[0196] ASM code
[0197] This example uses vector intrinsic functions (pseudo ASM instructions) to access double DP. repeat $0, $N - 1, $2, IDX_i, L_loop_start, L_ loop_end movx16s $0, %ACC2 movx16s $0, %ACC0 L_loop_start: vtcmp16s GT 0[BP_a + IDX_i], 0[BP_b + IDX_i] , %P1P0; cmov16 (%P1) 0[BP_a + IDX_i], $0, %R0 cmov16 (%P0) 1[BP_1 + IDX_i], $0, %R1 vmulaa16s %R0.D, 0[BP_c + IDX_i], %ACC2_ACC 0 L_loop_end: accadd %ACC0, $0, %ACC2 Cycles: 3 + N / 2(4) + 1 = 4 + 2N
[0198] HyperOp code (conditional DP slot execution - both slots) #define BP_a %B1 #define BP_b %B2 #define BP_c %B3 #define IDX_i %I1 repeat $0,$N - 1,$2,IDX_i,L_loop_start,L_ loop_end {|DP1| movx16s $0,%ACC2;|DP0| movx16s $ 0,%ACC0;} L_loop_start: {|A| ld32 0[BP_a + IDX_i],%A;|B| ld32 0[B P_b + IDX_i],%B; |DP1| tcmp16s GT %AH,%BH,%P1;|DP0| tcmp 16s GT %AL,%BL,%P0;} {|C| ld32 0[BP_c + IDX_i],%B; |DP1| mulaa16s (%P1) %AH,%BH,%ACC2;|DP0 | mulaa16s (%P0) %AL,%BL,%ACC0;} L_loop_end: accadd %ACC0,$0,%ACC2 Cycles: 1 + N / 2(2) + 1 = 2 + N
[0199] Note: ● Use DP1 and DP0 to process iterations i and i + 1 in parallel ● Split the sum into %ACC0 and %ACC2, and then combine them at the end ● Use predicate flags %P1 and %P0 to independently control the accumulation in %ACC2 and %ACC0 respectively
[0200] Conditional accumulation using double DP with double MUL respectively
[0201] The following is an example of where conditional HyperOp with four predicates can be used. C code int16 a[N], b[N], c[N]; int16 i; int32 sum = 0; for (int i = 0; i < N; i++) { if (a[i] > b[i]) sum += a[i] * c[i]; }
[0202] HyperOp code (using both DPs with quadruple conditions) #define BP_a %B1 #define BP_b %B2 #define BP_c %B3 repeat $0, $N - 1, $4, IDX_i, L_loop_start, L_ loop_end {|DP1| movx16s $0, %ACC2; |DP0| movx16s $ 0, %ACC0;} L_loop_start: {|A| ld64 0[BP_a + IDX_i], %AB; |C| ld64 0 BP_b + IDX_i], %C; |DP1| dtcmp16s GT %A, %CH, %P3P2; |DP0| dtcmp16s GT %B, %CL, %P1P0;} {|C| ld64 0[BP_c + IDX_i], %C; |DP1| dmulaa16s (%P3P2) %A, %CH, %ACC2; |DP0| dmulaa16s (%P1P0) %B, %CL, %ACC0;} L_loop_end: accadd %ACC0, $0, %ACC2 Cycles: 2 + N / 4(2) + 1 = 3 + 0.5N
[0203] Note: ●Process iterations i to i+3 in parallel: ● i and i+1 in DP1 ●i+2 and i+3 at DP0 ● DP0 performs dual operation, DP1 performs dual operation Split the total into %ACC0 and %ACC2, then combine at the end ● Use the predicate flags %P0~P3 to perform product accumulation independently in %ACC0 and %ACC2. Controlled b[] and c[] have different DM addresses than a[] because 64-bit access works. Must be R
[0204] Conditional Memory Using Dual DP
[0205] The following C code uses a conditional HyperOp to perform a conditional store: This is one example of what can be done.
[0206] C Code int16 a[N], b[N], c[N], d[N]; int16 i; for(int i=0;i <N;i++) { if(a[i]>b[i]) d[i]=a[i]*c[i]; }
[0207] ASM Code
[0208] This example uses vector eigenfunctions (pseudo AM instructions) to access the dual DP. #define BP_a %B1 #define BP_b %B2 #define BP_c %B3 #define BP_d %B4 #define IDX_i %I1 repeat $0,$N - 1,$2,IDX_i,L_loop_start,L_ loop_end L_loop_start: vtcmp16s GT [BP_a + IDX_i],[BP_b + IDX_i],% P1P0 vmul16s (%P1P0) [BP_a + IDX_i],[BP_c + IDX_ i],[BP_d + IDX_i] L_loop_end:
[0209] HyperOp code (double conditional memory) #define BP_a %B1 #define BP_b %B2 #define BP_c %B3 #define BP_d %B4 #define IDX_i %I1 repeat $0,$N - 1,$2,IDX_i,L_loop_start,L_ loop_end L_loop_start: {|A| ld32 0[BP_a + IDX_i],%A;|B| ld32 0[B P_b + IDX_i],%B; |DP1| tcmp16s GT %AH,%BH,%P1;|DP0| tcmp 16s GT %AL,%BL,%P0;} {|C| ld32 0[BP_c + IDX_i],%C; |DP1| mul16s %AH,%CLH,%D1;|DP0| mul16s %AL,%CLL,%D0; |D| dst16 (%P1P0) %D1_D0,0[BP_d + IDX_i]; } L_loop_end:
[0210] Note: ● Use DP1 and DP0 to process iterations i and i+1 in parallel ● Use predicate flags %P1 and %P0 to independently control 16-bit:16-bit memory (SIMD mode)
[0211] Example of conditional if-else-if using conditional jump
[0212] C code absq = abs(q); if (absq < qmin) { qmin2 = qmin; qmin = absq; imin = i; } else if (absq < qmin2) { qmin2 = absq; }
[0213] ASM code / / Assume imin and qmin are stored as packed data imin_qmin at an even address abs16s q, absq / / absq = abs(q) tcmp16 LT absq, qmin, %P1 / / P1 = (absq < qmin ) jmp (!%P1) L_else PNT / / If!P1 is true, skip qmin update tcmp16 LT absq, qmin2, %P0 / / P0 = (absq < qmi n2) - Delay slot L_if: / / Update qmin and qmin2: mov16 qmin, qmin2 / / qmin2 = qmin jmp L_end dmov16 i, absq, imin_qmin / / qmin = absq, imi n = i - Delay slot L_else: jmp (!%P0) L_end PNT DLY nop / / Delay the slot mov16 absq,qmin2 / / Update only qmin2 L_end:
[0214] ASM code using DLY optimization abs16s q,absq tcmp16 LT absq,qmin,%P1 jmp (!%P1) L_else PNT tcmp16 LT absq,qmin2,%P0 / / Executed in delay slot L_if: mov16 qmin,qmin2 jmp L_end dmov16 i,absq,imin_qmin / / Executed in delay slot L_else: jmp (!%P0) L_end PNT DLY mov16 absq,qmin2 / / Executed after JMP and not in delay slot L_end:
[0215] Example of conditional if-else-if using conditional move
[0216] C code absq=abs(q); if(absq<qmin) { qmin2=qmin; qmin=absq; imin=i; } else if(absq<qmin2) { qmin2=absq; }
[0217] ASM code / / imin and qmin are stored as packed data imin_qmin (even address) Assume so abs16s q, absq / / absq = abs(q) tcmp16s LT absq, qmin, %P1 / / P1 = (absq < qmi n) tcmp16s LT absq, qmin2, %P0 / / P0 = (absq < qm in2) cmov16 (%P1) qmin, qmin2 / / If P1 is true, qmi n2 = qmin cmov16 (%P1) absq, qmin / / If P1 is true, qmin = absq cmov16 (%P1) i, imin / / If P1 is true, imin = i cmov16 (!%P1 & %P0) absq, qmin2 / / If P0 is true in addition Then, qmin2 = absq Cycles: 7
[0218] HyperOp code {|A| ld16 q, AL; |B| ld16 i, %BL; |DP1| mov16 %BL, %ACC3; / / ACC3 = i |DP1| abs16s %AL, %ACC1; / / ACC1L = absq |D| dst16 %ACC3_ACC1, %ACC3;} / / ACC3H = i, ACC3L = absq {|A| ld32 imin_qmin, %A; / / AH = imin, AL = qm in |B| ld16 qmin2, %BL; / / BL = qmin2 |DP1| tcmp16 LT %ACC3L, %AL, %P1; / / P1 = (a bsq < qmin) |DP0| tcmp16 LT %ACC1L,%BL,%P0;} / / P0=( absq <qmin2) {|DP1| if(%P1) cmov32 %ACC3,%A,%ACC2; / If / P1 is true, {ACC2H=i,ACC2L=absq} / / Else, {ACC2H=imin,ACC2L=qmin} |DP0| if(%P1) cmov16 %AL,%BL,%ACC0; / / A CC0=(P1)?qmin2:qmin |D| st32 %ACC2,imin_qmin;} / / imin:qmin= Update to ACC2H:ACC2L {|DP0| if(!%P1&%P0) cmov16 %ACC3L,%ACC0 L,%ACC0; / / otherwise,ACC0L=(P0) ? absq:qm in |D| st16 %ACC0,qmin2;} / / Update qmin2=ACC0L do Cycle: 4
[0219] Note: Use P1 and %P0 to determine the Boolean results of the IF and ELSE IF tests. Hold the results imin and qmin are stored in memory as packed 16:16 Assume that Use %P1 and %P0 with CSEL to set state variables in pairs when possible. Conditionally update
[0220] Combining tests using predicate flags
[0221] C Code int16 a,b,c,d,e; void test() { a=(b<c)&&(d<e); }
[0222] ASM code tcmp16s LT b,c %P0 / / P0=(b<c) tcmp16s LT d,e,%P1 / / P1=(d<e) cmov16 (%P0&%P1) $1,$0,a / / a=(P0&P1)1:0
[0223] Note: ● The compiler replaces the && operator with the & operator: ● a=(b<c)&(d<e)
[0224] Combining tests using the register file
[0225] C code int16 a,b,c,d,e; void test() { a=(b<c)&&(d<e); }
[0226] ASM code tcmp16s LT b,c,%R0 / / R0=(b<c) tcmp16s LT d,e,%R1 / / R1=(d<e) and16 %R0,%R1,a / / a=R0&R1
[0227] Note: ● The compiler replaces the && operator with the & operator: ● a=(b<c)&(d<e)
[0228] Conditional jump to subroutine
[0229] C code int16 a,b,c,d,e,f; if((a<b)&(c<e)|(d>f)) foo();
[0230] ASM code tcmp16s LT a,b,%R1 / / R1=(a<b) tcmp16s LT c,e,%R2 / / R2=(c<e) tand16 NZ %R1,%R2,%P0 / / P0=(a<b)&(c<e) tcmp16s GT d,f,%P1 / / P1=(d>f) jsr (%P0|%P1) foo / / If P0|P1 is true, execute foo() execute
[0231] Note: ●Use TAND16 instead of AND16 ●Note that Pn cannot be dstD in ALU operations other than TEST be careful
[0232] Assignment of logical / test operation results
[0233] C code int16 a,b,c,d,e,f,result; result=((a<b)&(c<e)|(d>f));
[0234] ASM code tcmp16s LT a,b,%R1 / / R1=(a<b) tcmp16s LT c,e,%R2 / / R2=(c<e) and16 %R1,%R2,%R3 / / P3=(a<b)&(c<e) tcmp16s GT d,f,%R4 / / R4=(d>f) or16 %R3,%R4,result / / result=(R3|R4)
[0235] Any of the various forms described in this specification, for example, as a computer-implemented method and may be implemented in any form, such as a computer-readable storage medium or a computer system. The system may be implemented by one or more custom-designed hardware devices such as Application Specific Integrated Circuits (ASICs), by one or more programmable hardware elements such as Field Programmable Gate Arrays (FPGAs), by one or more processors that execute program-stored instructions, or by any combination of the foregoing.
[0236] In some embodiments, the non-transitory computer-readable storage medium may be configured to store program instructions and / or data, wherein the program instructions, when executed by a computer system, cause the computer system to perform a method, for example, any one of the embodiments of the methods described herein, or any combination of the embodiments of the methods described herein, or any subset consisting of any one of the embodiments of the methods described herein, or any combination of such subsets.
[0237] In some embodiments, the computer system may be configured to include a processor (or a set of processors) and a storage medium, wherein the storage medium stores program instructions, wherein the processor is configured to read and execute the program instructions from the storage medium, wherein the program instructions cause the processor to perform any of the various methods described herein (or any combination of the embodiments of the methods described herein, or any subset consisting of any one of the embodiments of the methods described herein, or any combination of such subsets). Any subset consisting of any of the embodiments of the methods described herein, or any combination of such subsets) is executable to implement. The computer system may be realized in any of various forms. For example, the computer system may be a personal computer (taking any form of various realizations of a personal computer ), a workstation, a computer on a card, an application-specific computer in a box , a server computer, a client computer, a handheld device , a mobile device, a wearable computer, a detection device, a television, a video capture device, a computer embedded in a living body , etc. The computer system may include one or more display devices. Any of the various calculation results disclosed herein may be displayed via a display device, or presented as an output via other methods through a user interface device .
[0238] The apparatus includes a plurality of processors and a plurality of data memory routers connected to the plurality of processors in a scattered array, and a specific data memory router relays a message received by at least one other data memory router among the plurality of data memory routers, and a specific processor among the plurality of processors is configured to set at least one predicate flag among the plurality of predicate flags and execute instructions conditionally using the plurality of predicate flags .
[0239] In the aforementioned apparatus, the plurality of predicate flags include at least a first set of predicate flags related to a data path included in a specific processor, and an address generator included in the specific processor It includes a second set of predicate flags associated with the unit.
[0240] In the aforementioned apparatus, in order to set at least one of the plurality of predicate flags, a specific processor is further configured to compare a first value and a second value in response to the execution of a test instruction, generate a result, and set at least one predicate flag based on the result. be configured.
[0241] In the aforementioned apparatus, in order to compare the first value and the second value, a specific processor is further configured to perform a logical operation using the first value and the second value to generate a result. be configured.
[0242] In the aforementioned apparatus, in order to set at least one of the plurality of predicate flags, a specific processor is further configured to set at least one predicate flag based at least in part on information indicating the timing operation of a data path included in the specific processor. be configured.
[0243] In the aforementioned apparatus, in order to set at least one of the plurality of predicate flags, a specific processor is further configured to set at least one predicate flag based at least in part on information indicating the timing operation of an address generator unit included in the specific processor. be configured.
[0244] In the aforementioned apparatus, in order to conditionally execute an instruction, a specific processor is further configured to conditionally execute one or more data path slots included in a data path included in the specific processor using a plurality of predicate flags.
[0245] A method includes a step of setting at least one predicate flag among a plurality of predicate flags by a specific processor among a plurality of processors, where the plurality of processors are connected to a plurality of data memory routers in a scattered array, and a step of conditionally executing an instruction using the plurality of predicate flags by the specific processor among the plurality of processors.
[0246] In the foregoing method, the plurality of predicate flags include at least a first set of predicate flags related to a data path included in a specific processor and a second set of predicate flags related to an address generator unit included in the specific processor.
[0247] In the foregoing method, the step of setting at least one predicate flag among the plurality of predicate flags includes a step of comparing a first value and a second value in response to executing a test instruction by a specific processor to generate a result, and a step of setting at least one predicate flag based on the result.
[0248] In the foregoing method, the step of comparing the first value and the second value includes a step of performing a logical operation using the first value and the second value to generate a result.
[0249] In the foregoing method, the step of setting at least one predicate flag among the plurality of predicate flags includes a step of setting at least one predicate flag based at least in part on information indicating a timing operation of a data path included in a specific processor by the specific processor.
[0250] In the foregoing method, the step of setting at least one of a plurality of predicate flags is, at least in part, based on information indicating the timing operation of an address generator unit included in a specific processor by the specific processor and includes the step of setting at least one predicate flag.
[0251] In the foregoing method according to claim 22, the step of conditionally executing an instruction by a specific processor includes the step of conditionally executing one or more data path slots included in a data path included in a specific processor using a plurality of predicate flags.
[0252] The apparatus includes a plurality of processors and a plurality of data memory routers connected to the plurality of processors in a scattered array, and a specific data memory router is configured to relay a message received by at least one other data memory router among the plurality of data memory routers, and a specific processor among the plurality of processors selectively activates a subset of a plurality of arithmetic logic circuits included in a specific data path included in the specific processor based on a received instruction and is configured to execute the received instruction using the subset of the plurality of arithmetic logic circuits.
[0253] In the foregoing apparatus, to selectively activate a subset of a plurality of arithmetic logic circuits, a specific processor is further configured to decode an instruction, generate a decoded instruction, and use the decoded instruction to selectively activate a subset of a plurality of arithmetic logic circuits.
[0254] In the foregoing apparatus, a specific processor among the plurality of processors, based on an instruction, a plurality Data is routed among individual arithmetic logic circuits included in a subset of arithmetic logic circuits and is further configured to do so.
[0255] In the foregoing apparatus, to route data among individual arithmetic logic circuits included in a subset of a plurality of arithmetic logic circuits a particular processor is further configured to selectively change the state of at least one of a plurality of multiplexing circuits included in a particular data path for routing data among individual arithmetic logic circuits included in a subset of a plurality of arithmetic logic circuits. and is further configured to do so.
[0256] In the foregoing apparatus, a particular arithmetic logic circuit among a plurality of logic circuits includes at least an adder circuit and is configured to do so.
[0257] In the foregoing apparatus, a particular arithmetic logic circuit among a plurality of logic circuits includes a look-up table configured to store an offset used when executing an instruction and is configured to do so.
[0258] In the foregoing apparatus, the instruction specifies a logarithmic probability operation.
[0259] The method includes a step of selectively activating a subset of a plurality of arithmetic logic circuits included in a particular data path among a plurality of data paths included in a particular processor among a plurality of processors, wherein the plurality of processors are connected to a plurality of data memory routers in a scattered array and a step of executing an instruction using a subset of a plurality of arithmetic logic circuits by a particular processor among the plurality of processors. and is configured to do so.
[0260] In the foregoing method, the step of selectively activating a subset of a plurality of arithmetic logic circuits includes decoding an instruction to generate a decoded instruction and using the decoded instruction to generate a plurality of selectively activating a subset of the arithmetic logic circuits;
[0261] The method further includes routing data between individual arithmetic logic circuits included in a subset of the plurality of arithmetic logic circuits based on an instruction. routing data between individual arithmetic logic circuits included in a subset of the plurality of arithmetic logic circuits based on an instruction.
[0262] In the method, the step of routing data between individual arithmetic logic circuits included in a subset of the plurality of arithmetic logic circuits includes selectively changing, by a specific processor, a state of at least one multiplexing circuit among a plurality of multiplexing circuits included in a specific data path. includes selectively changing, by a specific processor, a state of at least one multiplexing circuit among a plurality of multiplexing circuits included in a specific data path. routing data between individual arithmetic logic circuits included in a subset of the plurality of arithmetic logic circuits based on an instruction.
[0263] In the method, a specific arithmetic logic circuit among the plurality of logic circuits includes at least an adder circuit. In the method, a specific arithmetic logic circuit among the plurality of logic circuits includes at least an adder circuit.
[0264] In the method, a specific arithmetic logic circuit among the plurality of logic circuits includes a look-up table, and further includes storing, in the look-up table, an offset used when executing an instruction. In the method, a specific arithmetic logic circuit among the plurality of logic circuits includes a look-up table, and further includes storing, in the look-up table, an offset used when executing an instruction. routing data between individual arithmetic logic circuits included in a subset of the plurality of arithmetic logic circuits based on an instruction.
[0265] In the method, the instruction specifies a logarithmic probability operation.
[0266] Although the above embodiments have been described in connection with the preferred embodiments, it is not intended to be limited to the specific forms shown herein. On the contrary, such alternative forms, modifications, and equivalents are intended to be reasonably included within the spirit and scope of the embodiments of the invention as defined by the appended claims. modifications, and equivalents are intended to be reasonably included within the spirit and scope of the embodiments of the invention as defined by the appended claims. routing data between individual arithmetic logic circuits included in a subset of the plurality of arithmetic logic circuits based on an instruction.
Claims
Claim 1 An apparatus comprising: a plurality of processors including a specific processor including an address generator; a plurality of data memory routers connected to the plurality of processors in a scattered array; and wherein a specific data memory router is configured to relay messages received by at least one other data memory router among the plurality of data memory routers; wherein the specific processor among the plurality of processors: sets a specific predicate flag among a first set of predicate flags associated with a data path included in the specific processor, the first set of predicate flags being included in a plurality of predicate flags, and the setting of the specific predicate flag is performed based at least in part on information indicating a timing operation associated with the data path; uses the plurality of predicate flags to conditionally execute an instruction; sets a second predicate flag different from the specific predicate flag from among a second set of predicate flags associated with the address generator, based on information indicating a timing operation associated with the address generator, the second set of predicate flags being associated with the address generator and being included in the plurality of predicate flags, and the setting of the second predicate flag based on the information indicating the timing operation associated with the address generator is performed by executing a test operation on the address generator, verifying a resulting condition, selecting the second predicate flag corresponding to the verified condition, and setting the selected second predicate flag; An apparatus configured as such. Claim 2 The specific processor is further configured to: in response to executing a test instruction, compare a first value and a second value to generate a result; and set the specific predicate flag based on the result. The apparatus according to claim 1, configured as such. Claim 3 To compare the first value and the second value, the specific processor is further configured to: perform a logical operation using the first value and the second value to generate the result. The apparatus according to claim 2, configured as such.
4. The setting of the specific predicate flag based on the information indicating the timing operation related to the data path is performed by executing a test operation on the data path, checking the resulting conditions, selecting the specific predicate flag corresponding to the checked conditions, and setting the selected specific predicate flag, and is configured as such, the apparatus according to claim 1.
5. The data path includes a plurality of slots, and the instruction is included in a specific slot among the plurality of slots, the apparatus according to claim 1.
6. Conditionally executing the instruction includes selecting the specific slot based on the specific predicate flag, the apparatus according to claim 5.
7. The data path includes a plurality of arithmetic logic circuits, the apparatus according to claim 1.
8. A step of setting a specific predicate flag among a plurality of predicate flags including a first set of predicate flags related to a data path included in a specific processor by the specific processor, wherein the setting of the specific predicate flag is performed based at least in part on information indicating a timing operation related to the data path, the plurality of processors are coupled to a plurality of data memory routers in a scattered array, and the specific processor includes an address generator, the step A step of conditionally executing an instruction using the plurality of predicate flags by the specific processor; A step of setting a second predicate flag included in a second set of predicate flags different from the specific predicate flag and included in the plurality of predicate flags based on information indicating a timing operation related to the address generator by the specific processor, wherein the setting of the second predicate flag based on the information indicating the timing operation related to the address generator is performed by executing a test operation on the address generator, checking the resulting conditions, selecting the second predicate flag corresponding to the checked conditions, and setting the selected second predicate flag, and the second set of predicate flags is associated with the address generator, the step, A method comprising.
9. A step of generating a result by comparing a first value and a second value by the specific processor in response to the execution of a test instruction A step of setting the specific predicate flag based on the result by the specific processor; The method according to claim 8, further comprising.
10. The method according to claim 9, wherein comparing the first value and the second value comprises performing a logical operation using the first value and the second value to generate the result.
11. The setting of the specific predicate flag based on the information indicating the timing operation related to the data path is performed by executing a test operation on the data path, confirming the resulting conditions, selecting the specific predicate flag corresponding to the confirmed conditions, and setting the selected specific predicate flag. The method according to claim 8.
12. The method according to claim 8, wherein the data path includes a plurality of slots, and the instruction is included in a specific slot among the plurality of slots.
13. The method according to claim 12, wherein the step of conditionally executing the instruction includes selecting the specific slot based on the specific predicate flag.
14. The method according to claim 8, wherein the data path includes a plurality of arithmetic logic circuits.
15. A plurality of processors including a specific processor including a specific data path including a plurality of arithmetic logic circuits, the specific arithmetic logic circuit among the plurality of arithmetic logic circuits including a lookup table configured to store an offset, and a plurality of processors; A plurality of data memory routers connected to the plurality of processors in a scattered array, wherein a specific data memory router is configured to relay a received message to at least one other data memory router among the plurality of data memory routers. An apparatus comprising the plurality of data memory routers; The specific processor among the plurality of processors is Selectively activating a subset of the plurality of arithmetic logic circuits based on an instruction received via a data memory router; Executing the received instruction using the subset of the plurality of arithmetic logic circuits to generate a result; Adding the offset to the result to generate a final result; An apparatus configured as described above.
16. To selectively activate a subset of the plurality of arithmetic logic circuits, the specific processor further Decode the received instruction to generate a decoded instruction, and selectively activate a subset of the plurality of arithmetic logic circuits using the decoded instruction The apparatus according to claim 15, configured as such.
17. The specific processor is further configured to route data among predetermined arithmetic logic circuits of the subset of the plurality of arithmetic logic circuits based on the received instruction, the apparatus according to claim 15.
18. The specific data path includes a plurality of multiplexing circuits including a specific multiplexing circuit coupled between a first arithmetic logic circuit of the plurality of arithmetic logic circuits and a second arithmetic logic circuit of the plurality of arithmetic logic circuits, and the specific processor is further configured to selectively change the state of the specific multiplexing circuit to route the data, the apparatus according to claim 17.
19. A specific arithmetic logic circuit of the plurality of arithmetic logic circuits includes at least an addition circuit, The apparatus according to claim 15.
20. The received instruction specifies a logarithmic probability operation, the apparatus according to claim 15.
Citation Information
Patent Citations
Memory network processor with programmable optimization
JP2016526220A
Conditional branch execution
US20020199090A1
Memory controller operating method and memory controller
US20150178154A1
Multiprocessor fabric having configurable communication that is selectively disabled for secure processing
US9424441B2
Memory-network processor with programmable optimizations
US9430369B2