Execution of instructions requiring access to array register

By decomposing instructions into execution parts and issuing them based on hazard-free predictions, the control circuit optimizes array register utilization, enhancing throughput and reducing power consumption fluctuations in data processing systems.

JP2025100432APending Publication Date: 2025-07-03ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024218417
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-13
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing data processing systems face inefficiencies due to fluctuations in the utilization rate of array registers, leading to sub-optimal throughput and power consumption variations, particularly when instructions require access to multiple array regions.

Method used

A control circuit decomposes instructions requiring access to multiple array regions into multiple execution parts, delaying each part until it can be processed hazard-free, and issues them in different cycles based on prediction, optimizing the processing circuit's workload.

Benefits of technology

This approach improves overall instruction throughput and reduces fluctuations in power consumption by ensuring efficient utilization of array registers, allowing parallel execution of instructions that previously stalled due to varying access requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100432000001_ABST
    Figure 2025100432000001_ABST
Patent Text Reader

Abstract

To provide an apparatus, a system, a chip containing product, a method, and a medium that improve the throughput of instructions to access an array register.SOLUTION: An apparatus 10 includes control circuitry 16, a register 14, and processing circuitry 12. In addition to the register 14, the apparatus 10 includes an array register 18 that is divided into a plurality of array regions, including array regions 18(A)-18(D). In response to receiving an instruction requiring access to two or more array regions, the control circuitry decomposes the instruction into two or more execution parts, each corresponding to one of the two or more array regions, for each execution part, delays issuing the execution part until it is predicted that the execution part can be processed hazard free, and issues the two or more execution parts in different cycles based on when it is predicted that each execution part can be processed hazard free.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing. Further, the present invention relates to an apparatus, a system, a chip-containing product, a method, and a non-transitory computer-readable storage medium.

Background Art

[0002] Some apparatuses include a plurality of registers including an array register.

Summary of the Invention

[0003] According to a first aspect of the present technology, a plurality of registers configured to store data, the plurality of registers including at least one array register including a plurality of array regions; a processing circuit configured to receive an issued instruction and execute a processing operation identified in the issued instruction; a control circuit including a buffer circuit configured to store one or more instructions prior to execution by the processing circuit, the control circuit, in response to receiving an instruction that requires access to two or more of the plurality of array regions, decomposes the instruction into two or more execution parts, each of the two or more execution parts corresponding to one of the two or more array regions, for each execution part of the two or more execution parts, delays issuing the execution part as one of the issued instructions until it is predicted that the execution part can be processed hazard-free by the processing circuit, and the control circuit can issue the two or more execution parts of the instruction in different cycles selected based on when each of the two or more execution parts is predicted to be processed hazard-free. An apparatus including the control circuit is provided.

[0004] According to a second aspect of the present technology, an apparatus according to the first aspect mounted on at least one packaged chip; At least one system component, and a substrate, and the system provided comprises: The at least one packaged chip and the at least one system component are assembled on the substrate.

[0005] According to a third aspect of the present technology, a chip-containing product is provided that comprises a system of the second aspect assembled on a further substrate together with at least one other product component.

[0006] According to a fourth aspect of the present technology, storing data in a plurality of registers, the plurality of registers including at least one array register including a plurality of array regions; receiving an issued instruction using a processing circuit and executing a processing operation identified in the issued instruction; storing one or more instructions in a buffer circuit prior to execution by the processing circuit; using a control circuit to, in response to receiving an instruction that requires access to two or more of the plurality of array regions, decomposing the instruction into two or more execution parts, each of the two or more execution parts corresponding to one of the two or more array regions; for each execution part of the two or more execution parts, delaying issuing the execution part as one of the issued instructions until it is predicted that the execution part can be processed hazard-free by the processing circuit, the control circuit being capable of issuing the two or more execution parts of the instruction in different cycles selected based on when each of the two or more execution parts is predicted to be processable hazard-free. A method is provided that includes delaying.

[0007] According to a further aspect of the present technology, a non-transitory computer-readable storage medium is provided for storing computer-readable code for manufacturing an apparatus, the apparatus comprising: A plurality of registers configured to store data, the plurality of registers including at least one array register including a plurality of array regions, A processing circuit configured to receive an issued instruction and execute a processing operation identified in the issued instruction, A control circuit including a buffer circuit configured to store one or more instructions prior to execution by the processing circuit, the control circuit, in response to receiving an instruction that requires access to two or more of the plurality of array regions, Decompose the instruction into two or more execution parts, each of the two or more execution parts corresponding to one of the two or more array regions, For each execution part of the two or more execution parts, delay issuing the execution part as one of the issued instructions until it is predicted that the execution part can be processed hazard - free by the processing circuit, and the control circuit can issue the two or more execution parts of the instruction in different cycles selected based on when each of the two or more execution parts is predicted to be processed hazard - free. A control circuit is provided.

Brief Description of the Drawings

[0008] The present invention will be further described by way of example only with reference to those configurations shown in the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5a

Figure 5b

Figure 5c

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0009] Before considering the configuration with reference to the accompanying drawings, the following description regarding the configuration is provided.

[0010] According to some exemplary configurations, an apparatus is provided that includes a plurality of registers configured to store data. The plurality of registers includes at least one array register that includes a plurality of array regions. The apparatus also includes a processing circuit configured to receive issued instructions and execute the processing operations identified in the issued instructions. The apparatus also includes a control circuit that includes a buffer circuit for storing one or more instructions prior to execution by the processing circuit. The control circuit, in response to receiving an instruction that requires access to two or more of the plurality of array regions, decomposes the instruction into two or more execution parts, each of the two or more execution parts corresponding to one of the two or more array regions, and for each execution part of the two or more execution parts, delays issuing the execution part as one of the issued instructions until it is predicted that the execution part can be processed hazard-free by the processing circuit, and the control circuit is capable of issuing the two or more execution parts of the instruction in different cycles selected based on when each of the two or more execution parts is predicted to be able to be processed hazard-free.

[0011] (The apparatus may be a data processing apparatus) The apparatus includes a plurality of registers for storing data accessed by a processing circuit when processing instructions. The registers may be provided closer to the processing circuit than a data cache or main memory, and may store operands being operated on by the currently processed instruction and the results of the currently processed instruction. Each register may include a plurality of scalar registers and / or a plurality of vector registers each storing a plurality of data elements. The registers also include an array register which is a large register including a plurality of data elements which may be a plurality of scalar elements or a plurality of vector elements. The array register includes a plurality of array regions. These regions are separate from each other and each may include a storage device for a plurality of data elements. The processing circuit is provided to access the array register in response to a received instruction that requires (specifies) one or more data elements stored in the array register. Some instructions may require a single region of the array register, and other instructions may require a plurality of regions of the array register. The required regions of the array register may be used to store operands used in a processing operation or results generated by a processing operation.

[0012] Some processing instructions may specify a majority (e.g., all or half) of the array registers. Alternatively, or in addition, some processing instructions may specify a much smaller portion of the array registers, e.g., a processing instruction may specify a single element of the array register. To prevent underutilization of the device, scheduling techniques may be used that attempt to schedule instructions so that two consecutively issued instructions do not require access to the same element within the array register and can thus be processed in parallel. The inventors have recognized that even when using such scheduling techniques, instructions that require access to a large amount (e.g., all or half) of the array registers may still stall after instructions that require access to a small amount of the array registers, which means that most of the array registers and associated processing circuitry are not used while waiting for instructions that require a small amount of the array registers to complete. As a result, the utilization rate of the array registers may vary significantly between different instructions. This level of variation in the utilization rate of the array registers can result in both fluctuations in the power consumption of the array registers that can interfere with nearby processing circuitry and large redundancies within the device that result in sub-optimal throughput.

[0013] The apparatus includes a control circuit having a buffer circuit for storing one or more instructions prior to execution. The control circuit may be configured to identify instructions that require access to the array registers and, in response to the identification of those instructions, may be configured to decompose those instructions for execution on different portions of the array registers. For example, the decomposition may be a dynamic decomposition that occurs at execution time based on, for example, one or more fixed or variable criteria. The decomposition of the instructions is based on the region of the array registers that the instructions require access to. In other words, the instructions are decomposed into execution parts based on the location of the data that is required (used as an operand when processing that instruction or generated as a result when processing that instruction). Different execution parts may require access to one, a subset, or all of the array regions depending on the particular data specified by the instructions.

[0014] When an instruction is decomposed into a plurality of execution parts, the execution parts can be issued for independent execution from each other. In particular, when an instruction requires access to two or more of a plurality of array regions, the instruction is divided into two or more execution parts, each of which corresponds to one of the two or more array regions. The two or more execution parts can then be executed independently of each other. For example, the two or more execution parts may be issued for simultaneous execution. Alternatively, the two or more execution parts may be issued during different cycles. This alternative may occur, for example, because it is predicted that the processing operation before accessing one of the array regions will be completed before the processing operation before accessing another array region.

[0015] The timing of issuance of each of the execution parts is speculative and is based on a prediction of when that execution part can be executed hazard-free by the processing circuit. An instruction is assumed to be capable of hazard-free execution when issued in order with respect to other instructions that write to the same array region. Since the prediction is executed independently for each of the array regions, different execution parts may be issued for independent execution at different times from each other. Thus, an instruction that requires access to a large number of array registers that may have previously remained stalled after an instruction that accesses a single data element (or a single array region) of the array register (i.e., requires access to two or more array regions of the array register) can proceed in at least one of the array regions. As a result, the overall instruction throughput of the device can be improved, and fluctuations in power consumption of the processing circuit can be reduced.

[0016] The array register may be a one-dimensional register, for example, a vector register, and each of the array regions may correspond to a sub-vector of the array register. In some configurations, the array register is a two-dimensional array register that stores data elements identified by coordinates in a first array direction and a second array direction. The array register may be physically arranged as a two-dimensional structure having spatial coordinates corresponding to the coordinates used to identify the data elements, but this is not necessary. Rather, the dimensions referred to when describing the array register should be understood to refer to the number of indices required to identify elements within the array register. For example, a two-dimensional array register should be understood as an array register whose elements are accessed using two indices. Similarly, the coordinates referred to as being used to access the data elements should be understood to refer to the indices used to access an array register that can be physically configured in any form, for example, as a one-dimensional physical array of elements or as a two-dimensional physical array of elements. In some configurations, each of the array regions is a one-dimensional array region. In other configurations, each of the array regions is a two-dimensional array region, for example, that can be used to store a matrix representation of data elements.

[0017] In some configurations, the resubdivision of the array register to an element may only include resubdividing the array along a single one of the two array dimensions (directions). In some configurations, the array register is resubdivided into array regions in both the first array direction and the second array direction. The array regions may be the same size as each other, or may be different sizes. In some configurations, the array regions have the same number of data elements in the first array direction and the second array direction. In other configurations, the array regions have different numbers of data elements in the first array direction and the second array direction. In some configurations, the array region may be a contiguous group of data elements in at least one of the array directions. In some configurations, the array region may include a non-contiguous group of data elements in at least one of the array directions. In some configurations, the array region may include an interleaved grouping of data elements in at least one of the array directions.

[0018] In some configurations, the number of sub - divisions of the array register in the first array direction is different from the number of sub - divisions of the array register in the second array direction. Sub - dividing the array by different numbers of sub - divisions in each of the array directions can be beneficial for certain types of instructions. The array register may be capable of storing different types of data elements (e.g., integer data elements, double - precision data elements, floating - point data elements, etc.). Each of these types of data elements may require a different number of bits of the array register to correctly represent its data type. For example, the array register may be configured to store M adjacent 8 - bit elements (where M is a positive integer) in the first array direction and N rows of these M 8 - bit elements adjacent to each other in the second array direction. The array register may also be configured to store M / 2 adjacent 16 - bit elements in the first direction and N rows of these M / 2 16 - bit elements adjacent to each other in the second array direction. As a result, the utilization rate of the array in one direction may depend on the type of elements being stored. Thus, by having fewer array subdivisions in the first array direction, the control circuit may be able to better distribute the processing load while ensuring that a single data element is not split across multiple array regions.

[0019] In some configurations, an instruction identifies at least one range of positions within an array register, and a control circuit is configured to predict, for each execution portion, a delay before the execution portion can be processed hazard-free by a processing circuit, the delay including at least a predicted number of remaining instruction cycles for the processing circuit to process a currently executing instruction that specifies a position within the array register that overlaps with at least one range of positions. The range of positions may span one or more of the array regions and may vary for different instructions, depending on the use case. For example, sequential instructions each specifying different positions within the same array region may be issued for parallel execution (assuming no other instructions utilize these positions). As another example, sequential instructions that span multiple array regions and specify positions that overlap within one array region may be issued for parallel execution in some of the array regions, but may not be issued within the one array region where they overlap. The control circuit may be configured to record information identifying the predicted number of remaining instructions until each position within the array register becomes available for execution of the next instruction, and the prediction of the delay may then be calculated based on this information. This information may be stored on a per-element basis or based on groups of elements. As a result, the control circuit can estimate the delay until the execution portion corresponding to each array region becomes executable, for each array region. Sequential execution portions of different instructions each specifying non-overlapping positions may be assigned the same predicted delay since neither causes a data hazard to the other. This approach enables the control circuit to predict the delay before the next instruction is issued.

[0020] In some configurations, the control circuit is configured to store an indication of the number of pending instruction cycles required to process each execution portion, and the delay includes at least the predicted number of remaining instruction cycles and the number of pending instruction cycles required to process the previous execution portion stored in a buffer circuit that designates a position within an array register that overlaps at least one position range. In addition to making a prediction based on the predicted number of remaining instruction cycles, some configurations also utilize the number of pending instruction cycles required to process each execution portion that designates a position range that overlaps with an instruction. In other words, for each pending execution portion that designates a position range, the control circuit identifies the predicted number of remaining instruction cycles for that position range and the number of pending instruction cycles required to process the previous execution portion that also designates that position range. This approach enables the control circuit to predict the delay of some pending instructions, including the next pending instruction and subsequent instructions received by the control circuit while still being stored while waiting for the next pending instruction to be issued.

[0021] In some configurations, the control circuit is configured to determine an unnecessary array region of an array register that does not overlap at least one position range and to omit generating an execution portion corresponding to the unnecessary array region. In other words, the control circuit is configured to decompose an instruction only into execution portions corresponding to array regions that overlap at least one position range. For array regions that do not overlap at least one position range, the corresponding execution portions are not generated. This approach ensures that unnecessary execution portions that could otherwise be discarded at the issue stage are not generated.

[0022] In some configurations, the control circuit is configured to maintain a plurality of delay counters, the plurality of delay counters indicating a delay for one of two or more execution portions, and each of the plurality of delay counters maintaining a negative count of the predicted number of remaining instruction cycles in response to a determination that the predicted number of remaining instruction cycles is less than zero. The control circuit may have delay counters assigned to each execution unit within each array region. Thus, the number of counters required can be equal to the number of array regions multiplied by the number of execution portions that can be stored in each array region. The count may be represented by any numerical system capable of storing negative numbers. For example, the counter may use a one's complement representation, a two's complement representation, or a representation that includes a sign bit indicating whether the count value is positive or negative. By maintaining delay counters capable of representing negative values, the control circuit can select when to issue an execution portion based on additional delay information that may be available from the processing circuit or other locations in the device.

[0023] The additional delay may be a general delay representing a delay associated with the entire processing circuit (i.e., the additional delay information may not be specific to an individual array region). In some configurations, the delay includes a predicted array-region-specific delay for one of two or more array regions corresponding to the execution portion. The delay may represent additional delay information that is the result of one or more stall events within the array region, one or more power-related events within the array region, and / or one or more global events specific to the entire array register. By incorporating predictor array-region-specific delays into the delay, the control circuit can improve the scheduling of execution portions for each of the array regions.

[0024] In some configurations, the processing circuit is configured to issue feedback information indicating the number of instruction cycles between the reception of an issued instruction and the instruction cycle in which the issued instruction can be executed hazard-free, and the control circuit is configured to adjust the predicted array-region specific delay based on the feedback information. The predicted array-region specific delay is maintained by the control circuit and may represent an estimated value of the delay associated with the corresponding array region.

[0025] In some configurations, the processing circuit issues feedback information requesting a reduction in delay in response to an identification that an issued instruction may be executed hazard-free earlier than it is received. When the feedback information issued by the processing circuit indicates that the number of instruction cycles between the reception of the issued instruction and the instruction cycle in which the issued instruction is executed is small and that the issued instruction may have been issued previously (e.g., because the issue of the execution part is too late), and that the issued instruction may have been executed previously, the control circuit may reduce the predicted array-region specific delay. If the issue of the execution part as an issued instruction is too late, there may be too few instructions for the processing circuit to process, and the utilization of the processing circuit may be insufficient. By providing feedback between the processing circuit and the control circuit, the control circuit may be able to improve the utilization rate of the processing circuit.

[0026] In some configurations, in response to an identification that an issued instruction was received too early to be hazard-free and executed immediately, the processing circuit stalls the issued instruction and issues feedback information to request an increase in latency. If the feedback information issued by the processing circuit indicates the number of instruction cycles between the reception of the issued instruction (i.e., the execution part issued as an instruction issued to that array region) and the large instruction cycle of the issued instruction (for example, because the issuance of the execution part was too early), the control circuit may increase the predicted array-region-specific latency and decrease the number of execution parts issued as the issued instruction. Issuing the execution part as an issued instruction too early may result in an increase in the number of issued instructions queued in the processing circuit, which may lead to stalls in the processing circuit. By providing feedback between the processing circuit and the control circuit, the control circuit may be able to reduce the likelihood of stalls occurring.

[0027] In some configurations, the request for increased latency includes stall information indicating the number of stall cycles in which the issued instruction was stalled, and the control circuit determines whether to adjust the array region specific delay predicted based on the global congestion metric in response to stall information indicating that the number of stall cycles is less than a predetermined threshold, and the control circuit adjusts the array region specific delay predicted independently of the global congestion metric in response to stall information indicating that the number of stall cycles exceeds a predetermined threshold. The feedback information may be provided as a 2-bit value. The 2-bit value may represent four different feedback signals including an indication that the reception of the issued instruction was too late, an indication that the issued instruction was received on time, a weak indication that the reception of the instruction was too early (stall information indicating that the number of stall cycles is below a predetermined threshold), and a strong indication that the reception of the instruction was too late (stall information indicating that the number of stall cycles exceeds a predetermined threshold). The control circuit may be configured to ignore the weak indication, for example, based on the number of received instructions and how full the buffer circuit is. The control circuit may not be able to ignore the strong indication and may be configured to adjust the array region specific delay predicted independently of the fullness of the buffer circuit in response to the strong indication. In some configurations, this adjustment of the buffer circuit may cause the buffer circuit to fill and a stall signal may have to be sent to the circuit providing instructions to the control circuit.

[0028] The processing circuit may be provided as a global processing circuit capable of accessing each of the array regions. In some configurations, the processing circuit is configured as a processing array divided into a plurality of processing regions, each configured to perform a processing operation on a corresponding array region among the plurality of array regions. The processing array may include one or more common processing portions capable of receiving operands from other array regions and passing results or data as operands used in other array regions. The processing array may also be provided in combination with a global processing circuit and may be configured such that the processing load is shared between the processing array and the global processing circuit according to a specific access pattern and specific issued instructions being executed.

[0029] In some configurations, the processing circuit is capable of simultaneously executing different instructions in different array regions among the plurality of array regions. Due to various timing constraints associated with instructions issued in different array regions among the plurality of array regions, the delays associated with each of those array regions may be different. As a result, different issued instructions may be executed simultaneously with each other in different regions of the processing array.

[0030] In some configurations, when storing instructions in the buffer circuit causes the buffer to overflow, the control circuit signals a stall to the front-end circuit that provides instructions to the control circuit in response to the instructions. The front-end circuit may include a decoder circuit configured to decode instructions before they are passed to the control circuit. The decoder circuit may receive instructions of an instruction set architecture and may be configured to decode each of these instructions into one or more micro-operations passed to the control circuit.

[0031] Here, a specific configuration will be described with reference to the figures.

[0032] FIG. 1 schematically shows an apparatus 10 according to some configurations of the present technology. The apparatus 10 includes a control circuit 16, a register 14, and a processing circuit 12. In addition to the register 14, the apparatus 10 includes an array register 18 divided into a plurality of array regions including a first array region 18(A), a second array region 18(B), a third array region 18(C), and a fourth array region 18(D). The register 14 and the array register 18 are provided to store data values accessed by the processing circuit 12 during processing of instructions and to store the results of the processing operations by the processing circuit 12. The control circuit 16 is provided to receive instructions and decompose those instructions into execution parts. For example, the control circuit may receive an instruction that requires access to the first array region 18(A) of the array register 18, and the control circuit may decompose the instruction into a single execution part. As another example, the control circuit may receive an instruction that requires access to the third array region 18(C) and the fourth array region 18(D), and decompose the instruction into two execution parts. The control circuit 16 is configured to delay the issuance of the execution parts until it is predicted that the execution parts can be processed hazard-free by the processing circuit. As a result, the control circuit 16 may issue the execution parts of an instruction decomposed into two or more execution parts in different cycles based on when each of those execution parts is predicted to be able to be processed hazard-free.

[0033] The layout of the illustrated components is provided for illustrative purposes only, and it will be readily apparent to those skilled in the art that the physical arrangement of the circuits may be provided differently. In fact, the control circuit, the processing circuit, and the register may be provided as separate circuit blocks, or alternatively, one or more separate circuit blocks that provide the functions of the control circuit, the processing circuit, and the register together may be provided.

[0034] Figure 2 schematically shows an array register provided by some configurations of the present technology. The array register includes a plurality of different array positions. In the illustrated array register, the array register is a 32×32 array of positions, and each position within the array register may contain a data element. The positions within the array register are identified using two index values such that the positions within the array register can be specified by indices in a first direction and a second direction.

[0035] In the embodiment of Figure 2, the array register can be addressed in various ways as follows: - In the examples of ZA6H.D[0], ZA0H.H[7], ZA2H.S[5], ZA12H.Q[1], a horizontal slice of data from a single array row of the array register is accessible as a vector operand having different vector element sizes represented by.D,.H,.S,.Q notations, - In the example of ZA0V.B

[22] , a vertical slice of data from a given column position

[22] within the array register is accessible as a vector operand. Here too, similar to the horizontal slices shown in Figure 2, it is possible to provide different data element sizes for the vector operand accessed as a vertical slice. - As shown in the examples of ZA7V.D[3], ZA3V.S[4], ZA1V.H[1], ZA8V.Q[0], it is also possible to access each part of the tile as a single vector operand to a tile of elements extracted from different rows of the array register. For example, ZA7V.D[3] includes four sets of 8 elements selected from column positions [31:24] of each of rows 7, 15, 23, and 31, ZA3V.S[4] includes eight sets of 4 elements selected from column positions [19:16] of each of rows 3, 7, 11, 15, 19, 23, 27, and 31, ZA1V.H[1] includes 16 sets of 2 elements selected from column positions [3:2] of each of the odd rows of the array register, and ZA8V.Q[0] includes two sets of 16 elements selected from column positions [15:0] of rows 8 and 24 of the array register, respectively.

[0036] It will be understood that this is merely a subset of the available addressing patterns.

[0037] The array register is divided into a plurality of array regions. In the illustrated configuration, the array register is divided into two parts horizontally and eight parts vertically. Each array region can be identified based on a region index i indicating the horizontal array region position and a region index j indicating the vertical array region position. The array is divided into more regions in the vertical direction than in the horizontal direction. Since the array register can store data elements in different formats (which may require different numbers of bits), data elements stored in some formats may span multiple positions in the array register horizontally. By simply dividing the array register once horizontally, data elements of different sizes can be processed without the individual data elements being divided across different array regions.

[0038] During processing, each of the array regions is accessed when the processing instruction specifies a position within the array register that overlaps with that array region. As an example, the above-described exemplary addressing schemes may each be decomposed into a plurality of execution parts that access two or more of the array regions. Using the notation {i,j} to refer to a region, it is as follows. -ZA7V.D[3] overlaps with array regions {1,1}, {1,3}, {1,5}, and {1,7}, -ZA0V.B

[22] and ZA3V.S[4] overlap with array regions {1,0}, {1,1}, {1,2}, {1,3}, {1,4}, {1,5}, {1,6}, {1,7}, and {1,8}, -ZA8V.Q[0] overlaps with array regions {0,2} and {0,6}, -ZA6H.D[0] overlaps with array regions {0,1} and {1,1}, -ZA0H.H[7] overlaps with array regions {0,3} and {1,3}, -ZA0H.B

[20] and ZA2H.S[5] overlap with array regions {0,5} and {1,5}, -ZA12H.Q[1] overlaps with array regions {0,7} and {1,7}.

[0039] In response to receiving an instruction to access an array region using the above address specification method, the control circuit decomposes the instruction into a plurality of execution parts, one for each array region to be accessed. Each of the execution parts may be issued independently for execution (i.e., regardless of whether the other execution parts are ready for execution) if it is predicted that those execution parts can be executed hazard-free.

[0040] In an alternative configuration, the array region may be an interleaved region. For example, region j = 0 may include rows 0, 8, 16, and 24, region j = 1 may include rows 1, 9, 17, and 25, region j = 2 may include rows 2, 10, 18, and 26, region j = 3 may include rows 3, 11, 19, and 27, region j = 4 may include rows 4, 12, 20, and 28, region j = 5 may include rows 5, 13, 21, and 29, region j = 6 may include rows 6, 14, 22, and 30, region j = 7 may include rows 7, 15, 23, and 31, and region j = 7 may also include rows 7, 15, 23, and 31. Alternative methods of dividing the array register into array regions will be readily apparent to those skilled in the art and may be selected based on the usage case of a particular device.

[0041] FIG. 3 schematically shows details of the control circuit 16 and the processing circuit 12 according to some configurations of the present technology. The control circuit 16 is provided with a decomposition circuit 24, a buffer circuit 22 configured as a plurality of queues, an issue circuit 20, a region delay memory circuit 26, and a tracking circuit 28. The processing circuit 12 is configured to execute processing in each of a plurality of regions of the array register 18. In the illustrated configuration, the array register 18 includes eight different regions 18(A) to 18(H).

[0042] The area delay memory circuit 26 is configured to store information indicating a predicted additional delay associated with each of the array areas 18. The tracking circuit 28 is configured to store a predicted delay associated with the execution part issued to the processing circuit. The decomposition circuit 24 is configured to receive a series of instructions and identify instructions that require access to the array register. The decomposition circuit is configured to identify which array area of the array register the instruction requires access to, and is configured to decompose the instruction into a plurality of execution parts. Each execution part is stored in a queue corresponding to the identified array area in the buffer circuit 22, together with an indication of the number of predicted cycles required to process the execution part. The predicted execution cycle number may be obtained, for example, from a look-up table (not shown). The decomposition circuit is configured to identify a delay for each of the execution parts. The predicted delay for each execution part is equal to the maximum number of instruction cycles indicated in the tracking circuit 28 for the area overlapping the position identified within the execution part, added to the number of instruction cycles of the instruction associated with the execution part preceding the current execution part in the buffer circuit. The delay may be indicated using a counter.

[0043] The control circuit is configured to update (i.e., decrease) the delay indicated in the buffer circuit 22 and the delay indicated in the tracking circuit 28 for each cycle. Further, the control circuit updates the delay indicated in the area delay memory circuit 26 in response to feedback from the processing circuit.

[0044] The issue circuit 20 is configured to add the delay associated with the execution part in the buffer circuit to the region-specific delay stored in the region delay memory structure 26. When the total delay (the sum of the delay related to the execution part and the region-specific delay) is zero or less, the issue circuit issues the execution part to the processing circuit 12 as an issued instruction that requires access to the corresponding array region 18. When an instruction is issued, the issue circuit 20 also updates the delay stored in the tracking circuit to increase the delay associated with the position specified by the execution part. The delay increases by the number of cycles predicted to be required by the issued instruction.

[0045] By storing the tracking information 28, the delay information associated with the execution part, and the information indicating the predicted number of cycles of the execution part, the decomposition circuit 24 can calculate the estimated delay before each new execution part is issued as an instruction for execution. By combining this with the per-region delay obtained from the region delay memory structure 26, the control circuit can maintain a stable flow of instructions issued to each region of the processing circuit.

[0046] Figure 4 schematically shows the decomposition of instructions into execution parts by the decomposition circuit 24. The decomposition circuit receives a series of instructions, Instruction A, Instruction B, and Instruction C. The instructions are received in that order, and each instruction includes information identifying the position within the array register that needs to be accessed. Instruction A includes position information A 30(A), Instruction B includes position information B 30(B), and Instruction C includes position information C 30(C). As a first step, the decomposition circuit receives the instructions and decomposes them into execution parts. In the illustrated configuration, Instruction A specifies a single array region and is decomposed into a single execution part stored in the memory circuit 22. Instruction B specifies a row of the array region and is decomposed into a corresponding execution part arranged in the memory circuit 22 behind Instruction A. Instruction C specifies a pair of positions and is decomposed into two execution parts arranged in the memory circuit 22. Since Instruction A and Instruction B each specify one common array region, one of the execution parts associated with Instruction B is arranged behind the execution part associated with Instruction A in the memory circuit 22.

[0047] Figures 5a-5c schematically show a series of counters stored by a control circuit to track when each execution part is issued. Each of FIGS. 5a-5c shows a set of instructions queued in buffer circuit 22 for a single array region (array region {0,0}) and tracking information for region {0,0} stored in tracking circuit 28. The figure also illustrates information showing the per-region delay for each of the four array regions stored in per-region tracking circuit 26 (in the illustrated configuration).

[0048] FIG. 5a shows the tracked information in the first cycle, cycle n. In the first cycle, the per-region delay for region {0,0} is 2 cycles, the per-region delay for region {0,1} is 0 cycles, and the per-region delays for regions {1,0} and {1,1} are 1 cycle. Tracking information 28 stores information indicating the predicted number of remaining instruction cycles for the processing circuit to process the currently executing instruction. The tracking information is provided per position for each position within the array region. In the illustrated configuration, array region {0,0} is divided into 16 positions arranged in a 4×4 grid and addressed using horizontal indices X0-X3 and vertical indices X0-X3. Each of the positions [X0,Y0]-[X3,Y0] (corresponding to row Y0) has 4 predicted remaining execution cycles, each of the positions [X0,Y1]-[X3,Y1] (corresponding to row Y1) has 1 predicted remaining execution cycle, each of the positions [X0,Y2]-[X3,Y2] (corresponding to row Y0), as well as positions [X0,Y3] and [X1,Y3], has -2 predicted remaining execution cycles, and each of the positions [X2,Y3]-[X3,Y3] has -5 predicted remaining execution cycles. Note that a negative number of execution cycles is interpreted as indicating that issued instructions predicted to prevent hazard-free execution in those regions would have been expected (predicted) to have completed in the absence of per-region delay.

[0049] Memory circuit 22 is currently tracking three instructions: instruction F, instruction G, and instruction H. Instruction F indicates that access to the array region (X0,2,Y2,2) is required, starting at position [X0,Y2] and accessing positions within the array with a width of 2 horizontally and 2 vertically. The instruction is associated with a predicted delay until the specified region is predicted to be available, in this case -2 instruction cycles, and is stored. The instruction is also associated with the number of cycles predicted to be taken during execution, in this case 4 cycles, and is stored. Instruction G indicates that access to the array region (X0,4,Y1,1) is required, starting at position [X0,Y1] and accessing positions within the array with a width of 4 horizontally and 1 vertically. The instruction is associated with a predicted delay until the specified region is predicted to be available, in this case 1 instruction cycle, and is stored. The instruction is also associated with the number of cycles predicted to be taken during execution, in this case 2 cycles, and is stored. Instruction H indicates that access to the array region (X0,4,Y0,1) is required, starting at position [X0,Y0] and accessing positions within the array with a width of 4 horizontally and 1 vertically. The instruction is associated with a predicted delay until the specified region is predicted to be available, in this case 4 instruction cycles, and is stored. The instruction is also associated with the number of cycles predicted to be taken during execution, in this case 2 cycles, and is stored.

[0050] In the illustrated configuration, the delay associated with instruction F is negative, indicating that, conditional on the per-region delay, this execution portion may be eligible to be issued for execution as an issued instruction. The per-region delay for region {0,0} is 2, so the total predicted delay before the execution portion corresponding to instruction F can be issued is -2 + 2 = 0, indicating that instruction F can be issued as an issued instruction.

[0051] Figure 5b schematically shows the states of the memory circuit 22 for the region {0,0}, the per-region delay memory circuit 26, and the tracking circuit 28. When an instruction F is issued for execution, it is removed from the memory circuit 22. The region {0,0} tracking information stored in the tracking circuit 28 is updated to add a counter associated with the position specified by the instruction F by the number of cycles predicted for the instruction F to take. Instruction F specified the positions (X0,2,Y2,2), i.e., the four positions of [X0,Y2], [X0,Y3], [X1,Y2], and [X1,Y3]. Each of these positions is updated to indicate that the instruction takes 2 cycles - already storing a counter value of -2. The delay associated with each position within the tracked region (array region {0,0}) decreases by only 1 cycle as the instruction counter advances.

[0052] Figure 5c schematically shows cycle n+2, the next instruction cycle in which a new instruction, instruction I, is added to the memory circuit 22. Instruction I specifies the position (X0,4,Y1,1) and is predicted to require 4 cycles. The position accessed by instruction I is equal to the position accessed by instruction G. Thus, the total delay before instruction I can be issued as an issued instruction is equal to the predicted delay shown in the tracking circuit 28, i.e., -1 cycle + the total number of cycles predicted for instruction G, i.e., 2 cycles, resulting in a 1-cycle delay. The delay associated with the counter is decremented by 1 each as a result of the cycle increment.

[0053] In this way, the delay associated with each instruction can be tracked, maintaining the array-region-specific delay, and buffering instructions in the memory circuit for a longer or shorter time depending on how quickly the processing circuit can process those instructions.

[0054] In the illustrated configuration of FIGS. 5a to 5c, it can be seen that there is some redundancy in storing the number of delay cycles in both the memory circuit 22 and the trace circuit 28. This redundancy exists to clarify the description and to explain the delays present in each instruction. The delays shown in the memory circuit 22 can be omitted in some configurations, and it will be readily apparent to those skilled in the art that they can be inferred from the trace information 28 at each stage. Further, it will be readily apparent that instructions from the memory circuit can be issued in order or out of order in some configurations, based on the way the instructions are traced and information is stored. Further, the trace information 28 may be provided as coarse-grained tracing or fine-grained tracing. Alternatively, rather than providing trace information 28 for each region, the memory circuit 22 may hold the trace information associated with issued instructions over at least the predicted number of cycles associated with those instructions, and may estimate the number of predicted execution cycles from the held trace information.

[0055] Figure 6 schematically shows the details of instruction issuance by the control circuit 16 according to some configurations of the present technology. The control circuit 16 includes a storage circuit 22 that stores information indicating an execution portion that is delayed before being passed to the processing circuit to be processed within the processing area 18(A). When the delay associated with the execution portion ends, the processing area 18(A) receives the issued instructions issued by the control circuit 16. The processing area 18(A) undergoes power management and receives power management control information that triggers the processing area to operate in different power states. For example, the processing area 18 may respond to the power management information by moving to a lower throughput state. As a result, the speed at which the processing area 18(A) can process instructions decreases. In such a situation, the speed at which the issued instructions are processed by the processing area decreases, and the issued instructions can be queued in the buffer circuit 32 provided by the processing circuit. In the illustrated configuration, the processing circuit can buffer three issued instructions in three different buffer slots including a first buffer slot 32(A), a second buffer slot 32(B), and a third buffer slot 32(C). The processing area 18(A) transmits feedback information requesting an increase in the per-region delay in response to the determination that the issued instructions are buffered for one or more cycles before execution. Similarly, the processing area 18(A) transmits feedback information requesting a decrease in the per-region delay in response to the buffer circuit 32 becoming empty. In this way, the processing circuit and the control circuit 16 can cooperate to manage the instruction throughput. The processing circuit may be configured to transmit feedback information in each cycle. The feedback information may be used to maintain a predetermined number of instructions in the buffer slot. For example, the processing circuit may aim to keep one buffer slot full and may transmit a 2-bit value as feedback information to trigger the control circuit to update the per-region delay. When all of the buffer slots are empty, the processing circuit may transmit "00" as feedback information to indicate that the issued instructions are not being provided at a sufficiently fast speed.When one of the buffer slots is in use, the processing circuit may send "01" as feedback information to indicate that the issued instruction is being provided at the correct speed. When two of the buffer slots are in use, the processing circuit may send "10" as feedback information to provide a weak indication that the issued instruction is being provided too rapidly. When all three of the buffer slots are in use, the processing circuit may send "11" as feedback information to provide a strong indication that the issued instruction is being provided too rapidly. It will be readily apparent to those skilled in the art that more or fewer buffer slots may be provided and that the thresholds and timings for the feedback information may be adapted based on the implementation.

[0056] Figure 7 schematically shows a series of steps performed by an apparatus according to some configurations of the present technology. The flow begins at step S70, where an instruction designating two or more array regions within the array register is received. Next, the flow proceeds to step S72, where the instruction is decomposed into two or more execution parts, each of the two or more execution parts corresponding to one of the designated array regions of the array register. Next, the flow proceeds to step S74, where the execution parts are stored in a buffer circuit and a delay counter is set for each execution part to delay the issuance of the execution part until it is predicted that the execution part can be processed hazard-free.

[0057] Figure 8 schematically shows a series of steps executed by a device according to some configurations of the present technology. The flow starts at step S80, where the variable p is set to 0 corresponding to the array region of the array register. Next, the flow proceeds to step S82, where it is determined whether the delay count associated with region P indicates that the execution part within that array region can be executed hazard-free. In step S82, if it is determined that there is no execution part that can be executed hazard-free, the flow proceeds to step S90, where the execution part is delayed before the flow proceeds to step S86. In step S82, if the delay counter indicates that there is an execution part that can be issued hazard-free, the flow proceeds to step S84, where the execution part is issued regardless of whether other execution parts of the same instruction are being issued in the same cycle. Next, the flow proceeds to step S86, where it is determined whether there is still an array region. In step S86, if it is determined that there is still an array region, the flow proceeds to step S88, where after P is incremented, the flow returns to step S82. In step S86, if it is determined that there are no more array regions, the flow proceeds to step S92, where it is determined whether it is the next execution cycle. In step S92, if it is determined that it is not the next execution cycle, the flow proceeds to step S94, where it waits for the next cycle before returning to step S92. In step S92, if it is the next execution cycle, the flow proceeds to step S96, where the delay counter is decremented before the flow returns to step S80. Those skilled in the art will readily understand that considering different array regions sequentially is for illustrative purposes only, and the regions may be considered either sequentially or in parallel.

[0058] The concepts described herein may be embodied in a system comprising at least one packaged chip. The foregoing apparatus is implemented within at least one packaged chip (either implemented within one particular chip of the system or distributed across two or more packaged chips). The at least one packaged chip is assembled on a substrate together with at least one system component. The chip-containing product may comprise a system assembled on a further substrate together with at least one other product component. The system or chip-containing product may be assembled within a housing or on a structural support (such as a frame or blade).

[0059] As shown in FIG. 9, one or more packaged chips 400 having the foregoing apparatus implemented on one chip or distributed across two or more chips are manufactured by a semiconductor chip manufacturer. In some embodiments, the chip product 400 manufactured by the semiconductor chip manufacturer is provided as a semiconductor package comprising a semiconductor device implementing the foregoing apparatus and a connector such as a land, ball, or pin for connecting the semiconductor device to an external environment, and a protective casing (e.g., made of metal, plastic, glass, or ceramic). When two or more chips 400 are provided, these may be provided as separate integrated circuits (provided as separate packages) or packaged by a semiconductor provider (e.g., by using an interposer or by providing a multi-chip semiconductor package comprising two or more vertically stacked integrated circuit layers using three-dimensional integration) into a multi-chip semiconductor package.

[0060] In some embodiments, a collection of chiplets (i.e., small modular chips having a particular function) may itself be referred to as a chip. The chiplets may be individually packaged within a semiconductor package and / or packaged together with other chiplets into a multi-chiplet semiconductor package (e.g., by using an interposer or by providing a multi-chiplet product with two or more vertically stacked integrated circuit layers using 3D integration).

[0061] One or more packaged chips 400 are assembled on a substrate 402 together with at least one system component 404 to provide a system 406. For example, the substrate may include a printed circuit board. The substrate material may be made of various materials, such as plastic, glass, ceramic, or paper, a flexible substrate material such as plastic or a textile material. At least one system component 404 includes one or more external components that are not part of one or more packaged chips 400. For example, at least one system component 404 can include any one or more of the following. Another packaged chip (e.g., provided by a different manufacturer or produced on a different process node), an interface module, a resistor, a capacitor, an inductor, a transformer, a diode, a transistor, and / or a sensor.

[0062] A chip-containing product 416 is manufactured that includes a system 406 (including a substrate 402, one or more chips 400, and at least one system component 404) and one or more product components 412. The product components 412 include one or more additional components that are not part of the system 406. As a non-exhaustive list for this example, the one or more product components 412 can include user input / output devices such as a keypad, touch screen, microphone, loudspeaker, display screen, tactile device, wireless communication transmitter / receiver, sensor, actuator for actuating mechanical movement, thermal control device, additional packaged chips, interface module, resistor, capacitor, inductor, transformer, diode, and / or transistor. The system 406 and the one or more product components 412 may be assembled on a further substrate 414.

[0063] The substrate 402 or the further substrate 414 may be provided on or within a device housing or other structural support (e.g., a frame or blade) and provide a product that can be handled by a user and / or is intended for operational use by a person or enterprise.

[0064] System 406 or chip-containing product 416 may be at least one of an end-user product, a machine, a medical device, a computing or telecommunications infrastructure product, or an automation control system. For example, as a non-exhaustive list of this example, the chip-containing product can be any of the following. A telecommunications device, a mobile phone, a tablet, a laptop, a computer, a server (e.g., a rack server or a blade server), an infrastructure device, networking equipment, a vehicle or other automotive product, an industrial machine, a consumer device, a smart card, a credit card, smart glasses, an avionics device, a robotics device, a camera, a television, a smart TV, a DVD player, a set-top box, a wearable device, a household appliance, a smart meter, a medical device, a heating / lighting control device, a sensor, and / or a control system for controlling public infrastructure equipment such as a smart highway or a traffic signal.

[0065] The concepts described herein may be embodied in computer-readable code for manufacturing an apparatus that embodies the described concepts. For example, the computer-readable code may be used in one or more stages of a semiconductor design and manufacturing process, including an Electronic Design Atomation (EDA) stage, to manufacture an integrated circuit comprising an apparatus that embodies the concepts. The above computer-readable code may additionally or alternatively enable the definition, modeling, simulation, verification, and / or testing of an apparatus that embodies the concepts described herein.

[0066] For example, the computer-readable code for manufacturing an apparatus embodying the concepts described herein may be embodied in code that defines a hardware description language (HDL) representation of the concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus that embodies the concepts. The code may define an HDL representation of one or more logic circuits that embody the apparatus in an intermediate representation such as Verilog, SystemVerilog, Chisel, or Very High-Speed Integrated Circuit Hardware Description Language (VHDL), as well as FIRRTL. The computer-readable code may provide a definition for embodying the concepts using a system-level modeling language such as SystemC and SystemVerilog or other behavioral representations of the concepts that can be interpreted by a computer to enable simulation, functional and / or formal verification, and testing of the concepts.

[0067] Additionally or alternatively, the computer-readable code may define a low-level description of an integrated circuit component that embodies the concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. One or more netlists or other computer-readable representations of the integrated circuit components may be generated by applying one or more logic synthesis processes to the RTL representation to generate a definition for use in manufacturing an apparatus that embodies the present invention. Alternatively or additionally, one or more logic synthesis processes may generate a bitstream from the computer-readable code for configuring a field programmable gate array (FPGA) to embody the described concepts. The FPGA may be deployed for purposes of verification and testing of the concepts prior to manufacturing in an integrated circuit, or the FPGA may be deployed directly in a product.

[0068] The computer-readable code may include a mixture of code representations for manufacturing an apparatus, including, for example, one or more mixtures of RTL representations, netlist representations, or other computer-readable definitions used in semiconductor design and manufacturing processes for manufacturing an apparatus embodying the present invention. Alternatively or additionally, the concept may be defined in combination with a computer-readable definition used in semiconductor design and manufacturing processes for manufacturing an apparatus and computer-readable code defining instructions to be executed by the apparatus once manufactured.

[0069] Such computer-readable code may be disposed on any well-known transient computer-readable medium (such as wired or wireless transmission of code via a network), or a non-transient computer-readable medium such as a semiconductor, magnetic disk, or optical disk. An integrated circuit manufactured using the computer-readable code may include components such as a central processing unit, a graphics processing unit, a neural processing unit, a digital signal processor, or one or more of other components that individually or collectively embody the concept.

[0070] In summary, an apparatus, a system, a chip-containing product, a method, and a medium are provided. The apparatus includes a plurality of registers including at least one array register having a plurality of array regions. The apparatus includes processing circuitry that receives issued instructions and processes those instructions. The apparatus also includes control circuitry that, in response to receiving instructions that require access to two or more array regions, decomposes the instructions into two or more execution parts, each corresponding to one of the two or more array regions, and delays the issuance of the execution parts for each execution part until it is predicted that the execution part can be processed hazard-free. The control circuitry is capable of issuing two or more execution parts in different cycles based on when it is predicted that each execution part can be processed hazard-free.

[0071] In this application, the term "configured to" is used to mean that an element of a device has a configuration that enables it to perform a defined operation. In this context, "configuration" means a way of arranging or interconnecting hardware or software. For example, a device may have dedicated hardware that provides a defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not mean that some change must be made to the device element to provide the defined operation.

[0072] In this application, an enumeration of features preceded by the phrase "at least one of" means that any one or more of those features may be provided individually or in combination. For example, "at least one of: [A], [B], and [C]" encompasses any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), a combination of A and B (without C), a combination of A and C (without B), a combination of B and C (without A), or a combination of A, B, and C.

[0073] Exemplary configurations of the present invention have been described in detail herein with reference to the accompanying drawings, but it should be understood that the present invention is not limited to these precise configurations, and that various changes, additions, and modifications may be made herein by those skilled in the art without departing from the scope of the invention as defined by the appended claims. For example, various combinations of the features of the dependent claims may be made with the features of the independent claims without departing from the scope of the invention.

Claims

1. A plurality of registers configured to store data, including at least one array register including a plurality of array regions, the plurality of registers; A processing circuit configured to receive an issued instruction and execute a processing operation identified in the issued instruction; A control circuit including a buffer circuit configured to store one or more instructions prior to execution by the processing circuit, the control circuit, in response to receiving an instruction that requires access to two or more of the plurality of array regions, Decompose the instruction into two or more execution parts, each of the two or more execution parts corresponding to one of the two or more array regions, For each execution part of the two or more execution parts, delay issuing the execution part as one of the issued instructions until it is predicted that the execution part can be processed hazard-free by the processing circuit, and the control circuit is configured to issue the two or more execution parts of the instruction in different cycles selected based on when each of the two or more execution parts is predicted to be processed hazard-free. A control circuit, and an apparatus comprising the same.

2. The apparatus according to claim 1, wherein the array register is a two-dimensional array register that stores data elements identified by coordinates in a first array direction and a second array direction.

3. The apparatus according to claim 2, wherein the array register is re-divided into array regions in both the first array direction and the second array direction.

4. The apparatus according to claim 3, wherein the number of re-divisions of the array register in the first array direction is different from the number of re-divisions of the array register in the second array direction.

5. The instruction identifies at least one position range within the array register, The control circuit is configured to predict, for each execution part, a delay before the execution part can be processed hazard-free by the processing circuit, the delay including at least a predicted number of remaining instruction cycles for the processing circuit to process a currently executing instruction that designates a position within the array register that overlaps the at least one position range. The apparatus according to any one of claims 1 to 4.

6. The control circuit is configured to store an indication of the number of pending instruction cycles required to process each execution portion. The delay includes at least the sum of the predicted number of the remaining instruction cycles and the number of pending instruction cycles required to process the previous execution portion stored in the buffer circuit that specifies a position within the array register that overlaps with the at least one position range. The apparatus according to claim 5.

7. The control circuit is configured to determine an unnecessary array region of the array register that does not overlap with the at least one position range and omit generation of the execution portion corresponding to the unnecessary array region. The apparatus according to claim 5 or 6.

8. The control circuit is configured to maintain a plurality of delay counters, and the plurality of delay counters indicate the delay for one of the two or more execution portions. Each of the plurality of delay counters maintains a negative count of the predicted number of the remaining instruction cycles in response to a determination that the predicted number of the remaining instruction cycles is less than zero. The apparatus according to any one of claims 5 to 7.

9. The delay includes a predicted array region specific delay for one of the two or more array regions corresponding to the execution portion, and the apparatus according to any one of claims 5 to 8.

10. The processing circuit is configured to issue feedback information indicating the number of instruction cycles between reception of the issued instruction and an instruction cycle in which the issued instruction can be executed hazard-free. The control circuit is configured to adjust the predicted array region specific delay based on the feedback information. The apparatus according to claim 9.

11. The processing circuit issues the feedback information to request reduction of the delay in response to an identification that the issued instruction may be executed hazard-free earlier than it is received, and the apparatus according to claim 10.

12. The processing circuit stalls the issued instruction and issues the feedback information to request an increase in the delay in response to an identification that reception of the issued instruction is too early to be executed hazard-free immediately, and the apparatus according to claim 10 or 11.

13. The request for the increase in the delay includes stall information indicating the number of stall cycles in which the issued instruction was stalled, in response to the stall information indicating that the number of stall cycles is less than a predetermined threshold, the control circuit determines whether to adjust the predicted array region specific delay based on a global congestion metric, in response to the stall information indicating that the number of stall cycles exceeds the predetermined threshold, the control circuit adjusts the predicted array region specific delay regardless of the global congestion metric, The apparatus according to claim 11.

14. The processing circuit is arranged as a processing array divided into a plurality of processing regions, and each processing region is arranged to execute the processing operation for a corresponding array region among the plurality of array regions. The apparatus according to any one of claims 1 to 13.

15. The apparatus according to any one of claims 1 to 14, wherein the processing circuit is capable of simultaneously executing different instructions in different array regions among the plurality of array regions.

16. When storing the instruction in the buffer circuit causes the buffer to overflow, the control circuit signals a stall to a front-end circuit that provides instructions to the control circuit in response to the instruction. The apparatus according to any one of claims 1 to 15.

17. The apparatus according to any one of claims 1 to 16, implemented in at least one packaged chip, at least one system component, a substrate, and the at least one packaged chip and the at least one system component are assembled on the substrate. System.

18. A chip-containing product comprising the system according to claim 17, assembled on a further substrate together with at least one other product component.

19. Storing data in a plurality of registers, the plurality of registers including at least one array register including a plurality of array regions, storing, using a processing circuit to receive an issued instruction and execute a processing operation identified in the issued instruction, storing one or more instructions in a buffer circuit prior to execution by the processing circuit, In response to receiving an instruction that requires access to two or more of the plurality of array regions, using a control circuit, decomposing the instruction into two or more execution parts, each of the two or more execution parts corresponding to one of the two or more array regions; for each execution part of the two or more execution parts, delaying issuing the execution part as one of the issued instructions until it is predicted that the execution part can be processed hazard-free by the processing circuit, wherein the control circuit can issue the two or more execution parts of the instruction in different cycles selected based on when each of the two or more execution parts is predicted to be processed hazard-free; A method comprising.

20. A non-transitory computer-readable medium storing computer-readable code for manufacturing the apparatus according to any one of claims 1 to 18.