Performance measurement device and performance measurement method for VLIW processor

The proposed performance measurement technique for VLIW processors addresses the challenge of energy prediction by using a scale factor and a dedicated measurement device to accurately measure inter-instruction energy, enhancing prediction accuracy and efficiency.

WO2025143613A1PCT designated stage expired Publication Date: 2025-07-03SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019611
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2024-12-03
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing VLIW processors face challenges in accurately measuring energy consumption due to varying pipeline locations and lengths for different vector operations, making it difficult to predict energy usage for all possible combinations of instructions.

Method used

A performance measurement technique for VLIW processors that calculates inter-instruction energy by using a scale factor and a performance measurement device with components like a scale factor calculation unit, OP counter, NOP counter, scalar memory, and lookup table to account for the order of instructions and pipeline activation levels.

Benefits of technology

Enables fast and highly accurate energy prediction by considering the differences in pipeline stages, improving the accuracy of energy measurement and performance evaluation for VLIW processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019611_03072025_PF_FP_ABST
    Figure KR2024019611_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to performance measurement technology performed in a performance measurement device for a very long instruction word (VLIW) processor, the technology comprising the steps of: calculating a scale factor in an operation process based on operation instructions input to the performance measurement device; generating, on the basis of the scale factor, an operation (OP) count value corresponding to the operation instructions; and calculating the energy of the VLIW processor on the basis of the OP count value, wherein the energy may include inter-instruction energy for a plurality of vector operations of the operation instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Performance measurement device and performance measurement method for VLIW processors

[0001] The present invention relates to a VLIW (very long instruction word) type processor having multiple stages, and more particularly to a performance measurement technique for predicting energy usage of a VLIW processor.

[0002] This study is related to the Next-Generation Intelligent Semiconductor Technology Development (Design) (NO. 2710008363) project, which was supported by the National IT Industry Promotion Agency (NIPA) and funded by the Ministry of Science and ICT (Government) in 2024.

[0003] In addition, this study is related to the research project on information and communication broadcasting innovation talent training (R&D) (NO. RS-2023-00256081) supported by the Information and Communications Technology Planning and Evaluation Institute with funding from the Ministry of Science and ICT (government) in 2023.

[0004] For reference, this application claims priority to Korean Patent Application No. 10-2024-0082559, filed on June 25, 2024. The entire contents of that application, which serves as the basis for this priority claim, are incorporated herein by reference.

[0005] Performance measurement is typically conducted to evaluate the architecture of a developed processor and the efficiency of various programs running on the processor. Therefore, the market for performance measurement devices can be seen as growing in proportion to the steady development of next-generation processors.

[0006] Meanwhile, VLIW type processors can be said to be a type of processor that is constantly being designed in the current processor market where parallel operations have become widespread, and the market for this has also been active since its appearance in the 1980s.

[0007] In these VLIW processors, vector operations use different computational units depending on the instructions being executed, so the location and length of the pipeline they occupy vary. In this case, the activation level of each stage within the pipeline continuously changes, so the pipeline wire width tends to be very wide due to the characteristics of vector operations. In particular, when there are multiple vector operations, there is a difficulty in measuring the energy for all ordered pairs of combinations between instructions in order to accurately measure the energy.

[0008] The background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired in the process of deriving the present invention, and cannot necessarily be said to be publicly known technology disclosed to the general public prior to the application for the present invention.

[0009] (Prior art literature)

[0010] (Patent Document 0001) U.S. Patent Publication No. 11816490 (November 14, 2023)

[0011] In an embodiment of the present invention, by calculating inter-instruction energy for multiple vector operations based on a scale factor in an operation process based on an operation instruction of a VLIW (very long instruction word) processor and providing an energy prediction result in a VLIW type processor, a performance measurement technique is proposed that can take into account an operation environment that may occur depending on the order between instructions while adopting a method of multiplying the energy consumed per instruction by the number of times each instruction is executed.

[0012] The problems to be solved by the present invention are not limited to those mentioned above, and other problems to be solved that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below.

[0013] According to an embodiment of the present invention, a performance measurement method performed in a performance measurement device for a VLIW (very long instruction word) processor includes: a step of calculating a scale factor in an operation process based on an operation instruction input to the performance measurement device; a step of generating an operation (OP) count value corresponding to the operation instruction based on the scale factor; and a step of calculating energy of the VLIW processor based on the OP count value; wherein the energy includes inter-instruction energy for a plurality of vector operations of the operation instruction. The present invention can provide a performance measurement method for a VLIW processor.

[0014] Here, the method may further include, prior to the step of calculating the scale factor, a step of performing a NOP (no operation) operation on an unused slot to generate a NOP count value.

[0015] Additionally, the method may further include a step of storing the NOP count value and the OP count value in a scalar memory of the performance measurement device.

[0016] In addition, the method may further include, after the storing step is performed, a step of loading the OP count value and the NOP count value stored in the scalar memory; and a step of storing the loaded OP count value and the NOP count value in a vector register of the performance measurement device.

[0017] In addition, the step of calculating the energy may include a step of calculating the base energy and the inter-instruction energy of the VLIW processor by reflecting the OP count value and the NOP count value stored in the vector register into a look-up table of the performance measurement device.

[0018] In addition, the method further includes, prior to the step of calculating the scale factor, a step of determining whether a profiling mode signal applied to the performance measurement device is enabled and the operation command is committed, and the scale factor can be calculated when the operation command is committed.

[0019] Additionally, the loading step may be performed when the profiling mode signal is disabled.

[0020] Additionally, when the profiling mode signal is disabled, the OP count value and the NOP count value stored in the scalar memory can be stored in the vector register.

[0021] Additionally, in the scalar memory, the number of instructions for calculating the base energy and the number of instructions for calculating the inter-instruction energy can be allocated to separate spaces.

[0022] Additionally, the lookup table may store address information and a count value for recognizing the operation instruction of the vector register.

[0023] Additionally, the NOP counter may have the same number as the number of slots of the VLIW processor.

[0024] According to an embodiment of the present invention, a performance measurement device for a VLIW processor includes: a scale factor calculation unit for calculating a scale factor in an operation process based on an operation instruction input to the performance measurement device; and an OP counter for generating an OP count value corresponding to the operation instruction based on the scale factor; wherein, based on the OP count value, a performance measurement device for a VLIW processor can be provided for calculating inter-instruction energy of the VLIW processor for a plurality of vector operations for the operation instruction.

[0025] Here, the device may further include a NOP counter that performs a NOP operation for an unused slot and stores a NOP count value.

[0026] Additionally, the device may further include a scalar memory that stores the NOP count value and the OP count value.

[0027] Additionally, the device may further include a vector register that loads and stores the OP count value and the NOP count value stored in the scalar memory.

[0028] In addition, the device further includes a lookup table in which the OP count value and the NOP count value stored in the vector register are reflected, and the base energy and the inter-instruction energy of the VLIW processor can be calculated based on information of the lookup table.

[0029] Additionally, the device can determine whether a profiling mode signal applied to the performance measurement device is enabled and the operation command is committed, and calculate the scale factor when the operation command is committed.

[0030] Additionally, the device may store the OP count value and the NOP count value stored in the scalar memory in the vector register when the profiling mode signal is disabled.

[0031] Additionally, in the scalar memory, the number of instructions for calculating the base energy and the number of instructions for calculating the inter-instruction energy can be allocated to separate spaces.

[0032] Additionally, the lookup table may store address information and a count value for recognizing the operation instruction of the vector register.

[0033] Additionally, the NOP counter may have the same number as the number of slots of the VLIW processor.

[0034] According to an embodiment of the present invention, a computer-readable recording medium storing a computer program includes instructions for causing a processor to perform a performance measurement method performed in a performance measurement device for a VLIW processor, the method including: calculating a scale factor in an operation process based on an operation instruction input to the performance measurement device; generating an OP count value corresponding to the operation instruction based on the scale factor; and calculating energy of the VLIW processor based on the OP count value; wherein the energy may include inter-instruction energy for a plurality of vector operations of the operation instruction.

[0035] According to an embodiment of the present invention, a VLIW processor is equipped with its own performance measurement device, and based on this, the efficiency of the architecture and the efficiency of kernel programs can be evaluated simultaneously. In the case of existing analytical methods used to predict processor energy measurement, since the processor does not have an external counter, the number of times a specific instruction has been executed must be predicted before subsequent energy predictions can be made. Since the prediction is made in two steps, the accuracy must be considered to be low. Furthermore, energy predictions performed through HDL (hardware description language)-level or gate-level simulations can be very slow depending on the size of the processor, and despite this, the accuracy is also low, making it difficult to measure performance for various cases. The present invention is expected to enable fast and highly accurate energy predictions by utilizing an analytical performance prediction technique and the help of a performance measurement device within the processor.

[0036] FIG. 1 is a diagram illustrating one form of a VLIW processor that can be used in a performance measurement technique according to an embodiment of the present invention.

[0037] FIG. 2 is a block diagram illustrating the function of a performance measurement device for a VLIW processor according to an embodiment of the present invention.

[0038] FIG. 3 is a flowchart exemplarily illustrating a performance measurement method of a performance measurement device for a VLIW processor according to an embodiment of the present invention.

[0039] Figure 4 is a drawing structurally explaining the scale factor calculation process of Figure 3.

[0040] Figure 5 is a flowchart specifically explaining the process of storing a count value in the scalar memory of Figure 3.

[0041] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the scope of the present invention is defined solely by the claims.

[0042] In describing embodiments of the present invention, specific descriptions of known functions or configurations will be omitted unless actually necessary. Furthermore, the terms described below are defined based on their functions in the embodiments of the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.

[0043] The energy prediction method in the embodiment of the present invention includes calculating inter-instruction energy that may occur depending on the order between instructions, which can be easily overlooked, while adopting the conventional method of multiplying the energy consumed per instruction by the number of times each instruction is executed.

[0044] In addition, the embodiment of the present invention includes a function that considers the difference in energy consumption according to the degree of activation of a multi-stage pipeline, which was not considered in existing inventions, for the purpose of efficiency of implementation.

[0045] In the case of vector operations, the operation units used differ depending on the instruction being performed, so the location and length of the pipeline occupied are different. In this case, the degree of activation of each stage within the pipeline continuously changes, and due to the characteristics of vector operations, the width of the pipeline wire tends to be very wide.

[0046] If there are N vector operations in a combination of inter-instructions, then N 2 There are many combination pairs of pairs, and the difficulty is that the energy for all ordered pairs must be measured in order to accurately measure the energy.

[0047] Accordingly, in an embodiment of the present invention, a performance counter for a processor is proposed as a technology that can be applied to a VLIW (very long instruction word) type processor having a multi-stage pipeline.

[0048] It provides a two-faceted performance evaluation method when running a specific program on a target program: first, it counts the number of times a specific instruction is executed, and second, it provides a function to predict energy based on this count.

[0049] A differentiating feature of the performance measurement device in the embodiment of the present invention is that it calculates energy measurements by taking into account differences between instructions using different pipeline stages in a VLIW structure to increase prediction accuracy.

[0050] The processor in the embodiment of the present invention has a form called VLIW, which is also commonly referred to as a multi-slot form. A typical non-VLIW processor has instruction words that correspond one-to-one to instructions that perform a certain function, and the processor's decoder processes these sequentially. On the other hand, VLIW groups a certain number of instruction words together, which the decoder interprets simultaneously, then divides the long-length instruction into smaller chunks and delivers them to the appropriate computational unit.

[0051] A VLIW processor typically consists of a scalar slot, a vector slot, and a memory slot. If a particular VLIW processor includes one scalar slot, one vector slot, and two memory slots, decoding and computation can be performed in the form shown in Figure 1.

[0052] Programs that can run on the processor of Fig. 1 perform operations using scalars, vectors, and memory slots according to their respective purposes, and the form of energy consumption varies depending on the three types, and the method for resolving this is as follows.

[0053] First, scalar slots consume significantly less energy than other slots, making them less important for energy prediction accuracy. Simply multiplying the number of scalar operations performed by the energy consumption per operation provides a sufficient prediction.

[0054] Second, for memory slots, as with scalar slots, energy can be predicted by multiplying the number of memory operations performed by the energy consumption per operation. However, memory operations are divided into load and store series, and in all cases, energy consumption varies greatly depending on the arrangement of data bits. Therefore, it must be interpreted by dividing it into the following two parts: control path energy, which is fixedly generated by the operation itself, and data path energy, which varies depending on the arrangement of data bits.

[0055] Third, since the vector operation energy measured for a vector slot is also greatly affected by the arrangement of data bits, it is necessary to distinguish between control path energy and data path energy.

[0056] To further improve the accuracy of energy prediction, the following methods can be introduced.

[0057] First, memory operations consume almost the same amount of energy regardless of the combination of preceding and following operations, because operations corresponding to memory operations occupy pipelines of the same length within the processor.

[0058] On the other hand, in energy measurements for vector operations, the prediction error varies greatly depending on which vector operations preceded a specific vector operation (including NOP (no operation)). This is a problem that occurs because the location and length of the pipeline occupied by each vector operation are different.

[0059] The energy generated by the difference between operation combinations is defined as inter-instruction energy, and the scale factor was introduced to address this issue. The scale factor is a numerical value that reflects the additional energy consumption resulting from calculating the degree to which pipeline stages are activated and deactivated before and after performing different operations.

[0060] The number of all possible pairs of operations for N vector operations is N 2 Since it is a combination of N, it is difficult to measure the energy for all cases. Instead, if the inter-instruction energies occurring between a NOP operation and a specific vector operation are measured for N operations, the inter-instruction energy occurring between any two instructions can be estimated using this value and a scale factor. In the embodiment of the present invention, a performance measurement technique for a VLIW processor is proposed based on this point.

[0061] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.

[0062] FIG. 2 is a block diagram for explaining the function of a performance measurement device (100) for a VLIW processor according to an embodiment of the present invention.

[0063] As illustrated in FIG. 2, a performance measurement device (100) for a VLIW processor may include a decoder (102), a scale factor calculation unit (104), an OP counter (106), NOP counters (108-1 to 108-4), a scalar memory (110), a lookup table (112), and a vector register (114).

[0064] A VLIW operation instruction can be input to the decoder (102), and the operation instruction can be transmitted to the OP counter (106), NOP counters (108-1 to 108-4), and scalar memory (110) according to the profiling mode signal of the performance measurement device (100) for the VLIW processor.

[0065] The scale factor calculation unit (104) can calculate a scale factor in an operation process based on an operation instruction input to the performance measurement device (100) for a VLIW processor. Specifically, the performance measurement device (100) for a VLIW processor can determine whether the profiling mode signal applied to the decoder (102) is enabled and the operation instruction is committed, and when the operation instruction is committed, the scale factor can be calculated through the scale factor calculation unit (104). The scale factor calculation unit (104) calculates a scale factor according to the degree of pipeline activation when an operation process occurs in the pipeline. The scale factor can be calculated at the moment the operation instruction is committed. Thereafter, when counting, a value corresponding to a scale factor other than 1 is increased. The scale factor value cannot be greater than the number of pipeline stages.

[0066] The OP counter (106) can generate an OP count value corresponding to an operation instruction based on the scale factor calculated by the scale factor calculation unit (104). Here, the number of OP counters (106) is equal to the sum of the number of slots of the VLIW processor and the number of ports of the scalar memory (110). That is, although one OP counter (106) is illustrated in FIG. 1, this is merely an example, and the number of OP counters (106) can be designed differently depending on the number of slots of the VLIW processor.

[0067] The NOP counter (108-1 to 108-4) can be configured with, for example, four NOP counters (108-1 to 108-4), and can perform a NOP operation on unused slots of a VLIW processor to store the NOP count value. In the case of a fixed-length VLIW operation instruction, unused slots always involve a NOP operation, and the NOP counter (108-1 to 108-4) can be configured to accommodate such frequent occurrences. Here, the NOP counter (108-1 to 108-4) can be designed with the same number of NOP counters as the number of slots of the VLIW processor, for example.

[0068] The scalar memory (110) can store the OP count value of the OP counter (106) and the NOP count value of the NOP counter (108-1 to 108-4). Here, the VLIW processor loads and sets the previous count value stored in the scalar memory (110) at each commit point of the operation instruction for operation instructions other than the NOP count value, and then performs a procedure of adding 1 to the count value for the corresponding operation instruction. Here, the setting can be viewed as a concept of counting all operation instructions that appear in the program that is being run first by ISA type. However, since NOP appears frequently, it is counted through a separate counter (NOP counter), and other OPs share one counter based on the present invention. That is, when the decoder (102) decodes a specific OP A, it fetches the previous count value for OP A from the scalar memory (110), sets it in a register, and processes the count value as +1. Afterwards, when OP B is newly decoded, the counting value for OP B can be obtained from the scalar memory (110) in the same manner. At this time, since the number of ports that can access the scalar memory (110) at one time is equal to the number of slots of the scalar memory (110), the number of OP counters (106) for operations other than NOP operations can be designed equal to the number of ports of the scalar memory (110).

[0069] The lookup table (112) can reflect the OP count value of the OP counter (106) and the NOP count value of the NOP counter (108-1 to 108-4). For example, when program execution reaches the end point, energy is calculated based on the previously calculated count value. It is assumed that the base energy for each instruction and the inter-instruction energy occurring between the vector operation and the vector NOP have been measured in advance, and information about this can be stored in the lookup table (112).

[0070] The vector register (114) can load and store the OP count value of the OP counter (106) and the NOP count value of the NOP counters (108-1 to 108-4) stored in the scalar memory (110). The OP count value and the NOP count value stored in the vector register (114) can be reflected in the lookup table (112), and thus, the performance measurement device (100) for the VLIW processor can calculate the base energy and inter-instruction energy of the VLIW processor based on the information in the lookup table (112).

[0071] That is, at the end point of program execution, the profiling mode signal is disabled (changed from 1 to 0), at which time all count values ​​and NOP count values ​​stored in the scalar memory (110) can be loaded and sequentially stored in the vector register (114), and the base energy and inter-instruction energy can be sequentially calculated using the vector register (114) as an input to the lookup table (112).

[0072] The lookup table (112) may have values ​​stored in advance, and there may be two vector register values ​​(e.g., A and B) given as input. The values ​​stored in advance store information on energy consumed per OP, and the input value may correspond to an address (A; OP ID here) for accessing a lookup table entry, and in the case of B, may correspond to a value counted in advance. In other words, it can be viewed as a concept of retrieving the power consumption value per OP stored in the lookup table (112) through A, and then multiplying this value by the count value.

[0073] FIG. 3 is a flowchart exemplarily illustrating a performance measurement method of a performance measurement device (100) for a VLIW processor according to an embodiment of the present invention.

[0074] As illustrated in FIG. 3, when the profiling mode signal of the performance measurement device (100) for a VLIW processor is enabled, the performance measurement device (100) for a VLIW processor determines whether any operation instruction is committed (S100, S102).

[0075] When an operation instruction is committed, a performance measurement device (100) for a VLIW processor can calculate a scale factor in an operation process based on the operation instruction (S104).

[0076] Figure 4 is a drawing that structurally explains the scale factor calculation process.

[0077] As illustrated in Figure 4, when an operation instruction moves along a pipeline, a count number can also move along the pipeline.

[0078] If the pipeline stage where the operation instruction is located is active and the next instruction does not use the stage, the count value can be increased by one.

[0079] If the corresponding operation instruction commits past EX7, the result of the preceding instruction is located in counter #8, and the two values ​​can be used to compute the scale factor simultaneously.

[0080] Once such a scale factor is calculated, a performance measurement device (100) for a VLIW processor can generate a count value corresponding to an operation instruction based on the calculated scale factor (S106).

[0081] Afterwards, the performance measurement device (100) for the VLIW processor can store the generated count value in the scalar memory (110) (S108).

[0082] Figure 5 is a flowchart specifically explaining the process of storing a count value in the scalar memory (110) of Figure 3.

[0083] First, we assume the following about VLIW processors:

[0084] Let's assume that a VLIW processor consists of one scalar slot, one vector slot, and two memory slots, each of which receives a 32-bit instruction as input. Let's assume that the upper two bits are used to identify the slot, the next nine bits are used as the OPCODE, and the middle seven bits are used as the FUNC code.

[0085] The decoder (102) can distinguish commands by taking these 18 bits, which are referred to as OP_ID. The scalar memory (110) is configured as a space of appropriately large 4KB. Each address has a space of 4 bytes, and there are a total of 1,024 address spaces.

[0086] A VLIW processor has a pipeline with a total of eight stages. All scalar operations occupy the same number of pipeline stages, and all memory operations also occupy the same number of pipeline stages. Only vector operations are assumed to occupy different locations and numbers of pipeline stages.

[0087] A separate scalar memory space is designated for profiling, separate from the scalar memory area used by the program. This location is accessed using PROF_ADDR. PROF_ADDR can be implemented as a separate register or stored within scalar memory. The space for profiling can be defined as shown in [Table 1] below.

[0088] Address space (PROF_ADDR) Number of OPs (#OP) Number of inter-instruction OP traces (#IOP) (4 x 2) OP_ID & Counter for NOP (#OP x 2) OP_ID & Counter per OP ~ Reverse order from the end of memory address (#IOP x 2) OP_ID & Counter per OP

[0089] <Perform profiling mode>

[0090] When the program to be profiled is run, the decoder (102) extracts the OP_ID for each slot.

[0091] When the decoding process starts, all control passes are stalled (S200).

[0092] If a NOP word exists, the corresponding NOP counter is incremented by one in advance (S204). (Incremented during the decoding phase, not the commit phase.)

[0093] If N or more instructions are not NOPs, all control paths are stalled and space is allocated in the scalar memory (110). First, PROF_ADDR is accessed to obtain the number of OPs currently being tracked.

[0094] Afterwards, the loop is repeated as many times as the number of OPs and it is checked whether there is a counter with the same OP_ID as the current command (S206).

[0095] If the same OP_ID exists, the stall is released and the control pass is resumed (S208).

[0096] If it does not exist, access PROF_ADDR again, retrieve the current OP count from the main counter, increment it by one, and store it in the same address (S210). The location of the new address space can be calculated using the #OP before the increment, and the OP_ID value and the count value initialized to 0 are stored there.

[0097] If OP corresponds to a vector operation, a separate memory space must be allocated for inter-instruction calculations. In this case, space is allocated in reverse order from the last address of the scalar memory (110), and the initialization method is as above. #IOP is also increased by one along with #OP.

[0098] There may be cases where the number of OPs to be tracked becomes so large that the general OP area and the inter-instruction OP area intersect. In this case, profiling mode does not function properly, but assuming sufficient scalar memory, this case is not considered.

[0099] When an operation passes through the pipeline, the scale factor is calculated only for vector operations.

[0100] The scale factor is calculated according to the logic of the scale factor calculation unit (104), and the latter instruction among the instruction pairs for which inter-instructions must be calculated is used as the standard.

[0101] All non-NOP instructions stall the control path at the commit stage, then retrieve the counter stored from the scalar memory (110), increment it by 1, and store it back in memory. In one cycle, the stall by the decoder (102) precedes the stall by the commit.

[0102] At this time, for vector operations, the counter is increased by a scale factor other than 1.

[0103] Repeating the above process will cause the program to run very slowly in profiling mode. It should be noted that while accessing scalar memory (110) during execution, scalar slots are used, and this is not reflected in the profiling results.

[0104] Referring again to FIG. 3, the performance measurement device (100) for a VLIW processor determines whether the profiling mode signal is disabled (S110), and if the profiling mode signal is disabled, loads the count value of the scalar memory (110) and the count value of the NOP counter (108-1 to 108-4), and stores the loaded values ​​in the vector register (114) (S112, S114).

[0105] Thereafter, the performance measurement device (100) for the VLIW processor can calculate base energy and inter-instruction energy by reflecting the value of the vector register (114) in the lookup table (112). This is described in detail as follows.

[0106] <Performing energy calculation process>

[0107] Upon program execution, the algorithm applied to the embodiments of the present invention is executed to calculate the final result. The algorithm is a means for predicting performance using the program being measured in the embodiments of the present invention, and such an algorithm can be exemplified as follows.

[0108] 1 ENERGY_PROFILE_WRAP_UP:

[0109] 2 / Read from scalar memory

[0110] 3 base_count = SCALAR_MEM(PROF_ADDR)

[0111] 4 inter_count = SCALAR_MEM(PROF_ADDR+1)

[0112] 5 base_begin = PROF_ADDR + 2

[0113] 6 inter_begin = END_ADDR - inter_count * 2

[0114] 7 / Move NOP counter value to the scalar memory

[0115] 8 SCALAR_MEM(base_begin+1) = NOP_COUNTER [0]

[0116] 9 SCALAR_MEM(base_begin+3) = NOP_COUNTER [1]

[0117] 10 SCALAR_MEM(base_begin+5) = NOP_COUNTER [2]

[0118] 11 SCALAR_MEM(base_begin+7) = NOP_COUNTER [3]

[0119] 12 / Calculate performance for base energy

[0120] 13 idx = 0

[0121] 14 for i in Ceil(base_count+4, VEC_ELEM_SIZE):

[0122] 15 for j in VEC_ELEM_SIZE:

[0123] 16 VEC_REG0 [j] = SCALAR_MEM(base_begin+idx*2+0)

[0124] 17 VEC_REG1 [j] = SCALAR_MEM(base_begin+idx*2+1)

[0125] 18 LUT_SET(VEC_REG0)

[0126] 19 perf = LookUp(VEC_REG0, VEC_REG1)

[0127] 20 Store(perf)

[0128] 21 / Calculate performance for inter-instruction energy

[0129] 22 idx = 0

[0130] 23 for i in Ceil(inter_count, VEC_ELEM_SIZE):

[0131] 24 for j in VEC_ELEM_SIZE:

[0132] 25 VEC_REG0 [j] = SCALAR_MEM(inter_begin+idx*2+0)

[0133] 26 VEC_REG1 [j] = SCALAR_MEM(inter_begin+idx*2+1)

[0134] 27 LUT_SET(VEC_REG0)

[0135] 28 perf = LookUp(VEC_REG0, VEC_REG1)

[0136] 29 Store(perf)

[0137] In the above algorithm, lines 16 to 17 and 25 to 26 exemplify the process of taking each counter from the scalar memory (110) and sequentially inserting it into the vector register (114).

[0138]

[0139] The scalar memory (110) allocates separate space for profiling, with the first address space storing the number of base instructions and the second address space storing the number of instructions for which inter-instructions must be calculated (vector operations only).

[0140] A space is provided for storing a NOP counter in the subsequent location of the scalar memory (110). Assuming that the total number of slots of VLIW is four, two address spaces are allocated per slot, for a total of eight address spaces.

[0141] In the subsequent locations of the scalar memory (110), two memory spaces are also allocated for each new OP. The first address space stores an ID value so that the OP can recognize itself, and the second address space stores a count value.

[0142] The final performance calculation program transfers the count information of all stored OPs to a vector register (114) and multiplies the previously calculated energy per operation using a lookup table (112). At this time, it is assumed that the vector unit supports floating-point operations.

[0143] Before energy calculation, it is assumed that the base energy for N instructions and the inter-instruction energy for M vector instructions are pre-calculated.

[0144] The process for calculating base energy and inter-instruction energy is essentially the same. However, the two results are not combined in the final result.

[0145] The OP_ID and the corresponding counter value are retrieved from the scalar memory (110) and transferred to the vector register (114).

[0146] When the vector register (114) is filled, the lookup table (112) is first set using the vector register (114) filled with OP_ID. This means the process of obtaining the existing calculation result value (energy per instruction) corresponding to the OP_ID.

[0147] Once the lookup table (112) is set, the lookup table (112) is accessed using the counter value, and a multiplication operation is performed between the result from the lookup table (112) and the vector register (114) filled with the counter.

[0148] This result allows us to predict the final energy consumption.

[0149] According to the embodiments of the present invention described above, a VLIW processor can be equipped with its own performance measurement device, enabling simultaneous evaluation of the efficiency of the architecture and kernel programs. Existing analytical methods used to predict processor energy measurements require first predicting the number of times a specific instruction has been executed, as the processor lacks an external counter. This requires prior prediction of the number of times a specific instruction has been executed before subsequent energy predictions can be made. This two-step prediction process can result in reduced accuracy. Furthermore, energy predictions performed through hardware description language (HDL) or gate-level simulations can be significantly slower depending on the processor's size. Nevertheless, accuracy remains low, making it difficult to measure performance across a wide range of scenarios. The present invention utilizes an analytical performance prediction technique, utilizing the performance measurement device within the processor to enable fast and highly accurate energy predictions. The present invention was developed for an AI accelerator currently being designed or being developed for its next version, which is a VLIW processor. The proposed device is expected to provide an indicator for selecting the next version architecture by suggesting a performance evaluation method for the processor itself.

[0150] Meanwhile, the combinations of each block in the attached block diagram and each step in the flowchart may be performed by computer program instructions. These computer program instructions may be installed in a processor of a general-purpose computer, special-purpose computer, or other programmable data processing equipment, so that the instructions, executed by the processor of the computer or other programmable data processing equipment, create a means for performing the functions described in each block in the block diagram.

[0151] These computer program instructions may also be stored in a computer-usable or computer-readable recording medium (or memory) that can direct a computer or other programmable data processing equipment to implement a function in a specific manner, so that the instructions stored in the computer-usable or computer-readable recording medium (or memory) can also produce a manufactured item that includes instruction means for performing the function described in each block of the block diagram.

[0152] And, since the computer program instructions can also be installed on a computer or other programmable data processing equipment, a series of operation steps are performed on the computer or other programmable data processing equipment to create a process that is executed by the computer, so that the instructions that execute the computer or other programmable data processing equipment can also provide steps for executing the functions described in each block of the block diagram.

[0153] Additionally, each block may represent a module, segment, or portion of code that includes at least one executable instruction for performing a specific logical function(s). It should also be noted that in some alternative embodiments, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on their respective functions.

Claims

1. A performance measurement method performed on a performance measurement device for a VLIW (very long instruction word) processor, A step of calculating a scale factor in an operation process based on an operation command input to the above performance measurement device; A step of generating an OP (operation) count value corresponding to the operation instruction based on the scale factor; and A step of calculating the energy of the VLIW processor based on the OP count value; Including, The above energy includes inter-instruction energy for multiple vector operations of the above operation instructions. Performance measurement methods for VLIW processors.

2. In paragraph 1, Before calculating the above scale factor, It further includes a step of generating a NOP count value by performing a NOP (no operation) operation on an unused slot. Performance measurement methods for VLIW processors.

3. In paragraph 2, Further comprising a step of storing the NOP count value and the OP count value in a scalar memory of the performance measurement device. Performance measurement methods for VLIW processors.

4. In paragraph 3, After completing the above saving steps, A step of loading the OP count value and the NOP count value stored in the scalar memory; and further comprising a step of storing the loaded OP count value and the NOP count value in a vector register of the performance measurement device; Performance measurement methods for VLIW processors.

5. In paragraph 4, The steps for calculating the above energy are: A step of calculating the base energy and the inter-instruction energy of the VLIW processor by reflecting the OP count value and the NOP count value stored in the vector register to a look-up table of the performance measurement device. Performance measurement methods for VLIW processors.

6. In paragraph 5, Before calculating the above scale factor, Further comprising a step of determining whether a profiling mode signal applied to the above performance measurement device is enabled and the above operation instruction is committed, When the above operation command is committed, the scale factor is calculated. Performance measurement methods for VLIW processors.

7. In paragraph 6, The above loading steps are: The above profiling mode signal is disabled when it is in progress. Performance measurement methods for VLIW processors.

8. In paragraph 6, When the above profiling mode signal is disabled, the OP count value and the NOP count value stored in the scalar memory are stored in the vector register. Performance measurement methods for VLIW processors.

9. In paragraph 5, In the above scalar memory, the number of instructions for calculating the base energy and the number of instructions for calculating the inter-instruction energy are allocated to separate spaces. Performance measurement methods for VLIW processors.

10. In paragraph 5, The above lookup table stores address information and count values ​​for recognizing the above operation instruction of the vector register. Performance measurement methods for VLIW processors.

11. In paragraph 2, The above NOP counter has the same number of slots as the number of VLIW processors. Performance measurement methods for VLIW processors.

12. In a performance measurement device for VLIW processors, A scale factor calculation unit that calculates a scale factor in an operation process based on an operation command input to the above performance measurement device; and An OP counter for generating an OP count value corresponding to the operation instruction based on the scale factor; Based on the above OP count value, the inter-instruction energy of the VLIW processor for multiple vector operations for the above operation instruction is calculated. A performance measurement device for VLIW processors.

13. In paragraph 12, It further includes a NOP counter that performs a NOP operation on unused slots and stores the NOP count value. A performance measurement device for VLIW processors.

14. In paragraph 13, Further comprising a scalar memory for storing the above NOP count value and the above OP count value. A performance measurement device for VLIW processors.

15. In paragraph 14, Further comprising a vector register for loading and storing the OP count value and the NOP count value stored in the scalar memory. A performance measurement device for VLIW processors.

16. In paragraph 15, Further comprising a lookup table reflecting the OP count value and the NOP count value stored in the vector register, Calculating the base energy and the inter-instruction energy of the VLIW processor based on the information of the lookup table. A performance measurement device for VLIW processors.

17. In paragraph 16, The profiling mode signal applied to the above performance measurement device is enabled, the operation instruction is determined to be committed, and if the operation instruction is committed, the scale factor is calculated. A performance measurement device for VLIW processors.

18. In paragraph 17, When the above profiling mode signal is disabled, the OP count value and the NOP count value stored in the scalar memory are stored in the vector register. A performance measurement device for VLIW processors.

19. In paragraph 16, In the above scalar memory, the number of instructions for calculating the base energy and the number of instructions for calculating the inter-instruction energy are allocated to separate spaces. A performance measurement device for VLIW processors.

20. In paragraph 16, The above lookup table stores address information and count values ​​for recognizing the above operation instruction of the vector register. A performance measurement device for VLIW processors.

21. In paragraph 13, The above NOP counter has the same number of slots as the number of VLIW processors. A performance measurement device for VLIW processors.

Citation Information

Patent Citations

  • VLIW processor accepting branching to any instruction in an instruction word set to be executed consecutively

    US6615339B1