Performance analysis of architectural models using processor architectural designs

By generating pseudo-instructions based on software traces and using architectural models to simulate each stage of the processor architecture, the problems of inaccurate and time-consuming simulation in existing technologies are solved, and fast and accurate processor architecture performance analysis and on-chip system performance evaluation are achieved.

CN120826682APending Publication Date: 2025-10-21SYNOPSYS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480016164.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-20
Filing Date
2024-02-19
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing processor architecture performance analysis methods have problems such as inaccurate simulation, long time consumption, and difficulty in expansion and modification, especially for performance comparison of different architectures and new architecture design.

Method used

An architectural model is used to simulate the execution of pseudo-instructions. By generating pseudo-instructions based on software traces, the performance analysis of the processor architecture is achieved. The architectural model includes interconnected instances of model objects in the control flow graph, representing each stage of the processor architecture. It is independent of the specific instruction set architecture and is suitable for performance comparison of different processor architectures.

Benefits of technology

It provides accurate processor architecture performance information, reduces simulation time, improves scalability and modification efficiency, supports rapid performance research, and is suitable for performance analysis of on-chip systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120826682A_ABST
    Figure CN120826682A_ABST
Patent Text Reader

Abstract

One example is a method. The control flow graph is populated by instances of the model object. Each instance of the model object instances represents a respective stage of a processor architecture design. Instances of the model objects are interconnected in the control flow graph. An interconnect instance of model objects is an architectural model representing a processor architectural design. Comprising interconnected instances of model objects is output by one or more processors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to generating performance information for a processor architecture. Background Art

[0002] In processor design, determining the performance of a processor architecture can be one step in the design process. Simulations based on the processor design can be performed. The results of the simulations can indicate whether the processor design meets various design specifications or whether the processor design should be improved. An iterative process of creating or modifying the processor architecture and simulating the processor design can be implemented. This type of approach can lead to a satisfactory processor architecture design ready for tapeout and manufacturing, which can be very expensive. Summary of the Invention

[0003] An example is a method. A control flow graph is populated by instances of a model object. Each instance of the model object represents a corresponding stage of a processor architectural design. The instances of the model object are interconnected in the control flow graph. The interconnected instances of the model object constitute an architectural model representing the processor architectural design. The architectural model, including the interconnected instances of the model object, is output by one or more processors.

[0004] Another example is a non-transitory computer-readable medium. The non-transitory computer-readable medium includes stored instructions. When executed by a processor, the instructions cause the processor to: populate a control flow graph with instances of a model object, interconnect the instances of the model objects in the control flow graph, and output an architecture model including interconnected instances of the model objects. Each instance of the model object represents a corresponding stage of a processor architecture design. The interconnected instances of the model objects represent the architecture model of the processor architecture design.

[0005] Another example is a method. Using one or more processors, an architectural model is generated that represents a processor architecture. The architectural model includes interconnected instances of model objects in a control flow graph. Each of the instances represents a corresponding stage of the processor architecture. A software trace comprising instructions is received. The instructions are converted into instruction set architecture (ISA)-agnostic pseudo-instructions. Execution of the pseudo-instructions is simulated using the architectural model. Performance information for a first processor architecture is generated based on the simulation. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present disclosure will be more fully understood from the detailed description and exemplary drawings given below. The drawings are intended to provide knowledge and understanding of the examples and are not intended to limit the scope of the present disclosure to these specific examples. In addition, the drawings are not necessarily drawn to scale.

[0007] Figure 1 is an architectural model of a five-stage reduced instruction set computer (RISC) central processing unit (CPU) pipeline according to some examples.

[0008] Figure 2 is an architectural model of a superscalar processor according to some examples.

[0009] Figure 3 is a software trace generated by executing software for a processor according to some examples.

[0010] Figure 4 Based on some examples Figure 3 Pseudo-instructions generated by preprocessing the software trace.

[0011] Figure 5 and Figure 6 Based on some examples Figure 1 and Figure 2 The modified architectural model of the architectural model includes a model object mapped to a memory driver.

[0012] Figure 7 is a flow chart of a method for analyzing processor architecture performance using an architectural model, according to some examples.

[0013] Figure 8 Based on some examples from Figure 1 A graphical format of the performance information of the simulation of the architectural model.

[0014] Figure 9 is a flowchart of a method of generating an architectural model according to some examples.

[0015] Figure 10 is a flowchart of a method of generating pseudo-instructions from a software trace, according to some examples.

[0016] Figure 11 An architectural model of a simple RISC pipeline implemented in a SystemC environment according to an example.

[0017] Figure 12 is an architectural model of a first superscalar RISC pipeline implemented in a SystemC environment according to one example.

[0018] Figure 13 is an architectural model of a second superscalar RISC pipeline implemented in a SystemC environment according to one example.

[0019] Figure 14 is an example software trace from a cycle-accurate processor model according to an example.

[0020] Figure 15 is an example software trace from a loosely timed processor model according to an example.

[0021] Figure 16An example is shown Figure 13 Integration of the architectural model with the system-on-chip (SoC) hardware model.

[0022] Figure 17A and Figure 17B is a graphical format of performance information for architectural exploration indicating execution time according to an example.

[0023] Figure 18 is a graphical format of performance information for architectural exploration that indicates sequential and parallel execution of instructions, according to an example.

[0024] Figure 19 It is a representative diagram of the processor architecture.

[0025] Figure 20 is an architectural model of a processor architecture according to an example.

[0026] Figure 21A 、 Figure 21B 、 Figure 21C and Figure 21D SoC performance analysis of performance information and software benchmark traces of a processor on an architectural model of a processor architecture according to one example is shown in a graphical format.

[0027] Figure 22A 、 Figure 22B 、 Figure 22C and Figure 22D SoC performance analysis of performance information and software benchmark traces of an improved processor on an architectural model of the improved processor architecture according to one example is shown in a graphical format.

[0028] Figure 23 are flow charts of various processes used during the design and fabrication of integrated circuits, according to some examples.

[0029] Figure 24 is a diagram of an example computer system in which various examples may operate. DETAILED DESCRIPTION

[0030] Various aspects of the present disclosure relate to performing performance analysis using an architectural model of a processor architecture. Performance analysis can include, for example, simulating the execution of a software program or application on a processor model on a computer system to obtain performance information. Performance metrics of a processor architecture, such as cycles per instruction (CPI) or instructions per cycle (IPC), can be important for determining whether a processor architecture meets design specifications and / or for architectural exploration of different architectures. Various methods have been implemented for such performance analysis; however, such methods have technical limitations and problems.

[0031] One approach is based on untimed or loosely timed models. Untimed or loosely timed models may not have accurate timing in simulation and therefore may not be useful for performance analysis. Another approach uses statistical, or theoretical, models. Statistical models can be used for high-level architectural analysis, but may not be useful for performance analysis of processor architectures because the models cannot execute actual software programs or applications. Another approach implements a trace-driven model, which runs a trace of actual software execution instructions on the model. For some software applications, traces can be very large, which can result in long simulation times and slow down the simulation. Furthermore, traces may carry timing fingerprints specific to the processor model from which they were collected, and therefore may not be properly adapted to new processor architectures being analyzed. Another approach uses a task-based model, which represents processing time and memory traffic without implementing the processor's architectural details. Task-based models are generally not easily modified to represent different architectures. Furthermore, task-based models may not scale well to large traces. Another approach is based on cycle-accurate or approximate models. These models require detailed and complex modeling and are therefore difficult and time-consuming to implement. Cycle-accurate or approximate models typically require a software stack, which includes boot code and interrupt service routines, and typically require a compiler and tool chain. Furthermore, cycle-accurate or approximate models are available after the register transfer level (RTL) description of the processor has been implemented, which can defeat the purpose of architectural exploration.

[0032] This disclosure describes using an architectural model of a processor architecture to simulate the execution of pseudo-instructions to determine the performance of the processor architecture, wherein the pseudo-instructions are generated based on software traces and are agnostic to the instruction set architecture (ISA) of the processor architecture being modeled. The architectural model can include any number of model objects. The model objects include representations of stages of the processor architecture. Examples of stages include the fetch stage, decode stage, dispatch stage, execute stage, writeback stage, and completion stage. Examples of execution stages include the compute stage and the load-store stage. Depending on the type of stage represented, the model object can represent the control logic, timing, and / or memory access behavior of the stage. In some examples, the model object does not implement the non-memory access functional behavior of the stage, such as the computations of the compute stage. The model objects are connected in a control flow graph as an architectural model to represent the processor architecture. The processor architecture can include one or more pipelines, each of which includes various model objects. Traces (including executed instructions) obtained from another source are converted into pseudo-instructions, which can be architecturally agnostic. Then, execution of pseudo-instructions (e.g., instructions and / or pseudo-instructions flowing in a control graph) is simulated in the architectural model to determine performance information for performance analysis (e.g., CPI, IPC, number of parallel instructions executed, etc.) In some examples, the architectural model is extended to a system on a chip (SoC).

[0033] Technical advantages of the present disclosure include, but are not limited to, determining accurate performance information of a processor architecture, including control logic, timing, and memory access behavior, based on software traces that can be derived from different processor architectures. Using software traces from different processor architectures can eliminate the need for compilers and / or boot code for the processor architecture being modeled. In addition, implementing architectural models and pseudo-instructions can allow for lightweight simulations that are fast and easily scalable for large workloads. The architectural model can be easily reconfigured to create architectural variants for architectural exploration. In addition, the architectural model can be mapped into a hardware model of the SoC to determine workloads for the full system performance of the SoC. Other benefits and advantages can be achieved in various examples.

[0034] The architectural model of a processor is an execution model that can be created in any high-level hardware modeling environment or programming language (such as SystemC). The model objects in the architectural model represent the various pipeline stages of the processor architecture, such as acquisition, decoding, dispatch, calculation, load-store, writeback, and completion. The model objects represent the control logic, timing, and memory access behavior of the corresponding pipeline stage, and do not need to implement the non-memory access functional behavior of the corresponding pipeline stage. For example, in some examples, the computational model objects do not implement the actual calculations of the corresponding stage in the simulation. The lack of non-memory access functional behavior implementations may mean that the model objects cannot perform the actual processing or operation of the processor. However, the model objects can have accurate control logic, timing, and memory access information. Therefore, the model objects can be used for performance analysis of the processor.

[0035] For clarity, references to model objects herein, without a reference to a class qualifier or other context, refer to instances of model objects that can populate an architectural model. A model object class can broadly refer to a general definition of a class of model objects. Instances of model objects can include, for example, tasks in a task-based programming environment.

[0036] An architectural model can be agnostic about any particular instruction set architecture (ISA) implemented by the processor being simulated by the architectural model. Thus, an architectural model for a given processor can receive a software trace generated by executing software or processor instructions on another processor and can perform execution based on the software trace. This can allow a software trace obtained from one processor to be used as a workload for performance comparisons of different processor architectures, which may be existing or newly designed. Furthermore, architectural models can be easier to implement and modify in accordance with new architectural designs and can have much faster simulation speeds than fully functional, cycle-accurate models, which can allow for faster turnaround times for performance studies.

[0037] The control logic, timing, and memory access behavior of a processor pipeline can be modeled by one or more model objects in an architecture model. The control logic behavior of the pipeline can include instruction order, instruction dependencies, instruction parallelism, and stalls. Instruction ordering refers to the order in which instructions are fetched, decoded, executed, and / or completed. As defined by the processor architecture, instruction ordering can be ordered or unordered relative to the input program (e.g., trace) order. Instruction dependencies refer to the rules or conditions that govern when instructions can be executed based on the previous one or more instructions (one or more). Instruction parallelism refers to the number of instructions that are fetched, decoded, and / or executed simultaneously. Stalls can occur due to dependencies between instructions implemented by the ISA or the availability of hardware resources (e.g., buffers, execution units, etc.). The timing behavior of the pipeline can include, but is not limited to, the latency of each pipeline time segment in the pipeline stage within the pipeline, delays due to stalls, and delays due to memory access. The memory access behavior of the pipeline may include, but is not limited to, (i) the address, type (e.g., read or write), and size of memory accesses from a particular pipeline stage (e.g., fetch and / or load-store), and (ii) the number of parallel memory access requests.

[0038] Example stages of a processor pipeline include the fetch stage, the decode / dispatch stage, the execute stage, and the completion stage. Each stage can be modeled by a model object. The corresponding model object can represent control logic, timing, and / or memory access behavior and does not implement non-memory access functionality of the corresponding stage.

[0039] In a processor architecture, the fetch stage may read instructions from an instruction memory and maintain a program counter that points to the address of the memory from which the next instruction is to be read.

[0040] The acquisition model object reads instructions from the identified trace file (e.g., instead of reading from memory). Since the trace file provides a record of the program that has been executed, there is no need for an acquisition model object to maintain a program counter. For example, if it is necessary to model the memory traffic and latency of the acquisition phase, the acquisition model object can use the program counter given in the trace file to read the memory. However, in various examples, data read from memory may not be used. The acquisition model object reads instructions from the trace file and sends the read instructions to the decode and dispatch model object. The number of instructions sent per cycle can be defined by the processor architecture being modeled. If implemented, any delay between sending consecutive instructions can vary based on the read latency from memory.

[0041] In a processor architecture, the decode and dispatch stage typically decodes instructions fetched and sent from the fetch stage to determine the instruction type. If applicable, the decode and dispatch stage can read the values ​​of any operands. The decode and dispatch stage can check the dependencies of the instruction and, if any dependencies are met, send the instruction to the execution unit for execution. If the dependencies are not met, the decode and dispatch stage can save the instruction in a dispatch buffer and decode the next instruction, or can stall further decoding.

[0042] The decoding model object assigns a unique identifier (ID) to each instruction. The decoding model object also identifies the mnemonics in the received instructions according to the instruction set architecture (ISA) of the processor from which the trace was obtained. The decoding model converts the mnemonics into pseudo-mnemonics, and the decoding model attaches the pseudo-mnemonics to the instructions. The pseudo-mnemonics can be assigned based on the execution model object on which the instruction will be executed. This helps the dispatch model object determine which execution model object the instruction needs to be sent to. Because the model objects do not implement the non-memory access functionality of the corresponding stages, stages with the same control logic, timing, and memory access behavior can be represented by the same instance of the model object, even when the functionality implemented by these stages is different. Therefore, in some examples, instructions with different mnemonics in the trace file (and instructions that will be dispatched to different stages in actual execution) can be assigned the same pseudo-mnemonics and dispatched to the same instance of the execution model object.

[0043] Based on the incoming instruction, the decode model object also determines the dependency type of the instruction and identifies any other instruction(s) that the instruction depends on. The decode model object attaches to the instruction the ID(s) of the instruction(s) that the instruction depends on. The dependency IDs are used by the dispatch model object to determine when the instruction is ready to be dispatched for execution. For load / store type instructions, the memory address(es) and access size(s) are also determined by the decode model object from the instruction and attached to the instruction. The memory address(es) and access size(s) are used by the load / store unit to perform the memory operation.

[0044] The dispatch model object places the incoming instruction into an in-progress queue. The dispatch model object then checks the pseudo-mnemonics and the dependencies of the instructions in that queue, and if the dependencies have been satisfied (e.g., all other instructions that a given instruction depends on have completed), attempts to send the instruction to the specified execution model object. If the dependencies are not satisfied, the instruction waits in the in-progress queue until the instruction can be sent. Based on the processor architecture being modeled, an instruction may be sent out of order from the in-progress queue if the dependencies of the instruction preceding it are satisfied before the dependencies of the instruction preceding it are satisfied. Additionally, based on the processor architecture being modeled, more than one instruction may be sent to one or more execution units per cycle. After an instruction has been sent, it is marked as "SENT", but the instruction is not removed from the in-progress queue until the instruction completes execution as marked by the completion model object.

[0045] In processor architecture, the execution stage performs functional operations according to the instruction type and its parameters. A processor may include one or more execution units, which may be connected sequentially or in parallel.

[0046] Execution model objects can include or include various execution units based on the processor architecture. For simple processors like reduced instruction set computer (RISC) central processing units (CPUs), execution units can include arithmetic logic units (ALUs), multiplication units (MULs), load and store units (LOAD / STORE), branch units (BRANCH), floating-point units (FPUs), and so on. In the architectural model, the control logic and timing behavior of each unit are modeled. Non-memory access functional behaviors of the unit may not be modeled. For example, for an ALU or MUL, the actual calculation of operands and results is not modeled. Instead, the latency of the ALU or MUL is modeled. The latency of any unit can be fixed or variable. For example, for a LOAD / STORE operation, the latency of the load / store operation depends on the time it takes to read / write data from / to cache or memory, which can vary depending on the architecture and workload in the system. In such cases, the actual load / store operation can be performed by the LOAD / STORE, taking into account the latency encountered. For complex processors, such as graphics processing units (GPUs) or artificial intelligence (AI) engines, execution units can be a combination of various micro- or macro-operations (such as loads, computations, and stores), or can even include mini-pipelines within the execution unit. Such execution units can also be modeled as control and timing units.

[0047] In a processor architecture, the completion stage can store the results of executing instructions and retire instructions. The completion model object receives instructions from the execution model object and sends the instructions with a completion marker to the dispatch model object. Before sending, the instructions can be explicitly marked as "DONE" or a similar marker. The completed instructions are then retired (e.g., removed) from the in-progress queue in the dispatch model object. The retirement of instructions is typically performed in the order in which they are received from the acquisition model object to maintain correct program order.

[0048] Various model objects are connected into a control flow graph (e.g., a task graph) to represent the pipeline structure and topology of a given processor. Different types of topologies, from simple scalar pipelines (such as RISC CPUs) to complex superscalar pipelines (such as AI engines and GPUs), can be implemented using this technology.

[0049] Depending on the processor type and / or architecture being modeled, the number and type of pipeline stages may vary. As an example, Figure 1 1 is an architectural model 100 of a five-stage RISC CPU pipeline. The architectural model 100 includes an acquisition model object 102, a dispatch model object 104, a computation model object 106, a load / store model object 108, and a completion model object 110. In the architectural model 100, the execution unit is divided into two parallel model objects: the computation model object 106 (e.g., the ALU model object) and the load / store model object 108. As another example, Figure 2 2 is an architectural model of a superscalar processor. Architectural model 200 includes an acquisition model object 202, a dispatch model object 204, multiple parallel computation model objects 206, multiple parallel load / store model objects 208, and a completion model object 210. As shown in architectural model 200, a superscalar processor can have multiple parallel execution model objects of the same type (e.g., computation model objects 206 and load / store model objects 208). A superscalar processor and its corresponding architectural model can also have a branch prediction or prefetch stage and corresponding model objects to improve overall pipeline efficiency. An AI processor can treat each execution stage as a macro unit, which includes multiple subunits that perform basic load, compute, and store operations.

[0050] The architectural model presents a software trace as pseudo-instruction execution. A software trace can be obtained by executing the actual software binary on (i) any simulation model or development board with the desired processor, (ii) a variant of a processor with a similar or previous generation architecture, or (iii) a processor with a different architecture. The simulation model may or may not have actual or any instruction or pipeline timing for a specific architecture. The software trace may include the execution of relevant instructions without timing information.

[0051] The software trace can be preprocessed, and during the preprocessing, instructions of the software trace can be converted into pseudo-instructions based on the type of execution phase that the corresponding instructions will perform. For example, if the same control logic, timing and / or memory access are applied to different types of instructions, the different types of instructions can be converted into the same type of pseudo-instructions even without considering the underlying operations and / or operands of the instructions. Whether various types of instructions can be converted into the same type of pseudo-instructions can depend on the processor architecture being modeled. Two different types of instructions in one processor architecture being modeled can have the same control logic, timing and memory access, but the same two different types of instructions in another processor architecture being modeled can have different control logic, timing and / or memory access.

[0052] For example, for a RISC CPU, ALU and move instructions can be replaced by a single type of pseudo-instruction "ALU". Similarly, branch and jump instructions can be replaced by a single type of pseudo-instruction "BRANCH"; multiplication instructions can be replaced by pseudo-instructions "MUL"; load instructions can be replaced by pseudo-instructions "LOAD"; store instructions can be replaced by pseudo-instructions "STORE"; and so on. Similar conversions can be performed for various processor architectures, including AI and GPU processors, where each instruction can represent one or more compound operations such as convolution, matrix multiplication, load (e.g., data from double data rate (DDR) memory to static random access memory (SRAM)), store (e.g., data from SRAM to DDR), etc.

[0053] Furthermore, during preprocessing, instruction dependencies are determined. Dependencies can be attached to corresponding pseudo-instructions. As previously mentioned, preprocessing of software traces can be performed by decoding model objects.

[0054] Figure 3 is a software trace generated by executing software for a processor, and Figure 4 Shown are some examples from Figure 3 The corresponding pseudo instructions are generated by preprocessing the software trace. Figure 3 The software trace consists of seven instructions. Line 00 is a move instruction from a special register to a general register using the mnemonic MRS. Line 01 is a move instruction using the mnemonic MOV. Line 02 is a load register instruction using the mnemonic LDR. Line 03 is an unsigned bit field extract instruction using the mnemonic UBFX. Line 05 is a compare instruction using the mnemonic CMP. Line 06 is a branch if equal (BEQ) instruction using the mnemonic B.EQ.

[0055] refer to Figure 4 , preprocessed as Figure 3 The instruction at line 00 is assigned unique ID = 100, and the instruction is converted into an ALU pseudo-instruction. The instruction at line 01 is assigned unique ID = 101, and is converted into an ALU pseudo-instruction. The instruction at line 02 is assigned unique ID = 102, and is converted into a STORE pseudo-instruction with dependencies from pseudo-instructions with ID = 100, 101. The instruction at line 03 is assigned unique ID = 103, and is converted into a LOAD pseudo-instruction with dependencies from pseudo-instructions with ID = 101, 102. The instruction at line 04 is assigned unique ID = 104, and is converted into an ALU pseudo-instruction with dependencies from pseudo-instructions with ID = 103. The instruction at line 05 is assigned unique ID = 105, and is converted into an ALU pseudo-instruction with dependencies from pseudo-instructions with ID = 104. The instruction at line 06 is assigned unique ID = 106, and is converted into a BRANCH pseudo-instruction with dependencies from pseudo-instructions with ID = 105. Furthermore, pseudo instruction ID=102, 103 includes the corresponding address and size of the corresponding memory access of the pseudo instruction.

[0056] like Figure 3 and Figure 4 As shown in the pre-processing shown, the model object that performs pre-processing (e.g., decoding the model object) receives or determines the ISA of the software trace received by the model object and converts the instructions in the ISA into pseudo-instructions. In the example shown, Figure 3 The MRS, MOV, UBFX, and CMP instruction types in the _CMP_ instruction set are converted to _CMP_ instruction sets based on the control logic, timing, and memory access of the corresponding execution stages of the processor architecture being modeled. Figure 4 ALU pseudo-instructions in .

[0057] also, Figure 3 Software traces any non-memory access operation of the instruction, including any operands, in the Figure 4 The pseudo-instructions are stripped during the conversion. For example, Figure 3 The operands of the MRS, MOV, UBFX, and CMP instructions in Figure 4 The pseudo-instructions are removed, and the indicated operations of the mnemonics of the MRS, MOV, UBFX and CMP instructions are removed by converting them into general-purpose ALU pseudo-instructions.

[0058] The operands of instructions in the software trace can be used to determine dependencies. For example, Figure 3, the UBFX instruction on line 04 extracts data from the data stored in register w2, which was previously loaded by the LDR instruction on line 03. Therefore, in this example, the data to be extracted by the instruction on line 04 depends on the data loaded by the instruction on line 03. During preprocessing, the model object determines this dependency so that pseudo-instruction ID=104 (corresponding to the instruction on line 04) includes a dependency from pseudo-instruction ID=103 (corresponding to the instruction on line 03).

[0059] By stripping non-memory access operations and functionality from instructions when converting to pseudo-instructions, pseudo-instructions can be lightweight instructions, wherein execution simulation of pseudo-instructions can be faster. Simulation of execution of non-memory access functions can be avoided while maintaining appropriate control logic, timing, and memory access for performance analysis.

[0060] In some examples, memory read operations from the get model object and the load model object, and memory write operations from the store model object are fed to a bus driver that can be used as an application workload (e.g., traffic) for system-level performance analysis of the system interconnect, cache, and memory. Because the memory read and write operations use address and data size information from the corresponding instructions and are issued in an order and timing controlled by the software trace and hardware architecture, the memory read and write operations can fairly accurately represent the software workload. Figure 5 and Figure 6 The diagram shows architectural models 500 and 600, which are respectively Figure 1 and Figure 2 The architecture model 100, 200 is modified by virtual processing units (VPUs) 502, 504 communicatively coupled to the fetch model objects 102, 202 and load / store model objects 108, 208 as corresponding memory drivers. The VPUs 502, 504 are mapped to the memory drivers, which are connected to the interconnect and memory model to form the SoC system.

[0061] In addition, for SoC systems, the architecture model can be connected to a traffic driver to be simulated as an SoC platform. The traffic driver can be a model that can generate transactions that can be sent to access memory. Memory accesses from pipeline stages, such as fetch and load / store, can be converted into bus protocol transactions, such as Advanced eXtensible Interface (AXI) transactions. This transaction information is then sent to the traffic driver, which generates transaction-level modeling (TLM) transactions according to a given protocol (e.g., AXI). The traffic driver is connected to the SoC interconnect or bus, through which TLM transactions are sent to the cache or memory. Based on the delays encountered when accessing the cache or memory, delays and instruction stalls caused by fetch, load, and store can be accurately reflected in the simulation. Memory traffic from the architecture model can be used as a workload for SoC-level performance analysis and hardware redesign, such as cache size, interconnect topology, and memory configuration.

[0062] Implementing an architectural model as described above can allow for more efficient architectural exploration of processor architectures. The goal of architectural exploration can include determining how fast a given processor architecture can execute a given software program. This speed can be measured in instructions per cycle (IPC) or cycles per instruction (CPI). When architecturally exploring and comparing two or more processor architectures, an actual or representative software application (e.g., a benchmark) can be run on each of the processors and their IPC and / or CPI values ​​can be compared. The processor with the highest IPC or lowest CPI will be the fastest. The processors to be compared can belong to the same or similar architecture family, or can belong to very different architecture types.

[0063] When creating a new processor configuration, different hardware options may be explored, such as pipeline architecture, number and / or type of execution units, dispatch / retirement policy, memory / interconnect design, etc. In such cases, benchmark software may need to be executed or simulated on each of the processors. When the architecture is new and no compiler or bootloader is available for the architecture, or when the processor architectures are very different, in which case each processor may require a different compiler and bootloader, executing the benchmark software on each of the processors may be difficult or even impossible.

[0064] Using the architectural models described herein, performance analysis can be obtained without a compiler and without a bootloader for each different processor architecture or each variant of a processor architecture. Different pipeline architectures can be created by using the same pipeline stage model for different topologies and changing the execution model objects by adding more instances or changing the runtime parameters of the model objects. Because these models use pseudo-instructions instead of actual instructions based on the processor ISA, the same pseudo-instructions can be run as-is on architectural models of different architectures with different pipeline structures, different numbers of execution model objects, and / or different dispatch / retirement policies.

[0065] Figure 7 7 is a flow chart of a method 700 for analyzing processor architecture performance using an architectural model, according to some examples. At 702, an architectural model representing a processor architecture is obtained. The architectural model can be a control flow graph comprising interconnected instances of model objects representing the processor architecture. The architectural model can represent the control logic, timing, and memory access behavior of the processor architecture without implementing non-memory access functional behavior of the processor architecture.

[0066] At 704, pseudo-instructions generated from the software trace are obtained. The software trace can be generated from an execution or execution simulation of a software program on a processor having a processor architecture different from the processor architecture modeled by the architecture model. The software trace can be based on the ISA of the processor on which the execution or execution simulation is performed. The pseudo-instructions can be ISA agnostic. Each pseudo-instruction can include: (i) a mnemonic indicating which type of model object executes the corresponding instruction of the software trace, (ii) a unique ID, (iii) any (one or more) dependencies of other pseudo-instructions, and (iv) for any memory access pseudo-instructions, the memory address and access size of the memory access. In some examples, the pseudo-instruction may not include operands for non-memory access functional behavior of the corresponding instruction of the software trace.

[0067] At 706, the execution of the pseudo-instructions by the architecture model is simulated. In some examples, obtaining pseudo-instructions can be included in the simulation. In some examples, an acquisition model object of the architecture model obtains instructions based on the software trace. The acquisition model object passes the instructions to the decode and dispatch model object. The decode and dispatch model generates pseudo-instructions based on the instructions in the software trace. The decode and dispatch model then routes each pseudo-instruction to the appropriate instance of the execution model object based on the mnemonics of the corresponding pseudo-instruction and in an order based on applicable controls (e.g., any dependencies) and resource availability (e.g., execution model objects). In the instance of the execution model object, the execution latency is determined, including any latency for memory access. Once the execution latency is determined, the instance of the execution model passes the corresponding pseudo-instruction to the completion model object, which passes the completion of the pseudo-instruction to the decode and dispatch model object. The decode and dispatch model object retires the completed pseudo-instruction.

[0068] At 708, performance information is generated based on the simulation at 706. The performance information can be displayed in a graphical format (e.g., on a video display unit of a computer system). For example, the graphical format can include an axis listing instances of the model object and another vertical axis corresponding to execution time. The time and duration of each pseudo instruction at a given instance of the model object can be graphically indicated in the graphical format. Figure 8 Is from Figure 1 FIG800 shows an example graphical format 800 of performance information for a simulation of an architecture model 100. The vertical axis indicates instances of model objects, such as fetch, dispatch, compute (COMP), load / store (LOAD_STORE), and complete (complete), and the horizontal axis indicates time. The boxes in the graph indicate when the corresponding model object executed the corresponding pseudo instruction and the duration of the execution.

[0069] In some examples, based on the performance information obtained at 708, the processor architecture, and therefore the architecture model, can be redesigned or modified. Method 700 can be implemented again on the redesigned architecture. Thus, method 700 can be an iterative process whereby the architecture model is redesigned or modified until the architecture model produces satisfactory performance information.

[0070] Figure 9 is a flow chart of a method 900 for generating an architecture model according to some examples. The method 900 may be implemented as Figure 7 At 702, an architecture model is obtained. At 902, a class of a model object is defined using class parameters. The class of the model object with the class parameters represents the corresponding stage of the pipeline of the processor architecture to be modeled. The class parameters of a given model object class can be applied to each instance of the model object class. For example, Figure 2All instances of computation model objects 206 in the architecture model 200 may have the same or different latency, and thus, latency may be a class parameter of a computation model object class. Classes may be defined based on their corresponding functionality implemented by the architecture model as described above (e.g., including function calls that provide such functionality). The definition of a class of model objects may vary and may or may not include class parameters.

[0071] At 904, the control flow graph is populated with instances of the model object. Each phase of the processor architecture corresponds to an instance of the model object in the control flow graph. At 906, instance parameters of the instance of the model object are defined. The instance of the model object with the instance parameters represents the phase of the processor architecture being modeled. The instance parameters can allow different instances of the model object to have different characteristics. For example, if Figure 2 Instances of computation model objects 206 in architecture model 200 have different latencies, and the latency of each instance of computation model object 206 may be an instance parameter that can be defined separately for each instance of computation model object 206. At 908, the instances of the model objects are interconnected in a control flow graph to represent the processor architecture.

[0072] Figure 10 is a flow chart of a method 1000 for generating pseudo instructions from a software trace according to some examples. The method 1000 may be implemented as Figure 7 The pseudo instruction is obtained at 704. Figure 10 The method 1000 may be implemented by defining a class of a decoding model object and / or an instance of a decoding model object in an architecture model.

[0073] At 1002, a software trace generated for an ISA is received. The software trace includes instructions formatted according to the ISA. At 1004, each instruction in the software trace is converted into a corresponding pseudo-instruction, which is ISA-agnostic. For example, the mnemonic of the instruction can be used to identify the corresponding mnemonic of the pseudo-instruction. A lookup table or other functionality can be used to convert the instruction into a pseudo-instruction. At 1006, a unique ID is assigned and attached to each pseudo-instruction. At 1008, the corresponding memory address and access size are determined and attached to any memory access pseudo-instruction. At 1010, any dependencies indicated in the software trace are determined and attached to the corresponding pseudo-instruction.

[0074] According to one embodiment, a model object is created in a SystemC environment, and multiple architecture models are created in the SystemC environment using the model object. The architecture models simulate multiple RISC CPU architectures with different pipeline topologies and configurations. Each RISC CPU architecture includes a fetch model object (Fetch), a dispatch / decode model object (Dispatch), a compute model object (COMP), a load / store model object (LOAD_STORE), and a completion model object (Complete). The fetch model object reads a series of commands (e.g., instructions) from a JSON file, which are derived from software execution traces. The completion model object contains information about fetch, execution type, and command dependencies. The fetch model object uses a given program counter as an address to perform memory access for instruction fetch type commands and sends data access type commands to the dispatch / decode model object. The dispatch / decode model object converts the commands into pseudo-instructions and issues the pseudo-instructions to either the compute model object or the load / store model object, depending on the type of the pseudo-instruction. If the sequence on which the pseudo-instruction depends has not completed, the dispatch / decode model object stalls the issuance of the pseudo-instruction. The dispatch / decode model object can issue pseudo-instructions out of order if the pseudo-instructions have no dependencies or the pseudo-instructions they depend on have completed. The dispatch / decode model object maintains a list of issued and outstanding pseudo-instructions (in progress) and removes them in order upon receiving notification from the completion model object. The compute model object represents a compute unit and consumes a fixed latency. The load / store model object performs memory accesses for pseudo-instructions of the load / store model object type using the memory address and access size information in the corresponding pseudo-instruction. The load / store model object can issue multiple outstanding read / write pseudo-instructions. The completion model object collects pseudo-instructions when the pseudo-instructions are completed from the compute model object or the load / store model object and sends a notification to the dispatch / decode model object.

[0075] Figure 11 This is an architectural model of a simple RISC pipeline implemented in the SystemC environment. The architectural model consists of a fetch model object, a dispatch / decode model object, a compute model object, a load / store model object, and a completion model object interconnected in a control flow graph. Figure 12 This is the first superscalar RISC pipeline architecture model implemented in the SystemC environment. This architecture model consists of a fetch model object, a dispatch / decode model object, three parallel compute model objects, a load / store model object, and a completion model object interconnected in a control flow graph. These three compute model objects can execute up to three commands of the compute model object type in parallel. Figure 13This is an architectural model of a second superscalar RISC pipeline implemented in the SystemC environment. This architectural model includes a fetch model object, a dispatch / decode model object, three compute model objects, a load / store model object, and a completion model object interconnected in a control flow graph. The three compute model objects and one load / store model object can execute up to three compute model object-type commands and one load / store model object-type command in parallel.

[0076] A parsing and adaptation utility is created to read software instruction traces. Additionally, a parsing and adaptation utility is created to read software instruction traces from different types of simulators, such as loosely timed simulators and cycle accurate simulators. Figure 14 is an example software trace from a cycle-accurate model. Figure 15 is an example software trace from a loosely timed model.

[0077] The architectural model is integrated with the SoC hardware model in the SystemC environment. Figure 16 Pictured Figure 13 Integration of the architecture model and the SoC hardware model, wherein the acquisition model object 1600 and the load / store model object 1602 of the architecture model are mapped to the acquisition model object 1610 and the load / store model object 1612 of the SoC hardware model respectively.

[0078] The effects of different processor architectures, such as instruction dependencies and the number of execution model objects, are Figure 17A and Figure 17B The execution time and Figure 18 It is displayed on the instructions of each cycle in . Figure 17A and Figure 17B 1702 is shown with ID=23 for the first architecture model at line 17, and is shown with ID=26 for the first architecture model at line 20. Calculation model object #8 1704 is dependent on calculation model object #7 1702. The timing diagram shows how calculation model object #7 1702 is executed in the simulation before calculation model object #8 1704. In addition, Figure 17A and Figure 17BA load / store model object #7 1712 with ID=24 at line 18 for the second architectural model and a compute model object #8 1714 with ID=26 at line 20 for the second architectural model are shown. Compute model object #8 1714 is dependent on load / store model object #7 1712. The timing diagram shows how load / store model object #7 1712 is executed in simulation before compute model object #8 1714. The architectural model correctly represents command dependencies, and the simulation results show the impact of the dependencies as delays in the issuance of certain commands. Architectural choices were also analyzed, such as a single compute model object versus three compute model objects. For three compute model objects, it was observed that up to three compute model object type commands were executed in parallel, as shown in FIG. Figure 18 As shown, this improves execution performance because the parallel instructions 1802, 1804 take less time to execute than sequential execution 1812, 1814, respectively.

[0079] Additionally, modeling of processor architectures is presented. Figure 19 is a diagram of an example processor architecture. Figure 20 yes Figure 19 An architectural model of the processor architecture is shown. Software instruction traces obtained from benchmark software executed on a loosely timed model of the processor are converted into pseudo-instructions. These pseudo-instructions are simulated on the architectural model and a hardware model of the SoC that incorporates the architectural model. Processor and SoC performance are analyzed. Figure 21A 、 Figure 21B 、 Figure 21C and Figure 21D SoC performance analysis of the processor's performance information and software benchmark traces on an architectural model of the processor architecture is shown in a graphical format.For a number of execution model objects, multiple (eg, three) commands are observed to be issued per cycle.

[0080] The architectural model is modified using different processor configurations, such as the number of execution model objects and the number of dispatches per cycle, and the same software trace is used to simulate the different architectural models. Figure 22A 、 Figure 22B 、 Figure 22C and Figure 22D Performance information of the modified processor is shown in a graphical format along with SoC performance analysis of a software benchmark trace on an architectural model of the modified processor architecture.

[0081] Figure 23The diagram illustrates a set of example processes 2300 used during the design, verification, and manufacture of an article such as an integrated circuit for transforming and verifying design data and instructions representing the integrated circuit. Each of these processes can be structured and implemented as multiple modules or operations. The term "EDA" stands for the term "Electronic Design Automation." These processes begin with the creation of a product idea 2310 using information provided by a designer, which is transformed to create an article of manufacture 2312 using a set of EDA processes. When the design is finalized, the design is taped out 2334, which is when the artwork (e.g., geometric pattern) of the integrated circuit is sent to a manufacturing facility to produce a mask set, which is then used to manufacture the integrated circuit. After tapeout, semiconductor die 2336 are manufactured, and packaging and assembly processes 2338 are performed to produce a completed integrated circuit 2340.

[0082] The specification of a circuit or electronic structure can range from low-level transistor material layout to high-level description languages. Using a hardware description language (HDL), such as VHDL, Verilog, SystemVerilog, SystemC, MyHDL, or OpenVera, circuits and systems can be designed using a high-level representation. The HDL description can be transformed into a logic-level register transfer level (RTL) description, a gate-level description, a layout-level description, or a mask-level description. Each lower level of representation of more detailed description adds more useful detail to the design description, for example, including more detail about the modules being described. The lower-level representation of the more detailed description can be computer-generated, derived from a design library, or created by another design automation process. An example of a specification language for specifying a lower-level representation language of more detailed description is SPICE, which is used for detailed descriptions of circuits with many analog components. The description at each level of representation can be used by the corresponding system for that layer (e.g., a formal verification system). The design process can use Figure 23 The sequence depicted in . The described process is implemented by an EDA product (or EDA system).

[0083] During system design 2314, the functionality of the integrated circuit to be manufactured is specified. Simulation of the processor and / or SoC may occur during system design 2314, and system design 2314 may also include architectural exploration as previously described. The design may be optimized for desired characteristics such as power consumption, performance, area (physical and / or lines of code), and cost reduction. Partitioning the design into different types of modules or components may occur at this stage.

[0084] During logic design and functional verification 2316, modules or components in a circuit are specified in one or more description languages ​​and the functional accuracy of the specifications is checked. For example, components of a circuit can be verified to generate outputs that match the specifications of the designed circuit or system. Functional verification can use simulators and other programs such as testbench generators, static HDL checkers, and formal verifiers. In some embodiments, special component systems called "simulators" or "prototyping systems" are used to accelerate functional verification.

[0085] During synthesis and design for test 2318, the HDL code is converted into a netlist. In some embodiments, the netlist can be a graph structure, where the edges of the graph structure represent the components of the circuit, and where the nodes of the graph structure represent how the components are interconnected. Both the HDL code and the netlist are layered artifacts that EDA products can use to verify that the integrated circuit performs according to the specified design during manufacture. The netlist can be optimized for the target semiconductor manufacturing technology. Furthermore, the completed integrated circuit can be tested to verify that the integrated circuit meets the requirements of the specification.

[0086] During netlist verification 2320, the netlist is checked for compliance with timing constraints and for correspondence with the HDL code. During design planning 2322, the overall floor plan of the integrated circuit is constructed and analyzed for timing and top-level routing.

[0087] During layout or physical implementation 2324, physical layout (positioning of circuit components, such as transistors or capacitors) and wiring (connection of circuit components through multiple conductors) occur, and selection of cells from a library to implement a specific logic function may be performed. As used herein, the term "cell" may specify a group of transistors, other components, and interconnections that provide a Boolean logic function (e.g., AND, OR, NOT, XOR) or a storage function (such as a flip-flop or latch). As used herein, a circuit "block" may refer to two or more cells. Both cells and circuit blocks are referred to as modules or components and are implemented as physical structures and simulations. Parameters, such as size, are specified for the selected cell (based on standard cells) and accessed in a database for use by EDA products.

[0088] During analysis and extraction 2326, circuit functionality is verified at the layout level, which allows for refinement of the layout design. During physical verification 2328, the layout design is checked to ensure that manufacturing constraints, such as DRC constraints, electrical constraints, and lithography constraints, are correct and that the circuit device functionality matches the HDL design specifications. During resolution enhancement 2330, the layout geometry is transformed to improve the manufacturability of the circuit design.

[0089] During tape-out, data is created for producing photolithographic masks (after applying lithographic enhancements where appropriate).During mask data preparation 2332, the "tape-out" data is used to produce photolithographic masks, which are used to produce the finished integrated circuit.

[0090] The storage subsystem of a computer system (such as Figure 24 The computer system 2400 may be used to store programs and data structures used by some or all of the EDA products described herein, as well as products for developing library cells and physical and logical designs that use the library.

[0091] Figure 24 An example machine of a computer system 2400 is illustrated, in which a set of instructions may be executed for causing the machine to perform any one or more of the methodologies described herein. More particularly, the computer system 2400 may include stored instructions (e.g., stored on a non-transitory computer-readable medium) that, when executed by one or more processors of the computer system 2400, implements in whole or in part the methods described herein. Figure 7 、 Figure 9 and Figure 10 In various examples, the machine may be connected (e.g., using a network) to other machines in a LAN, an intranet, an extranet, and / or the Internet. The machine may operate in the capacity of a server or a client user machine in server-client network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client user machine in a cloud computing infrastructure or environment.

[0092] The machine may be a personal computer (PC), a tablet computer, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a network appliance, a server, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine. Furthermore, while a single machine is illustrated, the term "machine" should also be understood to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0093] The example computer system 2400 includes a processing device 2402, a main memory 2404 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM), a static memory 2406 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 2418, which communicate with each other via a bus 2430.

[0094] Processing device 2402 represents one or more processors, such as a microprocessor, a central processing unit, etc. More particularly, processing device 2402 can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. Processing device 2402 can also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. Processing device 2402 can be configured to execute instructions 2426 for performing the operations and steps described herein.

[0095] The computer system 2400 may also include a network interface device 2408 for communicating over a network 2420. The computer system 2400 may also include a video display unit 2410 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 2412 (e.g., a keyboard), a cursor control device 2414 (e.g., a mouse), a graphics processing unit 2422, a signal generating device 2416 (e.g., a speaker), a graphics processing unit 2422, a video processing unit 2428, and an audio processing unit 2432.

[0096] The data storage device 2418 may include a machine-readable storage medium 2424 (also referred to as a non-transitory computer-readable medium) having stored thereon one or more sets of instructions 2426 or software embodying any one or more of the methodologies or functionality described herein. During execution by the computer system 2400, the instructions 2426 may also reside, completely or at least partially, within the main memory 2404 and / or the processing device 2402, which also constitute machine-readable storage media.

[0097] In some implementations, the instructions 2426 include instructions that implement functionality corresponding to the present disclosure. Although the machine-readable storage medium 2424 is shown as a single medium in the example implementation, the term "machine-readable storage medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "machine-readable storage medium" should also be understood to include any medium that can store or encode a set of instructions executed by a machine and cause the machine and processing device 2402 to perform any one or more of the methods of the present disclosure. Therefore, the term "machine-readable storage medium" should include, but is not limited to, solid-state memory, optical media, and magnetic media.

[0098] Some portions of the foregoing detailed description have been presented in terms of algorithms within computer memory. These algorithmic descriptions are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm can be a sequence of operations leading to a desired result. These operations require physical manipulations of physical quantities. These quantities can take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. Such signals may be referred to as bits, values, elements, symbols, characters, terms, numbers, and the like.

[0099] It should be remembered, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, as will be apparent from this disclosure, it should be understood that throughout this specification certain terms refer to the actions and processes of a computer system or similar electronic computing device that processes and transforms data represented as physical (electronic) quantities in the computer system's registers and memories into other data similarly represented as physical quantities in the computer system's memories or registers or other such information storage devices.

[0100] The present disclosure also relates to an apparatus for performing the operations described herein. The apparatus may be specially constructed for the intended purpose, or it may comprise a computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including a floppy disk, an optical disk, a compact disk read-only memory (CD-ROM) and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, or any type of medium suitable for storing electronic instructions, each of which is coupled to a computer system bus.

[0101] The algorithms and displays described herein are not inherently related to any particular computer or other device. Various other systems may be used in conjunction with the program according to the teachings herein, or it may prove convenient to construct a more specialized device to perform the method. Furthermore, the present disclosure is not described with reference to any particular programming language. It should be understood that a variety of programming languages ​​may be used to implement the teachings of the present disclosure described herein.

[0102] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, which instructions may be used to program a computer system (or other electronic device) to perform the processes of the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium, an optical storage medium, a flash memory device, etc.

[0103] In the foregoing disclosure, the implementation of the present disclosure has been described with reference to specific example implementations thereof. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the implementation of the present disclosure as set forth in the following claims. Where the present disclosure refers to some elements in the singular, more than one element may be depicted in the accompanying drawings, and identical elements are marked with the same numerals. Therefore, the present disclosure and the accompanying drawings should be regarded as illustrative rather than restrictive.

Claims

1. A method comprising: populating a control flow graph with instances of a model object, each instance of the model object representing a respective stage of a first processor architectural design; interconnecting the instances of the model objects in the control flow graph, the interconnected instances of the model objects being an architectural model representing the first processor architectural design; as well as The architectural model including the interconnected instances of the model objects is output by one or more processors.

2. The method according to claim 1, further comprising: simulating execution of pseudo-instructions by the architecture model, the pseudo-instructions being generated based on the software trace; as well as Performance information is generated based on the simulation of the execution of the pseudo instruction by the architecture model.

3. The method according to claim 2, further comprising: receiving said software trace including instructions; as well as Each of the instructions of the software trace is converted into a corresponding pseudo instruction of the pseudo instructions. 4 . The method of claim 3 , wherein converting each instruction of the instructions of the software trace into the corresponding pseudo instruction of the pseudo instruction comprises excluding operands of the corresponding instruction in the corresponding pseudo instruction.

5. The method of claim 2, wherein the software trace is from execution or a simulation of execution of instructions by a processor having a second processor architecture different from the first processor architectural design.

6. The method of claim 2, wherein the pseudo-instruction is agnostic with respect to an instruction set architecture (ISA) of the first processor architectural design.

7. The method of claim 1, wherein each instance of the instance of the model object represents one or more of: one or more of control logic, timing, and memory access of the corresponding stage of the first processor architecture design. 8 . The method of claim 1 , wherein each instance of the instance of the model object represents a corresponding phase excluding non-memory access functional behavior of the corresponding phase.

9. The method according to claim 1, further comprising: connecting a first traffic driver to an acquire model object instance, said instance of said model object including said acquire model object instance, said first traffic driver being a model configured to generate a first transaction for accessing memory external to said architecture model; as well as A second traffic driver is connected to a load-store model object instance, the instance of the model object including the load-store model object instance, the second traffic driver being a model configured to generate a second transaction for accessing memory external to the architecture model.

10. A non-transitory computer-readable medium comprising stored instructions that, when executed by a processor, cause the processor to: populating a control flow graph with instances of a model object, each instance of the model object representing a respective stage of a first processor architectural design; interconnecting the instances of the model objects in the control flow graph, the interconnected instances of the model objects being an architectural model representing the first processor architectural design; and The architectural model including the interconnected instances of the model objects is output.

11. The non-transitory computer-readable medium of claim 10, wherein the instructions, when executed by a processor, cause the processor to: simulating execution of pseudo-instructions by the architecture model, the pseudo-instructions being generated based on the software trace; and Performance information is generated based on the simulation of the execution of the pseudo instruction by the architecture model.

12. The non-transitory computer-readable medium of claim 10, wherein each instance of the instance of the model object represents one or more of: control logic, timing, and memory access of the corresponding stage of the first processor architecture design. 13 . The non-transitory computer-readable medium of claim 10 , wherein each instance of the instance of the model object represents the corresponding phase excluding non-memory access functional behavior of the corresponding phase.

14. The non-transitory computer-readable medium of claim 10, wherein the instructions, when executed by a processor, cause the processor to connect a flow driver to an acquire model object instance, the instances of the model objects comprising the acquire model object instance, the flow driver being a model configured to generate transactions for accessing memory external to the architecture model.

15. A method comprising: generating, using one or more processors, an architectural model representing a first processor architecture, the architectural model comprising interconnected instances of model objects in a control flow graph, each instance of the instances representing a respective stage of the first processor architecture; receiving a software trace including instructions; Converting the instruction into a pseudo-instruction, the pseudo-instruction being instruction set architecture (ISA) agnostic; simulating execution of the pseudo-instruction by the architecture model; and Performance information of the first processor architecture is generated based on the simulation.

16. The method of claim 15, wherein each instance of the instances represents control logic, timing, memory access, or a combination thereof of the corresponding stage of the first processor architecture and does not implement non-memory access functional behavior of the corresponding stage of the first processor architecture.

17. The method of claim 15, wherein the software trace is from execution or a simulation of execution of the instructions by a processor having a second processor architecture different from the first processor architecture.

18. The method of claim 15, wherein for each of the instructions, converting the instruction into the pseudo instruction comprises: determining a pseudo-mnemonic for the pseudo-instruction based on the mnemonic for the corresponding instruction; attaching a unique identifier to the directive; attaching a memory access operand to said pseudo-instruction when said mnemonic of said corresponding instruction indicates that said corresponding instruction is a memory access instruction; determining whether the corresponding instruction is dependent on execution of a previous instruction; as well as When the corresponding instruction depends on the execution of a previous instruction, a dependency is attached to the pseudo instruction.

19. The method of claim 15, wherein each of the pseudo-instructions does not include a non-memory access operand.

20. The method of claim 15, wherein the instances of the model objects included in the architectural model comprise: obtaining an instance of a model object configured to obtain said instructions of said software trace; an instance of a decoding model object configured to convert the instruction into the pseudo instruction; dispatching an instance of a model object configured to route the directive; one or more instances of one or more execution model objects configured to simulate execution of corresponding pseudo instructions received from said instance of said dispatch model object; as well as An instance of a completion model object is configured to indicate when simulation of execution of a corresponding pseudo instruction is complete, the instance of the dispatch model object being further configured to retire pseudo instructions based on completion indications received from the instance of the completion model object.