Performance benchmarking of real-time software and hardware

By extracting the characteristics of the object code and combining with the test hardware processor, a performance benchmark specific to the real-time system is generated, which solves the problem of difficult to estimate or compare the performance of real-time system processing in the prior art, and achieves more accurate performance evaluation and hardware selection.

CN113778819BActive Publication Date: 2025-05-23GENERAL ELECTRIC CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110638417.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-10
Filing Date
2021-06-08
Publication Date
2025-05-23
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively estimate or compare the processing performance of specific real-time systems and applications through generalized benchmarks, especially in specific software applications running in special environments.

Method used

By extracting the features of the object code, the performance benchmark is generated and combined with specific test hardware processors to generate performance benchmarks related to the performance of real-time and time-critical computing systems.

Benefits of technology

Customized performance benchmarks for specific hardware systems are implemented, enabling more accurate estimates and comparisons of real-time performance of specific software, helping design engineers choose the best target hardware processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113778819B_ABST
    Figure CN113778819B_ABST
Patent Text Reader

Abstract

A system and method determines a unique performance benchmark for a particular computer object code for a particular microprocessor. By generating multiple unique benchmarks for a single identical code module on multiple different processors, the method determines which processor is optimal for the code module. By generating a performance benchmark for each of multiple code modules for a single specified processor, where the multiple modules have the same / similar functionality but the detailed code or algorithms differ, the system and method identifies the code changes that are optimal for the single specified processor. The system and method may require first extracting selected features of the object code (as actually executed) into a code profile, and then generating a performance benchmark based on the code profile and in machine-level timing data for the selected microprocessor. In this way, code security can be achieved by firewalling the object code from the second stage of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to design phase optimization of computing systems. More specifically, the present system and method relate to generating performance benchmarks for customized computer software running on a specific hardware system. The present system and method also relate to generating one or more performance benchmarks related to real-time, time-critical computing system performance. Background Art

[0002] Real-time or mission-critical programming

[0003] The field of real-time computing (RTC), sometimes also called "deterministic computing," concerns hardware and software systems that are subject to one or more "real-time constraints" (e.g., from event to system response). Real-time hardware and software programs must typically guarantee responses to system events within specified time constraints, which may be referred to as "deadlines."

[0004] Systems used in many mission-critical applications must be real-time, such as those used to control fly-by-wire aircraft or anti-lock brakes, both of which require immediate and accurate mechanical and electrical responses. More generally, in aviation systems or other types of transportation systems, engine and control systems may be required to respond to critical environmental events within a specified time to maintain safe operation of the aircraft or other vehicle. Other systems that do not require mechanical components may also be mission-critical. For example, in order to maintain uninterrupted and / or high-quality communication of mission-critical data, a communication system may require mission-critical packet routing, switching, data compression / decompression, data encryption / decryption, etc.

[0005] In a typical system, time-critical responses reflect automated operations (i.e., no real-time intervention by a human operator), and the deadlines for the system's response to an event can be on the order of milliseconds or microseconds. Although typical or expected response times can be given, systems that are not specified as real-time operations cannot generally guarantee responses within any time frame. Real-time processing fails if it is not completed within a specified deadline relative to an event; deadlines must always be met, regardless of system load, for optimal or even safe system performance.

[0006] Real-time systems can also be characterized as receiving environmental or system data (typically from system or environmental sensors), processing the sensed data, and returning results that are sufficiently responsive to affect system operation and / or the environment at substantially the same time as the data is received (i.e., without significant delay).

[0007] Software for real-time applications often needs to be carefully coded and fine-tuned to achieve optimal performance on a specific, specified hardware microprocessor. Real-time software applications can include either or both application-specific software and a real-time operating system.

[0008] Software Benchmarks

[0009] A software benchmark is a numerical value that indicates the performance speed of software, determined by testing software (code under development or commercially released software). When a specific program or software module has a higher numerical score than other software modules with similar functions, the high numerical rating generally indicates faster performance speed. In the alternative, the benchmark can also be configured so that a lower number indicates a reduction in the execution time of the task, therefore indicating better performance.

[0010] There are many benchmarks that measure the performance of hardware rather than software, usually the performance of hardware microprocessors. Some well-known benchmarks include Dhrystone, Whetstone, and some benchmarks developed by the Embedded Microprocessor Benchmark Consortium. However, in general, these are generalized benchmarks designed primarily to determine the performance of underlying hardware such as a microprocessor (possibly along with associated hardware such as a data bus). Their main purpose is to characterize the relative general performance of different microprocessors, regardless of the specific application software running on a given microprocessor.

[0011] Existing generalized benchmarks suffer from similar deficiencies, both with respect to real-time programming and with respect to programming of less mission-critical applications (e.g., general-purpose business software). Existing benchmarks measure what they were designed to measure on the selected hardware processor (e.g., integer performance or floating-point performance, or some combination of the two). For specific software applications running in special environments, especially for real-time systems, it is difficult to estimate or compare the processing performance of specific systems and applications via these generalized benchmarks. This is because specific real-time environments and real-time software applications have special and unique requirements for integer commands, floating-point commands, and other low-level memory commands.

[0012] Then, there is a need for systems and methods to profile existing applications (usually alpha or beta applications in development for real-time systems) when the applications are run in representative hardware environments. There is also a need for customized benchmarks that will better estimate the real-time performance of specific software in development for a specific target hardware system. Summary of the invention

[0013] In at least one aspect, embodiments of the present systems and methods determine a unique performance benchmark for a particular computer object code when the code is actually executed on a particular designated hardware microprocessor.

[0014] Benchmarking a Single Code Module on Multiple Processors: In an embodiment, the system and method can generate multiple unique benchmarks for a single identical object code module for each of multiple processors so that the system and method can determine which hardware processor (among multiple potential processors or "target platforms") is optimal for the code module.

[0015] Functionally Similar Software Modules: In another embodiment, the system and method can generate multiple performance benchmarks for a single specified processor, with each code module in the multiple target software modules having a unique benchmark, with the goal of selecting only one target software module for actual use and deployment. Each of the multiple software modules can have the same or similar functionality, but differ in the detailed code or algorithms. In this way, the system and method can identify those code changes that are optimal for a specific specified target platform (e.g., a specific specified hardware processor).

[0016] Functionally Distinct Software Modules: In another embodiment of the present system and method, multiple distinct target software modules have distinct functionality, and the multiple modules are intended to combine (i.e., integrate) performance, possibly operating in parallel or sequentially, or both. Each target module may then have a very different set of code. The present system and method may generate a first set of performance benchmarks for each target module when running on a first processor or first target hardware platform, then generate a second set of benchmarks for each target module when running on a second (different) target processor or different target hardware platform, and so on. In this way, guided by the performance benchmark results for multiple distinct cooperating target modules, the design engineer can select a target hardware processor that provides the desired balance of adequate execution time for most or all of the target software modules (thereby providing optimal performance for the integrated software system of all target modules).

[0017] Method: In some embodiments, the system and method may require a first phase for extracting selected features of the object code into a code profile and a second phase for generating a performance benchmark based on the code profile and based on the selection of a specific designated test hardware processor. Thus, code security can be achieved by firewalling the object code from the second phase of the method. In an embodiment, the code profile reflects the distribution of microcode commands that are executed by the object code when running on a selected target processor under specific conditions.

[0018] In embodiments, the system and method may also require a second stage of generating a processor profile for a particular processor, and a second stage of jointly analyzing the code profile and the processor profile to generate a performance benchmark. In such embodiments, a processor-specific performance benchmark for a particular object module may be generated without (i) actually running the object code on the particular processor (instead running the code on a different processor), or in some embodiments, (ii) without running the object code on any processor.

[0019] In some embodiments, the system and method involve not only profiling the object code, but further involve primarily identifying and isolating for analysis those portions of the object code that are most typically or frequently called when the target software module is actually running in real time on the processor.

[0020] In embodiments, the systems and methods may also be used to generate a performance benchmark for a specified computer program or code module when the code module is executed on a specific virtual processor or through an interpreter (eg, Java).

[0021] Embodiments of the present system and method may be particularly suitable for benchmarking object code that is intended for real-time, performance-critical applications. However, the present system and method may also provide advantages for software performance benchmarking of non-mission-critical systems.

[0022] In some embodiments, the present systems and methods may entail using a local or distributed computer configuration system, wherein the configuration system performs a method for benchmarking computer code, the method comprising:

[0023] (i) deconstructing first object code of a first software module to create a first configuration file of machine language instructions of the first object code for the first software module, which may indicate machine language instructions to be executed when the first object code module is executed on a given target processor or target hardware system;

[0024] (ii) benchmarking a first target hardware system including the first set of hardware to determine real-time execution timing data machine language instructions of the first target hardware system;

[0025] (iii) combining a first configuration file of the first software module with timing data of the first target hardware system to generate a first performance benchmark, wherein the first performance benchmark reflects the first object code of the first software module and the architecture of the first target hardware system.

[0026] In some embodiments, the deconstruction phase (phase (i)) may entail running the software module on a processor having an instruction set shared by other processors (e.g., different processors in a manufacturer's family of related processors, or different processors having a common hardware architecture, such as the well-known Intel instruction set). In alternative embodiments, the deconstruction phase (phase (i)) may entail analyzing the object code to determine the distribution of machine-level commands without actually running the object code on any processor.

[0027] In some embodiments of the present systems and methods, the stages or steps of generating a benchmark for a particular code module and a particular hardware processor may be performed by a single computer, and the stages may be performed substantially consecutively in time (i.e., each successive analysis stage is substantially initiated after the immediately preceding analysis stage is completed). In alternative embodiments, the stages of analysis may be performed by separate computers of a distributed system, and may be performed with considerable time gaps (e.g., minutes, hours, days, or longer periods) between the analysis stages. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Advantageous designs of embodiments of the invention are derived from the independent and dependent claims, the description and the drawings. In the following, preferred examples of embodiments of the invention are explained in detail with the aid of the drawings:

[0029] Figure 1 is an integrated system element / method diagram showing some exemplary system elements and some exemplary method steps for generating a performance benchmark for a target module intended to be executed on a target hardware platform.

[0030] Figure 2A A block diagram is presented of an exemplary real-time control system (RTS) for which the present systems and methods may be used to aid in the design / development process.

[0031] Figure 2B A simulation process that may be employed during the design / analysis phase of a real-time control system is shown.

[0032] Figure 3 An exemplary instruction mix decomposition is shown that may be generated by runtime analysis of target software modules when executed on an exemplary target hardware platform or platform emulation.

[0033] Figure 4 An exemplary process for determining the length of time required to execute a particular microcode instruction on a target hardware platform is shown.

[0034] Figure 5 Several alternative exemplary performance benchmark calculations in accordance with the present systems and methods are shown.

[0035] Figure 6 Exemplary comparative performance benchmark values ​​for different exemplary combinations of target hardware systems and target software modules are shown.

[0036] Figure 7 A block diagram of an exemplary computer for benchmarking software and hardware in accordance with the present systems and methods is shown. DETAILED DESCRIPTION

[0037] The following detailed description is merely exemplary in nature and is not intended to limit the systems and methods disclosed herein, the elements or steps of the systems and methods, the applications of the systems and methods, and the uses of the systems and methods. Furthermore, it is not intended to restrict or limit the scope to, or be bound or limited by, any theory presented in the preceding background or summary or the following detailed description.

[0038] Throughout the application, the description of various embodiments may use "comprising" language, indicating that the system and method may include certain elements or steps described; however, the system and method may also include other elements or steps not described, or may be described in conjunction with other embodiments, or may be only shown in the drawings, or those elements or steps that are known in the art to be necessary for the function of the processing system. However, those skilled in the art will understand that in some specific cases, the language "essentially consisting of" or "consisting of" may be used alternatively to describe the embodiments.

[0039] I. Terminology

[0040] In this document you will understand:

[0041] Object code is computer code containing machine-level commands (or, in some embodiments, bytecodes), also referred to as "microcode," "machine language instructions," "assembly language instructions," or simply "instructions" or "commands" in some cases for brevity. Object code containing machine-level commands is generated from source code, which may be written in any of a variety of general-purpose or specialized high-level programming languages. Such languages ​​include, by way of example, but not limited to: C, C++, C#, FORTRAN, Java, LISP, Perl, Python, Ruby, and synchronous programming languages ​​such as Argos, Atom, LabVIEW, Lustre, PLEXIL, and the like. Object code is typically generated from high-level language code by an automated process using an application called a "compiler," in a process called "compilation."

[0042] Machine-level code, microcode or assembly code commands can also be understood as the native machine instructions for a given hardware microprocessor, which are determined by the fixed hardware design of the processor. In many cases, there are "families" of different but similar microprocessors that share a common set of native machine instructions (or at least a common core set of local machine instructions, where advanced processors in a family may have additional native machine instructions that simpler siblings lack).

[0043] Target application, target module, target software module, target object code, target code, target program: The object code to be analyzed and benchmarked by the present systems and methods may be referred to equivalently in this document by any of the foregoing terms. Those skilled in the art will understand that some of the foregoing terms may have similar, slightly different, but overlapping meanings in general usage in the art and in various contexts. For example, a single application may include object code contained in one or more software modules or software libraries. For the purposes of the appended claims, the term "object code module" is used as an umbrella term for target executable program, target application, target software module, target library, etc. It will be understood that real-time system development often requires the generation and evaluation of multiple target modules or target applications.

[0044] Target processor, target microprocessor, target real-time processor, target hardware processor, target microcontroller, target controller, selected processor, selected microprocessor: In this document, a hardware processor designed to process and execute object code in a module may be referred to by any of the above terms. The target processor may generally be a hardware microprocessor with one or more cores, or may be a more specialized controller capable of interpreting and executing target object code. In some embodiments of the present systems and methods, the target processor may be a single large-scale integrated microchip. In some embodiments, the target processor may be an LSI or VLSI chip on-board circuit used for memory access, data bus access, cache memory or other processing purposes.

[0045] For purposes of the appended claims, the term "target hardware system" refers to and encompasses one or both of "hardware processor" or "target platform" discussed immediately below.

[0046] The target processor may be a dedicated processor for real-time applications; to name just one example (among many), the NXP XT is manufactured by NXP and used in avionics, industrial control, gateways, smart homes, and many other applications. P2020 or P2010 communications processor. In some embodiments, the target processor may alternatively be a general purpose processor, such as any of the numerous Pentium processors manufactured by Intel Corporation, or any of the general purpose processors manufactured by AMD, NXP, Qualcomm, Apple, TI, Renesas Electronics, Xilinx, etc.

[0047] As discussed further below, the present system method, in part and among other elements, benchmarks one or more target modules to operate relative to one or more target processors based in part on one or both of (i) actual test execution of the target modules on the target processor, or (ii) expected and desired execution of the target modules on the target processor.

[0048] Target Platform, Target Hardware, Target Hardware System: In this document, the terms "target platform", "target hardware platform", "target system", and "target hardware system" are used synonymously to refer to a processing system or computing system that is intended to be evaluated for benchmarking, where the processing / computing system typically includes a specific microprocessor and other supporting chips (e.g., memory chips, bus controllers, cache memory, I / O port chips, and other processing-related chips known in the art). The target platform may also be considered to include or contain a suitable operating system, which may be loaded or obtained from software or firmware.

[0049] Benchmark: A general term used to rate the performance of a hardware processor and / or data processing system; and more specifically, in this document, to generate for a target module (a target module may be a software program, an object code module component of a larger software program, a library or other object code compilation of a software program) a timing / speed performance rating that reflects the typical execution dynamics of the target module when running on a specified target processor or target computing system.

[0050] II. Overview of Performance Benchmarking Strategies and Tools

[0051] Benchmark suites (benchmarking applications)

[0052] In an embodiment, the present systems and methods may employ a tool suite 105 (software running on a suitable hardware processor or processors) to create a custom performance benchmark 150 that is tailored to a target application.

[0053] Benchmark Software Tools: In some embodiments, the present systems and methods may employ a software tool suite (software or firmware) running on a suitable computer or group of computers ("benchmark computer system" 110), which may or may not be networked together in real time. That is, different software applications 105 of the suite may be run at separate times on separate benchmark computers 110, with only the various output files needing to be shared (via network 190 or via other forms of file transfer) between the different programs in the software suite.

[0054] In an alternative embodiment, all required software tools 105.n may be integrated into a single benchmark application 105 running on a single computer 110. The exemplary software that may be used to implement the present systems and methods may be referred to herein as a "benchmark suite" 105 for convenience.

[0055] Benchmarking strategy

[0056] Figure 1 An integrated system / data flow / method diagram 100 is presented, illustrating some exemplary system elements 105 , 110 and exemplary method steps 170 of the exemplary system and method 100 to generate a performance benchmark 150 of a target software module 115 . Figure 1 Also shown are exemplary input sources 115, 120, 130, 135 and exemplary outputs 115, 120, 135, 150, 155. (Note that IMB 120 and PTD 135 defined below are both outputs and inputs.) The exemplary software elements of the application tool 105 and one or more computers 110 may also be referred to herein as a "benchmark suite" 102. In some cases, the term "benchmark suite" 102 may refer primarily or exclusively to the application tool 110.

[0057] In one or more embodiments of the present systems and methods, a benchmark strategy using a benchmark suite 102 may include software tools 105.n for:

[0058] Real-time instruction execution hybrid data collection: Instruction execution trace data is collected by running the target executable object code on real or simulated target hardware. As discussed further below, this can be done by the target module execution analyzer (TMEA) tool 105.1.

[0059] Object code decomposition of target module: The target object code is profiled to determine a representative low-level (i.e., machine language or bytecode) instruction mix decomposition (IMB) 120 of the target executable object code. In an embodiment, the instruction mix decomposition 120 is based on those instructions that are typically used when the target object code is actually executed in a simulation or operating environment. The IMB 120 can, for example, identify the percentage of use of different object code instructions during execution time, or can provide the raw number of times the object code instructions are executed. As discussed further below, this task can also be performed by a target module execution analyzer (TMEA) tool 105.1.

[0060] Timing data for target hardware (processor or hardware platform): Timing data is collected from the proposed target hardware, where the timing data indicates the time for each instruction to run for multiple different instruction sets for low-level instructions. The time can be clock cycles or analog mixed signal (AMS) time scales, such as nanoseconds. As discussed further below, this can be done through the hardware platform profiler (HPPR) tool 105.2.

[0061] It will be noted that each low-level machine instruction will have its own unique timing data; and further, such timing data will generally vary from one target microprocessor to another. As just one example, an instruction to retrieve a two-byte floating point number may require four cycles on a first hardware processor A, while requiring six hardware cycles on a second hardware processor B.

[0062] Note that in some embodiments, timing data may be obtained only for a target microprocessor or microcontroller. In alternative embodiments, timing data may be obtained for a wider range of hardware platforms that may include microprocessors / microcontrollers, data buses, attached memory caches, I / O ports, specified bus speeds or volatile memory capacity / speeds, specific system firmware, and possibly a selected operating system or operating system version / configuration.

[0063] (iv) Benchmark synthesis: synthesize intermediate values ​​from the timing data of the selected hardware processor / system and from the instruction mix profile of the target object code, which intermediate values ​​represent the performance of the instruction mix of the target object code on the selected hardware processor. As discussed further below, this can be done by the benchmark generator tool (BG) 105.3.

[0064] (v) Benchmark comparison and analysis: For a given real-time hardware / software system under development (e.g., for an aircraft control system with real-time operating constraints), the obtained performance benchmark scores are compared to the required scores in the hardware specifications for the real-time operating system. In this way, it can be determined whether a particular target module or target application can be executed on a test or potential hardware platform to have a sufficiently fast response time for the control task at hand. As discussed further below, this can also be accomplished by the Benchmark Generator Tool (BG) 105.3.

[0065] Example characteristics of a performance benchmark created using the benchmark suite

[0066] In an exemplary embodiment of the present system and method, some features of the object code performance benchmark 150 created using the benchmark test suite may include, for example but not limited to:

[0067] (A) Environment-related benchmarking: determining a performance benchmark 150 that takes into account software and / or hardware characteristics such as cache utilization and cryptography (i.e., encryption of code in an object module); and may also take into account physical environmental factors such as system operating temperature and / or CPU operating temperature;

[0068] (B) Software-specific benchmarks: Accurate benchmark measurements of proprietary software with fuzzy workflows;

[0069] (C) Generic Timing Profile of Target Hardware: Once the timing data of the potential target processor / platform is characterized, the throughput utilization of different software builds of the target module can be analyzed based on the generic timing profile of the potential target hardware;

[0070] (D) Realistic Profiling of Real-Time Execution: In embodiments of the present systems and methods, the benchmark suite profiles the target object module's real-time instruction mix, and the timing profile favors (i.e., provides a heavier weight to) those machine-level instructions that are most commonly used in actual operations, while reducing the weight of those machine language instructions that are not commonly used in actual execution.

[0071] For example, a target object code module may include: one or more input-output (I / O) operations that are used relatively infrequently (e.g., acquiring sensor data ten times per second), where these I / O operations are encapsulated in code segment 'A'; a variety of data analysis operations performed on current and historical sensor data, where the analysis operations are encapsulated in code segment 'B'; and system control (I / O) operations encapsulated in code segment 'C'. Assuming further that in typical real-time operation (as determined by the present system and method), code segment 'A' is called 10% of the time, code segment 'B' is called 88% of the time, and code segment 'C' is called only 2% of the time, then the object code profile will provide relative usage frequencies (for machine language instructions in code segments A, B, and C) of ratios of 0.10 / 0.88 / 0.02, respectively.

[0072] If a particular machine language instruction 'T' is identified k times in code segment A, m times in code segment B, and n times in code segment C, then when profiling the target module, the present system and method will weight or count instruction 'T' with an appropriate weight, such as (0.10)*k+(0.88)*m+(0.02)*n. Similar considerations will apply to other machine language instructions U, V, W, X, etc.

[0073] Note that the system for weighting and comparing the instructions identified in the previous paragraphs is exemplary only, and other methods or formulas for weighting, profiling, or ranking absolute machine instruction set usage and relative usage of machine instructions (all consistent with usage identified by the benchmark suite during execution of object code) may be employed consistent with the present systems and methods. For example, in some embodiments, performance benchmarks 150s may be generated to distinguish between data / instructions fetched from RAM and faster data / instructions fetched from cache memory.

[0074] The benchmark suite provides a highly automated throughput assessment tool and can be integrated into any continuous improvement / continuous development (CI / CD) workflow.

[0075] III. Example Elements and Data Flow of Benchmark Code

[0076] continue Figure 1 , the target software module 115 may be a software program, an object module component of a larger software program, a library or other object code compilation of a software program intended to be executed on the target hardware platform 130. The performance benchmark 150 values ​​reflect the expected execution dynamics of the target module 115 by the target platform 130 or a component of the target hardware platform, such as the target microprocessor 130.1 (see FIG. 2 below).

[0077] exist Figure 1 , solid black arrows indicate various sources / sinks and directions of exemplary data flows for some embodiments of the present systems and methods. Straight dashed black arrows illustrate optional data flows that may be employed in some embodiments of the present systems and methods. Curved dashed arrows indicate exemplary flows of exemplary methods 170 according to the present systems and methods.

[0078] The exemplary system 105 may include one or more computers 110 , such as desktop computers, servers, laptop or tablet computers, embedded test platforms or systems, or the like, which may be configured to run the benchmark tool 105 of the present systems and methods. Figure 1 Three exemplary computers (110.1, 110.2, and 110.3) are shown, but fewer or more computers 110 may be employed. Figure 1In the exemplary embodiment of the present invention, the computers 110 are not networked together as part of, for example, a local area network (LAN) 190 or a corporate network 190. In alternative embodiments, the computers 110 may be connected via a local area network (LAN) 190 or via a direct cable or wireless connection 190. In some embodiments, two or more computers 110 may be data coupled via a wide area network (WAN) 190 such as the Internet 190. In general, it will be understood that data may be transferred between the computers 110, possibly in real time or possibly with a time delay, via a network connection 190 or a tangible medium (e.g., a flash drive).

[0079] Running or configured to run on one or more computers 110 are exemplary software tools of benchmark suite 105. Figure 1 In the exemplary embodiment shown, the software tools 105 include a target module execution analyzer (TMEA) tool 105.1, a hardware platform profiler (HPPR) tool 105.2, and a benchmark generator (BG) tool 105.3. Those skilled in the art will recognize that in some embodiments, Figure 1 The exemplary tools shown may be combined into one or two software tools; and in alternative embodiments, one or more of TMEA 105.1, HPPR 105.2 and / or BG 105.3 may be implemented as two or more separate software tools, respectively, with appropriate data communications between them.

[0080] In an exemplary embodiment of the present system and method, the method 170 may entail a first step 170.1 of generating an instruction mix decomposition (IMB) 120 for a target module (i.e., executable object code to be analyzed) 115.1 + runtime environment 115.2. (Hereinafter, the combination of the target module 115.1 and possibly the runtime environment 115.2 may be referred to simply as the target module 115). The target module 115 may be provided to the TMEA 105 as input. This detail will be discussed further below (see FIG. 2).

[0081] Target module: Target module 115.1 is previously generated by a compiler or interpreter (not shown) based on computer source code (written in a known high-level programming language) typically generated by a human programmer or programming team (possibly in conjunction with automatic code generation tools and possibly integrated with third-party object code libraries). Target modules are typically intended to run on known or expected hardware platforms with certain input data or certain kinds of input data.

[0082] Runtime Environment: A known or expected hardware platform (or simulation software and parameters that characterize the platform) and sample input data for the target module together constitute the runtime environment 115.2 input to the TMEA 105. In an embodiment, the runtime environment 115.2 may include a representation or actual use of the target operating system for the environment.

[0083] Instruction Mix Breakdown: In an embodiment of the present system and method, the resulting output from TMEA 105 is an instruction mix breakdown (IMB) 120 (also referred to as an "object code breakdown"). Figure 3 The IMB 120, shown in pie chart form in FIG. 1 , provides (i) an absolute representation (i.e., a numerical count) of how many times each of the various machine-level (assembly language) instructions was actually executed by the target module during actual or simulated runtime (also referred to herein as a "machine code occurrence" or MC_A); and / or (ii) a relative representation (relative percentage) of the various machine-level (assembly language) instructions actually executed by the target module during actual or simulated runtime. Details of this will be discussed further below (see Figure 3 ).

[0084] Hardware Platform and Timing Data: In an exemplary embodiment of the present system and method, the method 170 may entail a second step 170.2 of generating platform timing data (PTD) 135 for the hardware platform 130. The hardware platform profiler (HPPR) tool 105.2 profiles the target hardware platform 130 to determine the intrinsic or inherent timing data of the microcode on the hardware platform 130.

[0085] As described above: During product development, engineers designing real-time systems, mission-critical systems, or other computing systems may choose to evaluate one or more target hardware platforms for use as system controllers. The target hardware platform 130 may be, for example, but not limited to: a microprocessor, a microcontroller, a controller system, a microprocessor or microcontroller with some additional microchips (e.g., memory, memory access chip, bus controller, I / O ports, etc.) mounted on a circuit board or backplane. In some embodiments of the present system and method, different or competing target hardware platforms 130 may be understood as the same hardware, but each platform runs different firmware. In some embodiments, different or competing target hardware platforms 130 may be understood as the same hardware, but each platform runs a different operating system or a different version of an operating system.

[0086] In an embodiment of the present system and method, the HPPR tool 105.2 may be executed on one of the benchmark suite computers 110.2, and the HPPR 105.2 may obtain appropriate operating data from the target hardware platform 130 via an I / O port and cable provided for data transfer (or via wireless means). In an alternative embodiment, the HPPR tool 105.2 may be executed directly on a processor of the target hardware platform 130 itself, so as to obtain real-time performance data from the target hardware platform 130. In an alternative embodiment, the HPPR tool 105.2 may be executed by a software / firmware process that runs simultaneously or sequentially on a processor of the target hardware platform 130 and on the benchmark suite computer 110.2.

[0087] The HPPR tool 105.2 obtains timing data for machine-level (also known as assembly language) instructions running on the target hardware platform 130. As is known in the art, machine-level instructions are typically those instructions that are elements of the instruction set of a hardware microprocessor or hardware microcontroller of the hardware platform 130. However, machine-level instructions may also include instructions from instruction sets of other chips in the hardware platform 130 (e.g., instruction sets of memory access chips, bus controllers, I / O hardware, digital signal processors, graphics processing chips, and other hardware-level instruction sets associated with the hardware platform 130). In some embodiments of the present systems and methods, machine-level instructions may be interpreted as including bytecode instructions to be interpreted by a byte-level interpreter (e.g., as may be used with Java and similar programming languages).

[0088] As is known in the art, machine-level instructions may include, for example, but are not limited to: instructions to retrieve data from memory for temporary storage in processor registers; instructions to store data from processor registers to memory; instructions to perform various arithmetic and / or logical operations on data currently in processor registers; instructions to trace and jump to various memory addresses; instructions to send data to or receive data from various system buses and / or I / O ports; and many other low-level processor instructions.

[0089] Processor Timing Data: Timing data for any given machine-level instruction may include, for example, but is not limited to: (a) the number of clock cycles required to execute the machine-level instruction; (b) the absolute time (in nanoseconds or other suitable timing units) to execute the machine-level instruction; (c) the system clock speed and / or system bus speed at which the timing data is obtained; (e) the number of data bytes operated on by the machine-level instruction; and (f) other numerical data related to the actual amount of time and / or relative time required to execute the machine-level instruction.

[0090] In some embodiments of the present systems and methods, HPPR 105.2 may be configured to identify whether a machine-level instruction has different timing data depending on, for example, the byte size of the data on which the instruction operates. In some embodiments, HPPR 105.2 may identify one or more of an upper timing limit for a machine-level instruction operation; a lower timing limit for a machine-level instruction operation; an average (mean or median) time for a machine-level instruction operation; and / or other statistical distribution data for machine-level instruction operation times.

[0091] In some embodiments of the present systems and methods, HPPR 105.2 may be configured to take into account that a single machine-level instruction may require different runtimes depending on variations in input parameters.

[0092] Platform timing data: In some applications, the actual timing data for a given microcode instruction may vary depending on whether the instruction timing is determined solely by processor performance or alternatively by full hardware platform performance (memory speed, processor speed, cache bus speed, ...). For example, the use of a memory cache or multiple levels of memory caches may affect the execution time of the entire hardware platform (depending in part on the size and speed of the cache memory).

[0093] The hardware platform profiler 105.2 may deliver as output platform timing data 135 for the entire target hardware platform. The platform timing data 135 may be enumerated in a variety of data formats including ASCII listings (directly human readable on standard displays and printouts), or in various hexadecimal or binary encodings.

[0094] The output platform timing data 135 will generally include the same types of data as described above for the hardware processor. That is, for each machine-level instruction to be characterized, the platform timing data may include: timing data 135, which may include, for example, but not limited to: (a) clock cycles; (b) the absolute time to execute the machine-level instruction (also referred to herein as "machine code timing" or MC_T); (c) system clock speed / bus speed; (e) the number of data bytes operated; and (f) other relevant digital data that characterizes the timing of the machine-level instruction.

[0095] Selective (target module specific) analysis of machine-level instruction timing: In some embodiments of the present systems and methods, the HPPR 110.2 may limit its timing analysis to only those machine-level instructions that were actually employed in a selected target object code module 115 during module execution. In such embodiments, the HPPR 110.2 may receive as input the IMB 120 of the target module 115. The IMB 120 identifies which instructions were actually executed by the target module 115, and thus the HPPR 110.2 may limit the timing analysis to only the listed machine-level instructions.

[0096] In an alternative embodiment of the present system and method, HPPR 110.2 profiles the timing of all machine-level instructions for the target hardware platform 130. This provides a generalized timing table for the target hardware platform 130, which can then be used for further analysis of many possible target modules 115.

[0097] In alternative embodiments of the present systems and methods, the platform timing data 135 may be provided by the hardware platform manufacturer as part of a data sheet or other supporting documentation for the hardware platform 130 .

[0098] Benchmark Generation: In an exemplary embodiment of the present system and method, the method 170 may entail a third step 170 . 3 of generating a performance benchmark 150 indicative of an expected performance level of the target module 115 if the target module 115 is executed on the target hardware platform 130 .

[0099] Benchmark generator (BG) tool 105.3 receives instruction mix breakup (IMB) 120 and receives platform timing data (PTD) 135. BG 105.3 then uses IMB 120 and PTD 135 to compute a performance benchmark 150 that is specific to a combination of at least (i) a target module and (ii) a target hardware platform.

[0100] In some embodiments of the present systems and methods, a given generated performance benchmark 150 may be further specific to a particular input test vector 115.2" for a particular simulation run of the real-time system 200. In some embodiments of the present systems and methods, and for a particular given target module and target hardware platform, multiple passes may be made to generate the performance benchmark 150 (e.g., with varying input test vectors), and a cumulative or average benchmark may be generated for a combination (i.e., pairing) of target modules and hardware platforms.

[0101] As an exemplary statement only and not intended to be limiting, the BG tool 105.3 may generate a performance benchmark 150 or a starting value for a benchmark according to an exemplary benchmark formula (subject to further refinement):

[0102] (B1) Performance benchmark [software module, hardware platform] (or PB[s,h]) =

[0103]

[0104] in:

[0105] CS = the number of different machine-level commands in the target module command set (CS), i.e., the number of different machine-level commands that can potentially be executed by the target module 115,

[0106] MC_Timing k = the timing (absolute or clock cycles) of the execution of the kth machine-level command (obtained from the platform timing data 135), and

[0107] MC_Appearances k =The number of times the kth machine-level instruction occurs during execution when the hardware platform 130 is executed on the target module 115 (obtained from the instruction mix decomposition 120).

[0108] Elsewhere below, MC_Timing k Sometimes abbreviated as MC_T, while MC_Appearances k Sometimes abbreviated as MC_A.

[0109] Throughout this application, especially in conjunction with the following Figure 5 , further discussing the generation of performance benchmarks 150. The steps of the exemplary method 170 are summarized as follows:

[0110] Step 170.1: Get the instruction mix decomposition (IMB) 120 of the target module (TM) 115.1 and the runtime environment (RTE) 115.2;

[0111] Step 170.2: Obtain platform timing data (PTD) 135 of the hardware platform (PT) 130;

[0112] Step 170.3: Calculate a single performance benchmark 150 based on the IMB and the PTD.

[0113] Optionally, method 170 may entail generating additional performance benchmarks 150 for different modules 115 running on different hardware platforms 130 , and further generating comparative performance benchmarks 155 .

[0114] In some embodiments of the present systems and methods, method step 170.2 may be performed before (or simultaneously with) method step 170.1.

[0115] Steps 170.1 through 170.3 are exemplary only, and alternative or additional steps may be performed consistent with the appended claims of the present systems and methods.

[0116] IV. Exemplary Real-Time System

[0117] Figure 2A is a block diagram of an exemplary real-time / mission-critical system (RTS) 200, which may be in the planning phase, design phase, or development phase (or upgrade or redesign phase), for which the present systems and methods may be used to aid in the design and development process.

[0118] Real-time system 200 is exemplary only and not limiting; RTS 200 is only one of many possible such systems that may use the present systems and methods, and RTS 200 has been selected in this document for illustration only. There is no implication that the present systems and methods are limited to application in engine design, or to the aviation field, or any other similar limitation, and such should not be inferred. Rather, the present systems and methods are applicable to the design and development of real-time systems and mission-critical systems across a wide range of technical fields.

[0119] Physical Real-Time Systems vs. Simulated Real-Time Systems: Those skilled in the art will appreciate that the development of a real-time system (RTS) 200 is typically intended to be implemented in a specific physical structure or system that may include a number of tangible hardware, electrical components, structural components, data processing components, and other parts 130, 210, 215, 217, 265, as well as the RTS software 115. During system development, one or more tangible physical prototypes may be developed. In addition, during system development, the entire physical real-time system 200 may be designed first and then simulated in whole or in part via one or more simulation programs 130", 210", 265" that may be run, for example, on one or more computers 110. Details are discussed further below.

[0120] In one exemplary application of the present system and method 100 for technology design, the RTS 200 may be an aircraft jet engine 210. The block diagram of the aircraft jet engine 210 omits many necessary elements of an actual jet engine design, and thus includes only a few selected elements for purposes of explanation and illustration.

[0121] The jet engine 210 may include an intake fan 215 or compressor blades 215 or the like, the fan rotation speed of which is controlled by a motor 217. The motor 217, and therefore the rotation speed of the fan 215, is in turn controlled by a hardware fan controller 130, which in this example constitutes a target hardware platform 130. In development, a number of different hardware fan controllers 130 may be considered for the final aircraft engine 210 design (i.e., a "target" for the purposes of this disclosure). The hardware fan controller 130 may include a microprocessor 130.1, a motor regulator 130.2, and a memory 130.3, as well as other digital and analog components not shown.

[0122] Memory 130.3 may be used to store fan controller application software 115.1, which is an exemplary target module 115 for benchmarking / evaluation according to the present systems and methods. The target fan controller application 115.1 or the target software module 115.1 of the entire application may regulate or help regulate the speed of the motor 217 and therefore the fan 215.

[0123] It should be noted that during engine development, typically only one target module 115.1 is employed or evaluated on a particular processor / computer 110.1 at a given time. In an alternative embodiment of the present system and method, the TMEA 105.1 may be configured to evaluate multiple TM+RTEs 115 simultaneously, with each TM+RTE 115 running in a separate evaluation process. Throughout the extended process of engine development, multiple similar but differently coded target modules 115.1 (e.g., 115.1.1, 115.1.2, ... 115.1.n) may be designed and tested for possible use with the hardware controller 130.

[0124] In order to perform performance benchmarking at a separate time (or to perform parallel benchmarking on separate processors / computers 110.1 or separate processes of TMEA105.1), each such different target module 115.n can be loaded into memory 130.3 for conventional testing and evaluation; in particular, each such different target module can be benchmarked by the TMEA 105.1 of the present systems and methods.

[0125] In this exemplary application, once the target module 115.1 is loaded into the memory 130.3, it may include object code that directs the hardware processor 130.1 to: (i) receive various local environmental data from the sensor 265; (ii) analyze and / or process the received sensor data in real time to determine or fine-tune the desired speed of the fan 215; and (iii) adjust the motor 217 in real time to ensure that the fan 215 rotates at the desired speed.

[0126] It will be apparent to those skilled in the art that, given the options of a number of different possible fan controllers 130 and the options of a number of different versions of the target module 115.1, the design engineer will seek to identify the best match / combination of a selected target module 115.1 and a selected fan controller hardware 130. The benchmarking system and method disclosed herein is configured to help identify one or more optimal hardware controllers 130, one or more optimal target modules 115.1, and the best combination of hardware controller 130 (possibly including the best choice of processor 130.1) and target module 115.1 code.

[0127] V. Exemplary Simulation Processing and Target Module Benchmarking

[0128] Figure 2B is a block diagram showing exemplary elements and data flow of a simulation process 290 that may be incorporated into a real-time system (including, but not limited to, Figure 1 and Figure 2A The simulation process 290 may be employed during the design / analysis phase of the exemplary real-time system discussed above. The simulation process 290 may be run, for example, on a computer 110 (such as the computer 110.1 of the benchmark test suite 102). In an embodiment of the present system and method, the simulation process 290 may be used to implement the method step 170.1 of the exemplary method 170 above to generate the instruction mix breakdown (IMB) 120.

[0129] The simulation 290 may include one or more software / firmware mechanical models 210'', such as, but not limited to, a simulated engine 210'' that simulates the structure and function of an aircraft engine. The simulated engine 210'' may also include one or more simulated sensors 265'' (which may simulate sensors 265 of an aircraft jet engine 210), in addition to other elements, functions, software objects and modules. Generally, hardware design simulations are known in the art, and further details are not provided herein.

[0130] In embodiments of the present systems and methods, simulation 290 may also provide a software-only simulated target hardware platform 130", which may simulate, for example, the target hardware platform controller 130 of exemplary system 200. In embodiments of the present systems and methods, the target hardware platform controller 130 may be integrated into the simulation process in the form of an actual controller 130, with appropriate communication links and data transfer means between the actual physical target hardware controller 130 and computer 110.1 ( Figure 2A not shown).

[0131] Combining Simulated and Physical Systems: In alternative embodiments of the present systems and methods, the target hardware platform controller 130 may be implemented as a combination of some actual hardware (e.g., the actual target microprocessor 130.1) and software objects that emulate other elements of the target hardware controller 130 (e.g., a software emulation of the motor regulator 130.2). In embodiments of the present systems and methods, implementations of the simulated hardware controller 130" may include a simulated operating system ( Figure 2B In alternative embodiments of the present systems and methods, simulation 290 may combine a real physical system 210 with a fully or partially simulated target hardware platform 130; or a fully physical target hardware platform 130 with a simulated mechanical model 210".

[0132] More generally, the target hardware platform controller 130 may be implemented as a combination of some actual hardware (e.g., an actual target microprocessor 130.1) and software objects that emulate other elements of the target hardware controller 130 (e.g., a software emulation of the motor regulator 130.2). In embodiments of the present systems and methods, implementations of the emulated hardware controller 130" may include an emulated operating system ( Figure 2B not shown).

[0133] In one embodiment of the present system and method, Figure 2B As shown, the RTS simulation process 290 may have integrated the target module execution analyzer (TMEA) 105.1 of the benchmark test suite 102, or execute the target module execution analyzer (TMEA) 105.1 of the benchmark test suite 102 as needed. In an alternative embodiment of the present system and method (not shown in FIG. 2), the TMEA 105.1 may be run as a main computer application or tool, and the TMEA 105.1 may be linked to the hardware simulation 290, integrate the hardware simulation 290 into itself, or call the hardware simulation 290 as a separate process. (Thus, as shown in FIG. Figure 2A The specific arrangement of software modules and functionality shown is to be understood as exemplary only and not limiting.)

[0134] In some embodiments of the present systems and methods, the RTS simulation process 290 is executed primarily or exclusively to perform the functions of the TMEA 105.1. In alternative embodiments, the RTS simulation 290 may also include or perform simulation tasks in addition to those of the TMEA 105.1 (e.g., simulating and determining other aspects of engine performance or efficiency).

[0135] During execution, the RTS simulation process 290 will accept as input, or integrate as stored data structures, a variety of input test vectors 115.2" representing the simulated environment 115.2. In the exemplary case of the current aircraft engine simulation 290, the input test vectors 115.2" may include, for example, but not limited to: (i) parameters describing / characterizing demands or control signals for the aircraft engine, which are related to or derived from desired aircraft speed, desired aircraft altitude, desired or necessary engine performance if a second engine fails, and other factors determined via aircraft pilot selection or physical environmental (e.g., external wind speed) requirements; and (ii) parameters describing / characterizing a hypothetical environment directly local to or internal to the aircraft engine, such as appropriate pressure data, engine internal air velocity data, temperature data, etc.

[0136] In operation, the RTS simulation 290 loads the target module object code 115 into the real and / or simulated target hardware platform 130 / 130" for execution by the target hardware platform; initiates simulated operation of the simulated aircraft jet engine (210") when controlled by the real / simulated target hardware platform 130 / 130"; and accepts appropriate input test vectors 115.2".

[0137] During a simulation run, a target module execution analyzer (TMEA) 105.1 monitors the actual / simulated target hardware platform (130 / 130") to determine which machine-level (assembly) instructions are executed by the target hardware platform; and to determine how many times any particular machine-level instruction is executed. The TMEA 105.1 provides as output an instruction-mixed data breakdown (IMB) 265 for the simulation run.

[0138] Those skilled in the relevant art will recognize that identifying which machine instructions are being executed in real time by a hardware processor or have been executed by a hardware processor 130.1 is known in the art. For example, for some hardware microprocessors 130.1, the currently running instructions may be read from a hardware level instruction register. The details may vary with different hardware platforms 130 and / or different hardware processors 130.1. The details are not discussed in this document.

[0139] In summary, in embodiments of the present systems and methods, an exemplary method 170.1 for obtaining an IMB 120 requires (see also below) Figure 3 ):

[0140] (i) executing the target module object code (115.1) and appropriate input test vectors 115.2 on a simulated target hardware platform 130" (to simulate the environment 115.2);

[0141] (ii) identifying each machine-level command 320.n executed by the simulated target hardware platform 130" during the simulation run; and

[0142] (iii) Maintaining a list 120 having an execution count 330 of how many times each machine-level command 320 was executed by the simulated hardware platform 130″.

[0143] Alternative Methods of Generating IMB 120: In alternative embodiments of the present systems and methods, an exemplary alternative method 170.1' (not shown) for obtaining IMB 120 may entail:

[0144] (i) Introduce one or more flags or markers in the source code (e.g., C++, Java, etc.) of the target module 115.1 that reflect the frequency estimates (determined by the software programmer or system engineer) at which some portion of the code is expected to execute. For example, each instruction loop may include a specific programmer-supplied flag that estimates the number of times the programmer estimates that loop may run. Those skilled in the art will understand that such flags are of course only estimates, and are necessarily independent of any input test vectors 115.2 that may simulate a real-world environment.

[0145] (ii) Generate a target object module 115.1 via a suitable compiler (not shown), the target object module 115.1 including an object module flag indicating a source code flag. For example, the object code implementing a loop may be preceded by an object module flag indicating the estimated number of times the loop is expected to be run.

[0146] (iii) Parsing the object code 115.1 (eg, by the TMEA 105.1) to identify the machine-level instructions 320 in the code and target module flags that estimate how many times a section of code may be executed.

[0147] (iv) Generate an IMB 120 based on the estimated execution values ​​reflected in the machine-level instructions and flags in the object module.

[0148] Those skilled in the art will recognize that this alternative method 170.1' may not be as accurate as a method that requires running the object module 115.1 in a simulation environment. For example, the alternative method 170.1' may be used during the design phase when a simulation mechanical model 210" and / or a simulation target hardware platform controller 130" is not available.

[0149] VI. Exemplary Instruction Mixing Decomposition

[0150] Figure 3An exemplary instruction mix breakdown (IMB) 120 is shown that may be generated by a target module execution analyzer (TMEA) tool 105 . 1 based on an exemplary runtime analysis of a target module 115 . 1 on an exemplary target hardware platform 130 of an exemplary simulation environment 115 . 2 .

[0151] During a simulation run of the target module 115.1, the exemplary IMB 120 identifies an instruction list 320 of machine-level instructions (also referred to as opcodes or assembly language instructions) that execute on the target hardware platform 130 given an input test vector 115.2 of the simulation environment 115.2. Figure 3 The illustrated list 320 is exemplary only, and depending on the details of the processor 130 . 1 and the target hardware platform 130 , the instruction list 320 may contain entirely different and / or additional or fewer opcodes.

[0152] For each respective instruction 320.1, 320.2, ... 320.n detected during execution, the IMB 120 may identify an absolute number of times the instruction was executed (ANTIE) 330, and / or a relative number of times the instruction was executed (RNTIE) 335 compared to a total count of machine-level instructions executed 340. In embodiments of the present systems and methods, the IMB 120 may also include a visual breakdown of relative or absolute instruction usage, such as an exemplary pie chart 325. Such a pie chart 325 or other visualization of instruction breakdown may assist computer programs in analyzing the performance of object code and improving source code for better performance.

[0153] As discussed elsewhere in this document, and in accordance with the present systems and methods, IMB 120 may be combined or integrated with platform timing data 135 via appropriate calculations to obtain a single performance benchmark 150 for a particular combination / pairing of target hardware platform 130 and target module 115 . 1 .

[0154] In one embodiment of the present system and method, a single simulation run may be used to generate the IMB 120 for a particular pairing of a selected target hardware platform 130 and a selected target module 115.1. However, it may be the case that even for a particular pairing of a selected target hardware platform 130 and a selected target module 115.1, the numerical decomposition of instructions 330, 335 may vary depending on the particular input test vectors 115.2" of the simulation run. Therefore, in some embodiments of the present system and method, the TMEA 105.1 may perform multiple simulation runs on the same pairing of a hardware platform 130 and a target module 115.1. The results of the multiple runs may be averaged or otherwise statistically analyzed to generate a more accurate numerical profile 330, 335 of instruction set usage.

[0155] Supplemental Hardware Platform Instruction Set: In some implementations, the hardware platform 130 may have other supplemental or additional low-level instruction sets 320 associated with hardware other than the primary microprocessor 130.1. For example, other microchips such as floating point units (FPUs), cryptographic processing units, digital signal processing (DSP) chips, or sensor control chips may have their own microcode commands. These additional microcode commands may be sent to the applicable microchip (DSP, sensor control, etc.) via various hardware means, including as parameters to a port call made via the target processor 130.1 or via direct memory access (DMA) between the memory 130.3 and the applicable microchip. In embodiments of the present systems and methods, the TMEA 105.1 may be configured to recognize such supplemental microcode commands in the target module 115.1 and determine the number of times such supplemental microcode instructions are executed during runtime.

[0156] Privacy of Target Module Operation: It will be noted that in some embodiments of the present systems and methods, the instruction mix decomposition IMB 120 contains data such as which specific machine-level instructions 320 have been executed by the real / simulated target hardware platform controller 130, and the frequency 330 at which each such instruction was executed (or in some embodiments, the relative percentage of times each instruction was called compared to other assembly instructions 335). However, the IMB (120) need not, and in some embodiments of the present systems and methods, does not, contain any indication of the order or sequence in which the instructions were called. As a result, in various embodiments, while the IMB 120 indicates which opcodes 320 of the controller instruction set were called by the operating target module 115 and at what frequency; the output of the IMB 120 effectively obscures the detailed operation and operating principles of the target module 115. In this way, the IMB 120 reveals data related to the operational efficiency of program execution while hiding or "firewalling" the details of code execution.

[0157] VII. Exemplary Platform Timing Method

[0158] In an embodiment of the present system and method, a hardware platform profiler (HPPR) tool 110.2 is used to analyze a target microprocessor 130.1 or a target hardware platform 130 (which typically includes a microprocessor 130.1) to obtain timing data 135 for some or all microcode instructions of the platform / microprocessor 130.1. In the following discussion, the term "target hardware platform 130" (or simply "target platform 130") is used to interchangeably refer to the entire platform 130 or one of its selected processing elements (e.g., the target microprocessor 130.1).

[0159] In one embodiment, HPPR tool 110.1 is communicatively coupled to target hardware platform 130 to control target hardware 130 via software to obtain timing data. In an alternative embodiment, HPPR tool 110.1 may be executed directly by microprocessor 130.1 of target hardware platform 130 to obtain platform timing data 135.

[0160] Figure 4 A flow chart of an exemplary method 170 . 2 of obtaining microcode timing for a target hardware platform 130 is shown.

[0161] The method 170.2 begins at step 405. In step 405, the method builds or obtains a list of M microcode commands 320 for which timing data is to be obtained. The list can be created based on a variety of different sources or methods, including, for example, but not limited to: (i) obtaining a complete list of microcode commands 320 for a target platform 130, such as from a manufacturer's specification sheet; (ii) creating (based on selections made by an engineer or programmer) a list with a subset of the complete platform microcode commands 320; (iii) using a target module specific list of microcode commands 320 for a specific hardware platform 130 based on the output of TMEA 105.1.

[0162] The method continues to step 410. In step 410, a specific microcode command m is selected from the microcode command list. i .

[0163] The method continues to step 420. In step 420, the method creates a software timing function or software timing routine to continuously execute each microcode instruction a number of 'N' times in a loop (an exemplary value of N may be in the range of 2 to 100, or even a value greater than 100). The timing function / routine may initially be generated in a high-level language or as a high-level macro command, but will be compiled (if necessary) into direct machine / microcode commands in order to be executed. In an alternative embodiment, a value of N=1 means that the selected microcode instruction is only executed once.

[0164] In an embodiment of the present system and method, generating timing microcode requires encoding the loop instruction execution function (LIEF) of the selected microcode command; and also using a compiler flag that directs the compiler to unroll the loop so that the selected microcode command is actually called multiple times in succession in the timing microcode. This can avoid / prevent measuring the time of jumps / branches in the loop.

[0165] For microcode commands that require one or more parameters, and in some embodiments of the present systems and methods, substep 420.1 may generate, encode, or select a single set of one or more appropriate values ​​for the parameters. In alternative embodiments of the present systems and methods, substep 420.2 may generate, encode, or select two or more sets of parameters, each set having different parameter values ​​or different parameter values ​​than the other sets. In this way, the microcode command may be run multiple times with different parameters to account for the fact that different parameter values ​​and data sizes may cause different execution times for the microcode command.

[0166] In an embodiment, the high-level code of LIEF (e.g., code written in C++) is coded to be general enough so that it can be compiled for and run on hardware processors with completely different instruction sets (e.g., Power PC, ARM, X86, and other processor families).

[0167] In pseudo-code form, the LIEF is exemplary and not limiting in any way, and is used for an exemplary "add and store" microcode command, which may contain the microcode:

[0168] (B2)timing_loop_for_'add-store'command{

[0169] addstore(A1,B1)

[0170] addstore(A1,B2)

[0171] addstore(A2,B3)

[0172] addstore(A2,B4)}

[0173] The exemplary microcode (B2) executes four (N=4) addstore() commands, where, for example, addstore(A, B) may retrieve a first value from memory address A, retrieve a second memory value from address B, add A+B, and store the resulting value back to memory address A. In an embodiment, each time the addstore(A, B) command is executed, different parameter values ​​A1, A2, B1, B2, B3, B4 may be used. In an alternative embodiment, each time the exemplary addstore(A, B) command is executed, the same parameter values ​​A, B may be used.

[0174] Exemplary method 170.2 continues to step 425. In step 425, a start time of LIEF execution is obtained.

[0175] In step 430, the LIEF is executed on the target platform 130, thereby (through the inherent encoding of the LIEF) causing N consecutive executions of the selected microcode instructions.

[0176] In step 435, the end time of the LIEF execution is obtained.

[0177] In step 440, the total execution time of LIEF is obtained as: total execution time (TET) = end time - start time.

[0178] In step 445 of exemplary method 170.2, for the selected microcode command, the microcode execution time (MC_T) is obtained as MC_T=total execution time / N.

[0179] The method then returns to step 410 where a different specific microcode command is selected from the microcode command list. The method is repeated until MC_T data has been generated for all commands in the list.

[0180] Table 1 (T1) immediately below gives exemplary output platform timing data (PTD) 135 for method 170.2 applied to an exemplary microprocessor 130.1. Output PTD 135 is exemplary only and includes only a small subset of commands for a typical microprocessor; is not limiting; and does not necessarily represent actual microcode commands or command execution times for any known microprocessor:

[0181] Table 1 - Platform Timing Data (PTD) (Example Microcode Timing)

[0182] Microcode Commands Microcode execution time (MC_T) (ns) C_Add 5 C_AddStore 10 C_Branch 3 C_Compare 7 C_IntMult 12 C_LoadRegister 5 C_Store 5 C_Subt 5

[0183] In an alternative embodiment of the present system and method, for some hardware microprocessors that have dedicated hardware registers to reflect execution time, the microcode execution time may be retrieved directly from the microprocessor.

[0184] Supplemental Hardware Platform Instruction Sets: As described above, the hardware platform 130 may have other, supplemental or additional low-level instruction sets 320, which may relate to, for example, but not limited to, a floating point unit (FPU), a cryptographic processing unit, a digital signal processing (DSP) chip, or a sensor control chip microcode set. In embodiments of the present systems and methods, the HPPR 105.2 may be configured to determine the amount of time required for the hardware platform 130 to execute such supplemental microcode instructions.

[0185] Comparing Target Hardware Processors: During the development of the present systems and methods, testing has demonstrated that the execution speeds of various comparable microcode instructions can vary significantly when a selected target module 115.1 is executed on different underlying target hardware platforms 130. Variations in the target microprocessor 130.1 design (e.g., the presence or absence of a floating point unit, the size of the data cache, and the speed of the RAM) can and do result in significantly different timing scores (e.g., for commands of the PPC instruction set versus the ARM instruction set). Functionally identical or similar commands show different times for the same instruction given on each platform. The present systems and methods are designed to detect these differences and identify their impact on the real-time performance of a given target module (which may have the same or nearly identical high-level language coding on different hardware platforms).

[0186] It will be noted that once a given target processor / platform 130 . 1 / 130 has been benchmarked to obtain MC_T, the same set of timing data MC_T (for the given target processor / platform 130 . 1 / 130 ) can be used to benchmark many different target modules / environments 115 of that platform 130 .

[0187] In an alternative embodiment of the present system and method, instead of employing the hardware platform profiler 105 . 2 , the platform timing data 135 may be obtained from the manufacturer's specifications for the target processor 130 . 1 and / or the target platform 130 .

[0188] VIII. Exemplary Benchmark Generation

[0189] As described above, the benchmark generator (BG) tool 105.3 is configured to accept as input the instruction mix decomposition 120 of the target module 115.1 and the platform timing data 135 of a given hardware platform 130; and provide as output a suitable benchmark score 150 indicative of the expected real-time performance of the target module 115.1 on the target hardware platform 130. The benchmark score 150 indicates the expected real-time performance, which reflects the expectation that the benchmark score generated by the present system and method can approximately predict the relative performance of different combinations of target module 115.1TM and hardware platform 130HP.

[0190] An exemplary formula (PB1) for generating a performance benchmark 150 (also referred to as a "benchmark calculation") is discussed above and is referred to herein for convenience. Figure 5 This exemplary formula (PB1) is reproduced in . Figure 5 Two other exemplary formulas (PB2) and (PB3) for generating performance benchmark 150 are also shown in FIG.

[0191] The exemplary formula B1 combines the timing MC_Timing of the processor microcode commands and the absolute number of times MC_Appearances 330 that each microcode command is executed to reach the performance benchmark value PB[s,h]. Using the exemplary performance benchmark formula (PB1) of PB[s,h], it will be apparent to those skilled in the art that:

[0192] (i) if a first target module TM.1 is executed on a first hardware platform 130HP.1 and then executed on a second hardware platform 130HP.2; and further (ii) if the execution time of the most frequently used machine-level commands on platform HP.2 is longer than that on platform HP.1, then (iii) the performance benchmark value PB2[s,h] of platform HP.2 will generally have a higher value than the performance benchmark value PB1[s,h] of platform HP.1. Thus, for the exemplary benchmark calculation (PB1), a slower target module / hardware platform combination will generally produce a higher performance benchmark value PB.

[0193] In some embodiments of the present system and method, it may be desirable to treat PB[s,h] generated by PB1 as an intermediate benchmark value; and then generate a final benchmark 150PB_final, such as a reciprocal value, so that PB_final[s,h] = 1 / PB[s,h]. In such an embodiment, if platform HP.2 executes slower than HP.1, then the performance benchmark value PB of platform HP.1 is higher than the performance benchmark value of platform HP.2. (That is, the lower performance benchmark value will reflect a lower ratio of performance).

[0194] Similarly, taking the exemplary formula PB1, and in an exemplary application of the benchmark suite:

[0195] (i) two target modules 115 TM.1 and TM.2 may be generated, both of which will run on the same hardware platform 130 HP;

[0196] (ii) each of modules TM.1 and TM.2 may be coded to have highly similar or identical functionality, but coded differently or optimized differently (perhaps using any or all of different source-level languages, different compilers, different code optimizations, or different code organization and design);

[0197] (iii) If module TM.1 uses the fast-executing machine-level commands of platform HP more frequently than module TM.2, and / or TM.1 uses the slower-executing machine-level commands of platform HP less frequently than module TM.2, then the above exemplary formula PB1 for PB[s,h] will generally indicate or reflect that module TM.1 will have a lower performance baseline value PB[s,h] compared to the target module TM.2.

[0198] Here again, and in some embodiments of the present systems and methods, it may be desirable to generate a parent / final performance benchmark such that if module X performs faster than module Y, then module X will have a higher performance benchmark value than module Y.

[0199] Figure 5 Another exemplary benchmark formula or calculation PB2 that may be employed with the present systems and methods is also defined.Performance benchmark calculation PB2 is similar to PB1, but instead employs the relative number of times RMC_A 335 that a particular microcode command is executed by the target hardware platform while running the target module 115.

[0200] Figure 5 Another exemplary performance benchmark formula or calculation PB3 that may be employed with the present systems and methods is also defined. Benchmark calculation PB3 is similar to B1, but further includes one or more digital weighting factors 510 associated with some or all microcode commands. In the overall calculation of benchmark PB3, weighting factors that may be established or determined by a system design engineer (for example) may favor certain microcode commands and their timing or execution frequency over others. For example, if the design engineer believes that certain types of microcode commands (e.g., memory access commands, cached commands or cache control commands, branch commands, or certain arithmetic commands) should be prioritized over others in the benchmark calculation, then these factors may be selected.

[0201] Other Benchmark Embodiments: It will be understood that the benchmark calculations PB1, PB2, and PB3 are exemplary only, and that other benchmark formulas, calculations, or algorithms may be employed consistent with the scope of the appended claims. For example, the benchmark calculations may employ various adjustments designed to scale and / or cluster benchmark scores; thus, for example, but not limited to, a first pairing of a slower hardware platform and a better-coded (faster) object module produces approximately the same benchmark value as a second pairing of a faster hardware platform and a less efficiently-coded (slower) software module. Alternative adjustments may be employed to provide different relative weights between hardware platform performance and object module performance.

[0202] Those skilled in the art will further understand that, although multiplication operations are used in all exemplary formulas PB1, PB2 and PB3, this is only exemplary. Other arithmetic operations and / or other mathematical functions may be used to generate effective benchmark formulas PB[s,h] based on a combination of platform timing data (PTD) 135 and instruction mix decomposition (IMB) 120.

[0203] Normalization: In some embodiments of the present systems and methods, it may be desirable to normalize the benchmark formula 150 to a common standard. For example, a first formula may produce an intermediate value that is then normalized to a common standard. In some embodiments, when the present systems and methods are put into use, appropriate normalization factors may be identified over time. For example, it may be desirable to normalize the benchmark formula so that the final benchmark result 150 is substantially the same as the actual time taken for the target software module 115 to execute on the target hardware platform 130. Over time, it may be found that once the target module 115 is actually running on the target platform 130, the initial formula will typically deliver a value that indicates an execution time that is shorter than the actual execution time. In this case, appropriate normalization may be introduced so that the final benchmark value 150 is typically very close to the actual execution time. Other normalizations are also contemplated.

[0204] Generic Benchmark Formulas for System Analysis: Above, several exemplary performance benchmark formulas 150 (PB1, PB2, PB3) have been proposed as alternative benchmark calculations. In some embodiments of the present system and method, and for design evaluation purposes of the object module 115 (software) and the hardware platform 130, it is expected that design engineers will utilize a specific generic performance benchmark (e.g., one of PB1, PB2, or PB3) to evaluate a variety of different potential hardware / software combinations.

[0205] IX. Benchmarks

[0206] In an embodiment of the present system and method, the goal of the benchmark suite 102 is to aid in the development of an optimal design for a real-time, often mission-critical system 200. During the development process, a design engineer may consider multiple alternative hardware platforms 130 for use as a controller for (and integration into) a real-time system. Also during the development process, for any one of the potential hardware platforms 130, a design engineer may suggest and evaluate multiple alternative software architectures with multiple alternative software modules 115.

[0207] A single hardware platform 130 may be capable of running several different software modules 115; similarly, a single potential software module 115 may be capable of executing on several different potential hardware platforms 130. (In some cases, a single software module written in a high-level language may be compiled into different target object modules 115, each of which uses different machine-level code 320 for a different target hardware platform 130.) The benchmark suite 102 of the present system and method can help design engineers evaluate and compare different combinations of software modules 115 and hardware platforms 130 to determine which combinations are likely to deliver the best performance (usually the fastest performance) for a given task.

[0208] In an embodiment of the present system and method, and based on multiple simulation runs with varying hardware platforms 130 and / or varying target modules 115, the BG tool 105.3 can generate comparative baseline values ​​155 and / or comparative performance statistics 155 for different combinations / pairings of the target hardware platform 130 and various target modules 115 in actual use.

[0209] Figure 6 An exemplary comparison benchmark 155 is shown in diagram form 155 comparing benchmarks 150 obtained via the present system and method when evaluating their different target modules 115 (in the figure, applications, e.g., "APP-A," "APP-B," "APP-C") and their performance when executed on two different target hardware processors 130 ("Processor 1," "Processor 2"). It should be understood that in order to obtain a valid comparison, the benchmarks 150 obtained via application of a single type of benchmark calculation 150 (e.g., only B1, only B2, or only B3) will be compared. Figure 6 All six benchmarks (150.1, 150.2, ... 150.6) are shown in FIG. In the exemplary embodiment shown, the units of benchmarks 150.n are milliseconds, and higher benchmarks reflect longer execution times, ie, slower performance.

[0210] Clearly, in the example shown, executing APP-B on hardware processor 2 yields the best performance benchmark 150.4 (at 0.328638 milliseconds), executing APP-B on processor 1 yields the second best performance benchmark 150.3 (at 0.649929 milliseconds), and overall, APP-B appears to deliver better task performance (i.e., higher speed performance) than either APP-A or APP-B. Based on such an analysis and comparison, a system design engineer can make a decision about which application / processor combination to use for a controller.

[0211] It will be apparent to those skilled in the relevant art(s) that generating only six benchmark values ​​150 is merely exemplary, and that many more combinations of application modules 115 and hardware processors 130 may be evaluated and compared.

[0212] Exemplary Comparison Applications: It will be apparent to those skilled in the relevant art that in alternative embodiments, more fine-grained or detailed benchmark comparisons 155 may be generated. These may include, for example but not limited to:

[0213] (A) Comparison of benchmarks generated for triple combinations (eg, different hardware processors 130, different object modules 115.1, and different input test vectors 115.2) (simulating different environmental and operating conditions 115.2).

[0214] (B) Comparison of benchmarks generated for triple combinations such as different hardware processors 130, different object modules 115.1, and different configurations of supporting hardware for processing (e.g., different configurations of cache memory or different amounts of memory in controller memory 130.3).

[0215] (C) comparing benchmarks for object modules generated using different compilers (from the same source code); and / or benchmarks generated for multiple object modules, where the multiple object modules are all generated from common source code and using a common compiler, but with different compiler optimization settings.

[0216] (D) In ​​an embodiment, the system and method can generate multiple unique benchmarks for a single identical object code module for each of multiple processors so that the system and method can determine which hardware processor (among multiple potential processors or "target platforms") is optimal for the code module.

[0217] (V) In another embodiment, the system and method can generate multiple performance benchmarks for a single specified processor, with each code module in a plurality of target software modules having a unique benchmark, with the goal of selecting only one target software module for actual use and deployment. Each of the plurality of software modules can have the same or similar functionality, but differ in the detailed code or algorithms. In this way, the system and method can identify those code changes that are optimal for a particular specified target platform (e.g., a particular specified hardware processor).

[0218] (E) In another embodiment of the present system and method, multiple different target software modules have different functions (e.g., control different hardware components of a jet engine), and the multiple target modules are intended to combine (i.e., integrate) performance (possibly operating in parallel or sequentially, or both). Then, each target module may have a very different set of code.

[0219] The present systems and methods can generate a first set of benchmarks for each target module when running on a first processor or first target hardware platform, and then generate a second set of benchmarks for each target module when running on a second (different) target processor or different target hardware platform, etc. In this way, guided by the benchmark results for multiple different cooperating target modules, the design engineer can select a target hardware processor that provides sufficient performance time for all target software modules (thereby providing optimal performance for the integrated software system of all target modules).

[0220] Other combinations of analog variations and pairings of modules 115 and processors 130 are also contemplated within the scope of the appended claims.

[0221] X. Example Computers Used for Benchmark Testing

[0222] Figure 7 A block diagram or system level diagram of an exemplary benchmark computer 110 (e.g., any of computers 110.1, 110.2, and / or 110.3) that may be employed in accordance with the present systems and methods is shown. The computer 110 may implement or execute, for example, any of the benchmark application tools 105. The computer 110 typically has a motherboard (not shown) that typically holds and interconnects various microchips 715 / 720 / 725 and volatile and non-volatile memory or storage devices 730 / 735 that together implement the hardware level operation of the computer 110 and also implement the operations of the present systems and methods 102, 170. The computer 110 may include, for example, but not limited to:

[0223] The hardware microprocessor 715, also referred to as a central processing unit (CPU) 715, provides overall operational control for the computer 110. This includes, but is not limited to, receiving data from data files or from connections with other computers 110, receiving data from a target hardware platform 130, and sending data or files to a target hardware platform 130. The microprocessor 715 is also configured to perform the arithmetic and logic operations necessary to implement the present systems and methods 102, 170.

[0224] Those skilled in the relevant art will appreciate that the hardware microprocessor 715 is distinct from the target processor 130 . 1 of the target hardware platform 130 , and similarly, the memories 720 , 730 , 735 are distinct from the memory 130 . 3 of the target hardware platform 130 or controller 130 .

[0225] Static memory or firmware 720 can store non-volatile operating code, including but not limited to operating system code, computer code for local processing and analysis of data, and computer code that can be specifically used to enable the computer 110 to implement the methods described in this document and other methods within the scope and spirit of the appended claims. The CPU 715 can use the code stored in the static memory 720 and / or the dynamic memory 730 and / or the non-volatile data storage device 735 to implement the methods described in this document and other methods within the scope and spirit of the appended claims.

[0226] The control circuit 725 can perform various tasks, including data and control exchanges and input / output (I / O) tasks, network connection operations, control of the bus 712, and other tasks generally known in the art of processing systems. The control circuit 725 can also control or interface with the non-volatile data storage device 735.

[0227] The control circuit 725 may also support functions such as external input / output (eg, via a USB port, an Ethernet port, or wireless communication, not shown).

[0228] Volatile memory 730 , such as dynamic RAM (DRAM), may be used to temporarily store data or program code. Volatile memory 730 may also be used to temporarily store some or all code from static memory 720 .

[0229] The non-volatile memory may take the form of a hard drive, a solid state drive (including flash drives and memory cards), recording on magnetized tape, storage on a DVD or similar optical disk, or other forms of non-volatile storage now known or to be developed.

[0230] XI. Further Embodiments

[0231] Benchmarking Object Code Compilers and Compiler Optimization Settings: In certain embodiments of the present systems and methods, and in certain applications of the benchmark tool 105, the same target source code may be compiled into multiple object code modules (115.1.1, 115.1.2, ... 115.1.n), each of which is generated by one or both of (i) a different object code compiler and / or (ii) a generic object code compiler with different optimization settings. All of the object code modules may then be executed on a generic target hardware platform (130, 130") and utilizing the same input test vectors (115.2"). The resulting benchmark values ​​150 will then indicate how different compilers and / or different optimization settings will affect the actual performance of the generic source code running on the generic hardware platform.

[0232] Automatic Benchmarking During Code Development: In an embodiment of the present system and method, the benchmark suite 102 (or some elements of the benchmark suite, such as TMEA 105.1 and BG 105.3) can be integrated into a source code development environment (IDE). In an embodiment, an IDE employing the present system and method can generate object code on the fly as new code algorithms are developed. The object code can then be analyzed / benchmarked against a target processor to determine, for example, whether a new or modified algorithm throws an operation timing outside a specified threshold.

[0233] Single Computer and Distributed Computers: In some embodiments of the present systems and methods, the phases or steps of generating a benchmark for a particular code module and a particular hardware processor may be performed by a single computer 110, and the phases may be performed substantially consecutively in time (i.e., each successive analysis phase may be initiated substantially after the immediately preceding analysis phase is completed). In alternative embodiments, the phases of analysis may be performed by separate computers 110.n of a distributed system, and may be performed with considerable time gaps (e.g., minutes, hours, days, or longer periods) between analysis phases.

[0234] In some embodiments, some application tools 105 of the present systems and methods may be remotely available to programmers or design engineers via the Internet 190 (or via a company or organization's intranet 190) for time-distributed or spatially-distributed execution of the present systems and methods. Thus, for example, an instruction mix breakup (IMB) 120 may be generated by a programmer at a first location using a first computer 110.1; platform timing data (PTD) 135 may be generated by a hardware platform manufacturer at a second location via a second computer 110.2; and then a single performance benchmark (PB) 150 may be generated for the pair of IMB 120 and PTD 135 by a third computer 110.3 at a third location (e.g., by a performance benchmark service provider, such as General Electric). It will be further understood that the present benchmarking system and method 100 may be understood to be implemented as a whole via one or more application tools 105 and / or via one or more computers 110; and it will be further understood that any one or more application tools 105.n, and / or any one or more related method steps 170.n, and / or any one or more computers 110.n that may perform aspects or portions of the present system and method may themselves be understood to be their own systems and / or methods within the scope of the present disclosure.

[0235] Non-transitory storage media for instructions: In some embodiments of the present system and method 100, and as part of implementing the present method, the system 102 may include, incorporate or obtain processing instructions for the application tool 105 via one or more non-transitory computer-readable media (sometimes also referred to as "non-transitory computer-readable storage media," "tangible computer-readable storage media," and other similar phrases) that store one or more software programs 105, software applications 105, application tools (105), and application instructions that cause (or when executed can cause) the computer 110 of the benchmark computer system 102 to perform the processes or methods described in this document. Such non-transitory computer-readable media may include, for example, but are not limited to: floppy disk drives, hard disk drives, solid-state drives, flash drives, optical computer disks (CDs), digital video disks (DVDs), read-only memories (ROMs), programmable read-only memories (PROMs), field programmable gate arrays (FPGAs), and holographic memories. In some embodiments of the present systems and methods, a processor or microcontroller (not shown) of computer 110 may have an integrated memory (e.g., in the form of a ROM, PROM, or FPGA) (not shown) that serves as a “non-transitory computer-readable storage medium” for the present systems and methods.

[0236] in conclusion

[0237] Those skilled in the art can make alternative embodiments, examples and modifications that the present disclosure will still include, especially based on the above teachings.In addition, it should be understood that the terms used to describe the present disclosure are intended to have the nature of descriptive words rather than limiting words.

[0238] Those skilled in the art will also recognize that various modifications and variations of the above preferred and alternative embodiments may be configured without departing from the scope and spirit of the present disclosure. Therefore, it should be understood that within the scope of the appended claims, the present disclosure may be practiced in a manner different from that specifically described herein.

[0239] Further aspects of the invention are provided by the subject matter of the following clauses:

[0240] 1. A computer-readable non-transitory storage medium storing instructions, which, when executed by one or more computers of a benchmarking computer system, cause the one or more computers to perform a method for benchmarking, the method comprising: receiving a first instruction mix breakdown (IMB) of a plurality of machine language instructions of a first object code module, the first IMB indicating an execution count of each of the plurality of machine language instructions; receiving first platform timing data (PTD) of a first target hardware platform, the first PTD including single instruction execution timing data for each of the plurality of machine language instructions of the first target hardware platform; and calculating a first performance benchmark (PB) indicating an expected performance speed of the first object code module when executed on the first target hardware platform based on the first IMB and the first PTD.

[0241] 2. A computer-readable non-transitory storage medium according to any preceding clause, wherein the step of receiving the first IMB comprises: receiving a first IMB, for which the execution count of each machine language instruction is determined from actual execution of the first object code module on a physical or simulated first target hardware platform.

[0242] 3. A computer-readable non-transitory storage medium according to any preceding clause, further comprising generating the first IMB, the generating comprising: loading the first object code module for execution on a physical or simulated first target hardware platform; providing an input test vector representing a simulated operating environment as input to the physical or simulated first target hardware platform; executing the first object code module on the physical or simulated first target hardware system; and generating the first IMB based on the execution of the first object code module using the input test vector, wherein: when the first object code module is executed on the first target platform in a specified physical or simulated environment, the first IMB substantially reflects the execution count of each of the plurality of machine language instructions.

[0243] 4. The computer-readable non-transitory storage medium according to any preceding clause, further comprising generating said first platform timing data for said target hardware platform.

[0244] 5. A computer-readable non-temporary storage medium according to any of the preceding items, wherein generating the first platform timing data of the target hardware platform comprises: continuously executing a selected microcode command of the target hardware platform N times; continuously identifying a microcode execution time MET required to execute the selected microcode command of the target hardware platform N times; and dividing the MET by N to determine a single-pass execution time of the selected microcode command.

[0245] 6. A computer-readable non-transitory storage medium according to any preceding clause, wherein generating the performance benchmark for the first object code module and the first target hardware platform comprises: for each corresponding machine language instruction in the IMB, generating a plurality of corresponding products by multiplying real-time execution timing data of the corresponding instruction by a number of occurrences of the corresponding instruction during the execution of the first object code module; and generating the performance benchmark as a sum of the corresponding products.

[0246] 7. A computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises: receiving second platform timing data (PTD) for a second target hardware platform, the second PTD comprising real-time execution timing data for each of a plurality of machine language instructions of the second target hardware platform; and calculating a second performance benchmark (PB) based on the first IMB and the second PTD, the second performance benchmark (PB) indicating an expected performance speed of the first object code module when executed on the second target hardware platform.

[0247] 8. A computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises: receiving a second instruction mix breakdown (IMB) of a plurality of machine language instructions of a second object code module, the IMB indicating an execution count of each of the plurality of machine language instructions; and calculating a second performance benchmark (PB) based on the second IMB and the first PTD, the second performance benchmark (PB) indicating an expected performance speed of the second object code module when executed on the first target hardware platform.

[0248] 9. A computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises: receiving a second object code module that is different from the first object code module and provides the same functionality as the first object code module, wherein the first performance benchmark and the second performance benchmark indicate relative performance of two different object code modules configured to perform the same functionality on the first target hardware platform.

[0249] 10. A computer-readable non-transitory storage medium storing instructions that, when executed by one or more computers of a benchmark computer system, cause the one or more computers to perform a method for benchmarking, the method comprising: generating a first instruction mix breakdown (IMB) of a plurality of machine language instructions of a first object code module, wherein: the first IMB indicates a respective execution count of each respective machine language instruction of the plurality of machine language instructions on a first target hardware platform; and determining the first IMB based on actual execution of the first object code module on at least one of a physical first target hardware platform and a simulation of the physical first target hardware platform (P / S first target hardware platform).

[0250] 11. A computer-readable non-transitory storage medium according to any of the preceding items, wherein the method for generating the first IMB further comprises: providing an input test vector representing a physical or simulated operating environment as input to the P / S first target hardware platform; executing the first object code module on the P / S first target hardware platform; and generating the first IMB based on the execution of the object code module using the input test vector.

[0251] 12. A computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises: receiving first platform timing data (PTD) for the first physical target hardware platform, the PTD comprising individual instruction execution timing data for each of a plurality of machine language instructions of the first physical target hardware platform; and calculating a first performance benchmark (PB) indicating an expected performance speed of the first object code module when executed on the first target hardware platform based on the first IMB and the first PTD.

[0252] 13. A computer-readable non-transitory storage medium according to any preceding clause, wherein calculating the performance benchmark when the first object code module is executed on the first target hardware platform comprises: for each corresponding machine language instruction in the IMB, obtaining a corresponding product by multiplying the real-time execution timing data of the corresponding instruction by the number of occurrences of the corresponding instruction during the execution of the first object code module, thereby generating a plurality of corresponding products; and generating the performance benchmark as the sum of the corresponding products.

[0253] 14. The computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises generating the first platform timing data for the target hardware platform.

[0254] 15. A computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises: receiving second platform timing data (PTD) of a second target hardware platform, the second PTD comprising real-time execution timing data of each of a plurality of machine language instructions of the second target hardware platform; and calculating a second performance benchmark (PB) based on the first IMB and the second PTD, the second performance benchmark (PB) indicating an expected performance speed of the first object code module when executed on the second target hardware platform.

[0255] 16. A computer-readable non-transitory storage medium according to any preceding clause, wherein the method further comprises: generating a second instruction mix breakdown (IMB) of a plurality of machine language instructions of a second object code module, the IMB indicating an execution count of each of the plurality of machine language instructions; and calculating a second performance benchmark (PB) based on the second IMB and the first PTD, the second performance benchmark (PB) indicating an expected performance speed of the second object code module when executed on the first target hardware platform.

[0256] 17. A computing system for determining a performance benchmark, the computing system comprising one or more computers, each of the one or more computers comprising: a central processing unit (CPU), a memory, and data input and output resources, wherein the computing system is configured to execute instructions via the CPU, the instructions causing the computing system to: store a first instruction mix breakdown (IMB) of a plurality of machine language instructions of a first object code module in the memory, the IMB indicating an execution count of each of the plurality of machine language instructions when executed by a processor of a target hardware platform; store first platform timing data (PTD) of the first target hardware platform in the memory, the PTD comprising individual instruction execution timing data for each of the plurality of machine language instructions of the first target hardware platform; and calculate a first performance benchmark (PB) based on the first IMB and the first PTD, the first performance benchmark (PB) indicating an expected performance speed of the first object code module when executed on the first target hardware platform.

[0257] 18. A computing system according to any preceding clause, wherein the computing system is further configured to execute instructions via the CPU, the instructions causing the computing system to: for each corresponding machine language instruction in the first IMB, generate a plurality of corresponding products by multiplying real-time execution timing data of the corresponding instruction by a number of occurrences of the corresponding instruction during the execution of the first object code module to obtain a corresponding product; and calculate the performance benchmark as a sum of the corresponding products.

[0258] 19. A computing system according to any of the preceding clauses, wherein the computing system is further configured to execute instructions via the CPU, the instructions causing the computing system to: provide input test vectors representing an operating environment as input to a physical or simulated first target hardware platform; execute the first object code module on the physical or simulated first target hardware platform; and generate the first IMB based on the execution of the first object code module on the physical or simulated first target hardware platform using the input test vectors.

[0259] 20. A computing system according to any of the preceding clauses, wherein the computing system includes a plurality of distributed computers, wherein: the first IMB can be obtained on a first computer in the plurality of distributed computers; and the first performance benchmark can be generated on a second computer in the plurality of distributed computers, thereby generating the first performance benchmark on the second computer without the second computer obtaining access to machine-level code of the first object code module.

Claims

1. A computer-readable non-transitory storage medium storing instructions, which, when executed by one or more computers of a benchmarking computer system, causes the one or more computers to perform a method for benchmarking, the method include: receiving a first instruction mix decomposition IMB of a plurality of machine language instructions of a first object code module, the first IMB indicating an execution count of each machine language instruction of the plurality of machine language instructions; receiving first platform timing data PTD for a first target hardware platform, the first PTD comprising single instruction execution timing data for each of a plurality of machine language instructions of the first target hardware platform; as well as calculating, based on the first IMB and the first PTD, a first performance benchmark PB indicating an expected performance speed of the first object code module when executed on the first target hardware platform; Characterized in that the method further comprises: Generating the first PTD for the first target hardware platform, wherein the generating comprises: executing a selected microcode command of the first target hardware platform N times in succession, wherein the selected microcode command requires at least one parameter, wherein two or more sets of the at least one parameter are determined, the two or more sets of the at least one parameter having different values ​​for the at least one parameter, and executing the selected microcode command a plurality of times using the two or more sets of the at least one parameter; identifying a microcode execution time required to execute the selected microcode command of the first target hardware platform N times in succession; and The microcode execution time is divided by N to determine a single pass execution time for the selected microcode command.

2. The computer-readable non-transitory storage medium according to claim 1, It is characterized in that in, The step of receiving the first IMB includes: A first IMB is received for which the execution count of each machine language instruction is determined from actual execution of the first object code module on a physical or simulated first target hardware platform.

3. The computer-readable non-transitory storage medium according to claim 1 or 2, It is characterized in that Further comprising generating the first IMB, wherein the generating comprises: loading the first object code module for execution on a physical or simulated first target hardware platform; providing as input to the physical or simulated first target hardware platform an input test vector representing a simulated operating environment; executing the first object code module on the physical or simulated first target hardware platform; and generating the first IMB based on the execution of the first object code module using the input test vector, wherein: The first IMB reflects an execution count of each of the plurality of machine language instructions when the first object code module is executed on the first target hardware platform in a specified physical or simulated environment.

4. The computer-readable non-transitory storage medium according to claim 1, It is characterized in that in, Generating the performance benchmark of the first object code module and the first target hardware platform includes: generating, for each corresponding machine language instruction in the IMB, a plurality of corresponding products by multiplying real-time execution timing data for the corresponding instruction by a number of occurrences of the corresponding instruction during the execution of the first object code module to obtain a corresponding product; and A sum of the corresponding products is generated as the performance benchmark.

5. The computer-readable non-transitory storage medium according to claim 1, It is characterized in that in, The method further comprises: receiving second platform timing data PTD for a second target hardware platform, the second PTD comprising real-time execution timing data for each of a plurality of machine language instructions of the second target hardware platform; and A second performance benchmark PB is calculated based on the first IMB and the second PTD, wherein the second performance benchmark PB indicates an expected performance speed when the first object code module is executed on the second target hardware platform.

6. The computer-readable non-transitory storage medium according to claim 1, It is characterized in that in, The method further comprises: receiving a second instruction mix decomposition IMB of a plurality of machine language instructions of a second object code module, the IMB indicating an execution count of each of the plurality of machine language instructions; and A second performance benchmark PB is calculated based on the second IMB and the first PTD, the second performance benchmark PB indicating an expected performance speed of the second object code module when executed on the first target hardware platform.

7. The computer-readable non-transitory storage medium according to claim 6, It is characterized in that in, The method further comprises: A second object code module different from the first object code module and providing the same functionality as the first object code module is received, wherein the first performance benchmark and the second performance benchmark indicate relative performance of two different object code modules configured to perform the same functionality on the first target hardware platform.

8. A computer-readable non-transitory storage medium storing instructions, It is characterized in that When the instructions are executed by one or more computers of a benchmark computer system, the one or more computers are caused to execute the method for benchmarking according to claim 1, wherein the method further comprises: A first instruction mix IMB of a plurality of machine language instructions of a first object code module is generated, wherein: The first IMB indicates a respective execution count of each respective machine language instruction of the plurality of machine language instructions on a first target hardware platform; and The first IMB is determined based on actual execution of the first object code module on at least one of a physical first target hardware platform and a simulated first target hardware platform.

9. The computer-readable non-transitory storage medium according to claim 8, It is characterized in that in, The method of generating the first IMB further comprises: providing as input to the physical first target hardware platform an input test vector representing a physical or simulated operating environment; executing the first object code module on the physical first target hardware platform; and The first IMB is generated based on the execution of the first object code module using the input test vector.

10. A computing system for determining a performance benchmark, It is characterized in that The computing system includes one or more computers, each of the one or more computers including: Central processing unit CPU, memory and data input and output resources, The computing system is configured to execute instructions via the CPU, the instructions causing the computing system to: storing a first instruction mix decomposition IMB of a plurality of machine language instructions of a first object code module in the memory, the first IMB indicating an execution count of each of the plurality of machine language instructions when executed by a processor of a target hardware platform; storing first platform timing data (PTD) for a first target hardware platform in the memory, the first PTD comprising individual instruction execution timing data for each of the plurality of machine language instructions for the first target hardware platform; and calculating a first performance benchmark PB based on the first IMB and the first PTD, the first performance benchmark PB indicating an expected performance speed of the first object code module when executed on the first target hardware platform; The computing system is configured to execute instructions via the CPU so that the computing system generates the first PTD for the first target hardware platform, wherein the generating includes: executing a selected microcode command of the first target hardware platform N times in succession, wherein the selected microcode command requires at least one parameter, wherein two or more sets of the at least one parameter are determined, the two or more sets of the at least one parameter having different values ​​for the at least one parameter, and executing the selected microcode command a plurality of times using the two or more sets of the at least one parameter; identifying a microcode execution time required to execute the selected microcode command of the target hardware platform N times in succession; and The microcode execution time is divided by N to determine a single pass execution time for the selected microcode command.

11. The computing system according to claim 10, It is characterized in that in, The computing system is further configured to execute instructions via the CPU, the instructions causing the computing system to: for each corresponding machine language instruction in the first IMB, obtaining a corresponding product by multiplying real-time execution timing data for the corresponding instruction by a number of occurrences of the corresponding instruction during the execution of the first object code module, thereby generating a plurality of corresponding products; as well as The sum of the corresponding products is calculated as the performance benchmark.

12. The computing system according to claim 10, It is characterized in that in, The computing system includes a plurality of distributed computers, wherein: The first IMB is available on a first computer among the plurality of distributed computers; and The first performance benchmark can be generated on a second computer of the plurality of distributed computers, such that the first performance benchmark is generated on the second computer without the second computer gaining access to machine level code of the first object code module.

13. A computing system according to any one of claims 10 to 12, It is characterized in that in, The computing system is further configured to execute instructions via the CPU, the instructions causing the computing system to: providing as input to a physical or simulated first target hardware platform an input test vector representing an operating environment; executing the first object code module on the physical or simulated first target hardware platform; as well as The first IMB is generated based on the execution of the first object code module on the physical or simulated first target hardware platform using the input test vectors.

14. A computing system according to any one of claims 10 to 12, It is characterized in that in, The computing system is configured to execute instructions via the CPU, the instructions causing the computing system to: A first instruction mix IMB of a plurality of machine language instructions of a first object code module is generated, wherein: The first IMB indicates a respective execution count of each respective one of the plurality of machine language instructions on a first target hardware platform; and The first IMB is determined based on actual execution of the first object code module on at least one of a physical first target hardware platform and a simulated first target hardware platform.

Citation Information

Patent Citations

  • Application execution profiling in conjunction with a virtual machine

    EP1331565A1