Identification of Performance Hot Spots in Neural Networks

The method addresses the challenge of identifying performance hotspots in neural network models by using mapping files to translate performance hotspots from lower-level to higher-level representations, allowing for efficient optimization and improved execution efficiency.

JP7695756B2Active Publication Date: 2025-06-19INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023558441
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-30
Filing Date
2022-03-08
Publication Date
2025-06-19
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

Identifying performance hotspots within neural network models is challenging, especially in models compiled using modern compilers, as it requires determining which instructions or operations consume the most execution time or CPU resources.

Method used

A computer-implemented method that collects sample data with instruction addresses related to a neural network model, determines performance hotspots, and uses mapping files to map these hotspots from lower-level intermediate representations to higher-level representations, facilitating optimization.

Benefits of technology

This method enables efficient identification and optimization of performance hotspots at different levels of compilation, improving the execution efficiency of neural network models by reducing execution time and CPU usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695756000002
    Figure 0007695756000002
  • Figure 0007695756000003
    Figure 0007695756000003
  • Figure 0007695756000004
    Figure 0007695756000004
Patent Text Reader

Abstract

An implementation for identifying performance hot spots includes collecting sample data having instruction addresses, the sample data being for a neural network model, and determining instructions within the instruction addresses that are performance hot spots. A list file is used to map instructions of the sample data that are performance hot spots to locations in a lower level intermediate representation. A mapping file is used to map locations of the lower level intermediate representation that are performance hot spots to operations in one or more higher level representations, one or more of the operations corresponding to the performance hot spots, the mapping file being generated from compiling the neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to computer systems, and more particularly to a computer-implemented method, a computer system, and a computer program product configured and arranged to identify performance hot spots in neural networks.

Background Art

[0002] An artificial neural network, commonly referred to as a neural network, is a computing system inspired by the biological neural networks that make up the brains of animals. An artificial neural network is based on a collection of connected units or nodes called artificial neurons that roughly model the neurons in a biological brain. Each connection, like a synapse in a biological brain, can transmit a signal to other neurons. An artificial neuron that receives a signal can then process the signal and send the signal to neurons connected to that artificial neuron. The "signal" at a connection is a real number, and the output of each neuron is calculated by some non-linear function of the sum of the inputs. The connections are called edges. Neurons and edges typically have weights that adapt as learning progresses. The weights increase or decrease the strength of the signal at a connection. A neuron may have a threshold such that a signal is transmitted only if the aggregated signal exceeds that threshold. Typically, neurons are grouped into layers. Different layers may perform different transformations on the inputs to those layers. A signal moves from the first layer (input layer), through one or more hidden layers, possibly traversing a layer multiple times, and then to the last layer (output layer).

[0003] A neural network can be made very complex, consisting of a large number of compiled instructions. Sometimes, there may be performance hotspots for one or more instructions. A performance hotspot in computer science is most commonly defined as a region of a computer program where a high percentage of the executed instructions occur, or where the most time is spent during program execution, or both. When a program is randomly interrupted, it is often found that the program counter (a pointer to the next instruction to be executed) contains the addresses of instructions within a specific range, which may indicate code that requires optimization. However, it can be difficult to determine, identify, or both, performance hotspots that require optimization within a neural network model, especially within a neural network model compiled using a modern compiler, and thus improvement is needed. Summary of the Invention

[0004] Embodiments of the present invention are directed to a computer-implemented method for identifying performance hot spots of a neural network for optimization. A non-limiting, exemplary computer-implemented method includes collecting sample data having instruction addresses, where the memory sample data relates to a neural network model. The method includes determining instructions within the instruction addresses that are performance hot spots and using a list file to map the instructions of the sample data that are performance hot spots to locations within a lower-level intermediate representation. The method also includes using a mapping file to map the locations of the lower-level intermediate representation that are performance hot spots to operations within one or more higher-level representations, where one or more of these operations correspond to the performance hot spots and the mapping file is generated from compiling the neural network model.

[0005] This enables more efficient determination and identification of performance hot spots in a higher-level intermediate representation related to the neural network model, resulting in an improvement over known methods for performance hot spots. The higher-level representation is easier for a human user to read, understand, and modify, and thus optimization can be performed on the performance hot spots that affect the neural network model, thereby improving the execution of the neural network model. Further, in one or more embodiments, improvements for identifying performance hot spots at any suitable level can be utilized. One or more embodiments can find the largest hot spots such that a minimal amount of effort can be utilized to optimize the performance of the hot spots.

[0006] In addition to, or as an alternative to, one or more of the features described above or below, in further embodiments of the present invention, the method may include mapping operations within one or more higher level representations to nodes within a neural network model using a mapping file. Thus, the improvement advantageously identifies which nodes within the neural network model are performance hotspots.

[0007] In addition to, or as an alternative to, one or more of the features described above or below, in further embodiments of the present invention, the method may include determining which nodes within a neural network model represent performance hotspots based on mapping from one or more higher level representations mapped from a lower level intermediate representation mapped from sample data. Thus, the improvement advantageously tracks performance hotspots from a more difficult machine-readable code such as binary code to operations at a level that is easier for a human user to read, understand, and modify. Compilation is a top-down process, but one or more embodiments provide a bottom-up mapping process that maps relationships from a lower level up using a mapping file.

[0008] In addition to, or as an alternative to, one or more of the features described above or below, in further embodiments of the present invention, the method may include classifying performance hotspots at different levels of compilation of a neural network model using a mapping file, where the different levels include from a lower level intermediate representation to one or more higher level representations. Thus, the improvement advantageously provides options at different levels for optimizing performance hotspots within the neural network model.

[0009] In addition to, or as an alternative to, one or more of the features described above or below, in further embodiments of the present invention, sample data further includes information from one or more counters.

[0010] In addition to, or as an alternative to, one or more of the features described above or below, in further embodiments of the present invention, a performance hot spot includes one or more metrics that meet or exceed one or more thresholds.

[0011] In addition to, or as an alternative to, one or more of the features described above or below, in further embodiments of the present invention, the method includes determining one or more operations within one or more higher levels of representation that optimize a performance hot spot and thereby address the performance hot spot.

[0012] Other embodiments of the present invention implement the features of the foregoing method in a computer system and a computer program product.

[0013] Other technical features and advantages are realized by the technology of the present invention. Embodiments and aspects of the present invention are described in detail herein and are considered part of the claimed subject matter. For a better understanding, please refer to the detailed description and the drawings.

[0014] The details of the proprietary rights described herein are specifically pointed out and clearly claimed in the claims at the end of this specification. The foregoing and other features and advantages of each embodiment of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0016] One or more embodiments of the present invention provide a computer-implemented method, a computer system, and a computer program product arranged and configured to identify and optimize performance hot spots of a neural network. One or more embodiments are configured to solve the problem of identifying performance hot spots in neural network nodes, including using multi-level intermediate representation (MLIR) instruction address (IA) mapping files in a neural network model to identify performance hot spots at different compilation levels in order to optimize the performance of the neural network. According to one or more embodiments, the MLIR IA mapping file is utilized to map samples of instruction addresses to name / position elements in MLIR compilation, and the MLIR compilation can include neural network nodes in addition to MLIR dialect operations and MLIR passes. The MLIR IA mapping file and MLIR dialect are utilized to annotate data flow graphs and elements of MLIR using performance hot index information. Further, one or more embodiments search for and find the range of the most hot performance instructions in addition to the most hot neural network operations and the most hot operations at each compiler level.

[0017] In performance analysis, for example, there are comprehensive practical questions such as which part of the program takes the most execution time, showing the actual performance problems to be solved and the solutions to the performance problems. The code area that takes the most execution time (also called the central processing unit (CPU) time) is called a performance hot spot. A performance hot spot can also be called a part of the code that requires the most CPU usage (e.g., the CPU percentage of instructions over a specific period compared to other parts of the code). Since the performance in a neural network model can be greatly improved with little effort in making the performance hot spot faster, the performance hot spot is the best place to tune and optimize. It may be difficult to find the performance hot spots within a neural network model. There may be a single module without function-level symbols existing within a neural network module, and there may be hot instructions (i.e., instructions that cause performance hot spots) distributed among more than 10,000 instructions. Even if it is possible to find the most hot basic block patterns, it is difficult to know which parts should be optimized.

[0018] As a technical solution and advantage for improving the determination and identification of performance hotspots of a neural network model, one or more embodiments are configured to identify which neural network nodes in a graph are performance hotspots, or which operations in MLIR are performance hotspots, or both, thereby identifying the performance hotspots of the neural network. One or more embodiments provide opportunities for optimizing (i.e., improving) performance hotspots at different levels, including different levels of compilation of the neural network model, such as the compilation level of MLIR. It should be understood that performance hotspots, hotspots, hotspot information, hot index information, hot, etc. can be used synonymously to refer to one or more instructions and / or operations, or both, that utilize more execution time (i.e., CPU time) or require more CPU usage / rate than other instructions / operations or a predefined threshold, or both. Optimization can be performed at any level by determining performance hotspots in one of the lower-level or higher-level or intermediate representations between the two to improve the execution of the neural network model and thereby improve the functionality of the computer system (itself) that executes the neural network model. By determining performance hotspots and optimizing the neural network model, it becomes possible to reduce execution time (i.e., runtime), reduce CPU usage, reduce memory usage, reduce bandwidth, etc.

[0019] Referring now to FIG. 1, a computer system 100 is generally shown in accordance with one or more embodiments of the present invention. The computer system 100 can be an electronic computer framework that comprises, or employs, or both, any number and combination of computing devices and networks that utilize various communication technologies, as described herein. The computer system 100 can be modular, easily scalable and extensible, and can have the ability to change for different services or to reconfigure some functions independently of other functions. The computer system 100 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, the computer system 100 can be a cloud computing node. The computer system 100 can be described in the general context of executable instructions by a computer system, such as program modules executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc., that perform particular tasks or implement particular abstract data types. The computer system 100 can be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media, including a memory storage device.

[0020] As shown in FIG. 1, computer system 100 has one or more central processing units (CPUs) 101a, 101b, 101c, etc. (collectively or generally referred to as processor 101). The processor 101 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The processor 101, also called a processing circuit, is coupled via a system bus 102 to a system memory 103 and various other components. The system memory 103 can include a read-only memory (ROM) 104 and a random access memory (RAM) 105. The ROM 104 is coupled to the system bus 102 and may include a basic input / output system (BIOS) that controls certain basic functions of the computer system 100 or its successor such as the Unified Extensible Firmware Interface (UEFI). The RAM is a read-write memory coupled to the system bus 102 for use by the processor 101. The system memory 103 provides a temporary memory space for the operation of the aforementioned instructions during operation. The system memory 103 can include a random access memory (RAM), a read-only memory, a flash memory, or any other suitable memory system.

[0021] The computer system 100 includes an input / output (I / O) adapter 106 and a communication adapter 107 coupled to the system bus 102. The I / O adapter 106 can be a small computer system interface (SCSI) adapter that communicates with a hard disk 108 or any other similar component or both. The I / O adapter 106 and the hard disk 108 are collectively referred to herein as mass storage 110.

[0022] Software 111 for execution on computer system 100 may be stored in mass storage 110. Mass storage 110 is an example of a tangible storage medium readable by processor 101, and software 111 is stored as instructions for execution by processor 101 to cause computer system 100 to operate as described hereinafter in this specification with respect to various figures. Examples of computer program products and execution of such instructions are described in more detail herein. Communication adapter 107 interconnects system bus 102 with network 112, which may be an external network, enabling computer system 100 to communicate with other such systems. In one embodiment, a portion of system memory 103 and mass storage 110 collectively stores an operating system, which may be any suitable operating system, to coordinate the functions of the various components shown in FIG. 1.

[0023] Other input / output devices are shown as being connected to system bus 102 via display adapter 115 and interface adapter 116. In one embodiment, adapters 106, 107, 115, and 116 may be connected to one or more I / O buses, which are connected to system bus 102 via an intermediate bus bridge (not shown). A display 119 (e.g., a screen or display monitor) is connected to system bus 102 by display adapter 115, which may include a graphics controller to improve the performance of graphics-intensive applications and video controllers. A keyboard 121, a mouse 122, a speaker 123, etc. can be interconnected to system bus 102 via interface adapter 116. For example, interface adapter 116 may include a Super I / O chip that integrates multiple device adapters into a single integrated circuit. I / O buses suitable for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe). Thus, as configured in FIG. 1, computer system 100 includes processing capabilities in the form of processor 101, storage capabilities including system memory 103 and mass storage 110, input means such as keyboard 121 and mouse 122, and output capabilities including speaker 123 and display 119.

[0024] In some embodiments, communication adapter 107 can transmit data using any suitable interface or protocol, such as the Internet, Small Computer System Interface, etc. Network 112 can be, in particular, a cellular network, a wireless network, a wide area network (WAN), a local area network (LAN), or the Internet. An external computing device may be connected to computer system 100 via network 112. In some examples, the external computing device may be an external web server or a cloud computing node.

[0025] It should be understood that the block diagram of FIG. 1 is not intended to show that computer system 100 includes all of the components shown in FIG. 1. Rather, computer system 100 can include any suitable fewer components or additional components not shown in FIG. 1 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Further, the embodiments described herein with respect to computer system 100 may be implemented using any suitable logic, which, when referred to herein, can include any suitable hardware (e.g., in particular, a processor, an embedded controller, or an application specific integrated circuit), software (e.g., in particular, an application), firmware, or any suitable combination of hardware, software, and firmware in various embodiments.

[0026] Referring now to FIG. 2, an exemplary neural network model / architecture 200 is shown in accordance with one or more embodiments. The neural network model / architecture 200 may be implemented using one or more software applications 111 on a computer system 100. In one or more embodiments, the computer system 100 may include one or more special hardware, such as an accelerator, for use with the neural network model / architecture. Here, the operation of the exemplary neural network model 200 will be described. During the feedforward operation, each of the set of input neurons 202 transmits a corresponding input voltage in parallel to each row of weights 204. The output current flows from the weights 204 to each hidden neuron 206, and each of the weights 204 has a settable resistance value so as to represent a weighted input. The current output by a particular weight is determined as I = V / r, where V is the input voltage from the input neuron 202 and r is the set resistance of the weight 204. The currents from each weight are added column-wise and flow into the hidden neurons 206. A set of reference weights 207 have fixed resistances and couple their outputs to a reference current supplied to each of the hidden neurons 206. Since the conductance value can only be a positive number, some reference conductance is required to encode both positive and negative values within the matrix. The current generated by the weights 204 is continuously evaluated and is positive. Thus, the reference weights 207 are used to supply a reference current, currents above the reference current are considered to have positive values, and currents below the reference current are considered to have negative values. In some embodiments, each array of weights may include one or more reference weights having static resistances.

[0027] As an alternative to using the reference weights 207, one or more embodiments may use another array of weights 204 to capture negative values. Each approach has advantages and disadvantages. In some embodiments, using the reference weights 207 is more efficient in terms of chip area, but the reference values need to match exactly with each other. In one or more embodiments, the use of another array for negative values does not involve an exact match because each value has a pair of weights for comparison. However, the matrix approach with negative weights uses approximately twice the chip area compared to a single reference weight column. Additionally, the reference weight column generates a current that needs to be copied to each neuron for comparison, while the negative matrix array directly provides the reference value for each neuron. In the negative array embodiment, the weights 204 of both the positive and negative arrays are updated, which also increases the signal-to-noise ratio because each weight value is the difference between two conductance values. These two embodiments provide the same function in encoding negative values, and one of ordinary skill in the art can select the embodiment suitable for the current application.

[0028] The hidden neurons 206 use the currents from the arrays of weights 204 and reference weights 207 to perform some calculations. Next, the hidden neurons 206 output their own voltages to another array of weights 207. This array operates in the same way, where the columns of weights 204 receive the voltages from their respective hidden neurons 206, generate a weighted current output, and this current output is added row by row and supplied to the output neuron 208.

[0029] It should be understood that any number of these stages may be implemented by inserting additional layers of arrays and hidden neurons 206. Note also that some neurons may be constant neurons 209 that supply a constant voltage to the array. The constant neurons 209 can exist between the input neurons 202 or the hidden neurons 206 or both, and are used only during the feed-forward operation.

[0030] In one or more embodiments, during backpropagation, the output neuron 208 supplies voltage in the reverse direction across the array of weights 204. The output layer compares the generated network response with the training data and calculates an error. This error is applied to the array as a voltage pulse, and the height or duration or both of the pulse are adjusted in proportion to the error value. In this example, the rows of weights 204 receive voltage in parallel from each output neuron 208, convert this voltage to current, and this current is summed for each column to provide an input to the hidden neuron 206. The hidden neuron 206 combines the weighted feedback signal with the derivative of the feedforward calculation, stores the error value, and then outputs the feedback signal voltage to each column of weights 204. This backpropagation moves through the entire neural network model 200 until all hidden neurons 206 and input neurons 202 store the error value.

[0031] In one or more embodiments, during weight update, the input neurons 202 and the hidden neurons 206 apply a first weight update voltage forward through the neural network model 200, and the output neurons 208 and the hidden neurons 206 apply a second weight update voltage backward. The combination of these voltages causes a change in state within each weight 204, causing the weight 204 to acquire a new resistance value. In this way, the weights 204 can be trained to adapt the neural network model 200 to the error during processing. It should be noted that the three operating modes of feedforward, backpropagation, and weight update do not overlap with each other.

[0032] FIG. 3 is a block diagram of a system 300 for identifying and optimizing performance hot spots within a neural network model in accordance with one or more embodiments of the present invention. FIG. 3 shows one or more computer systems 302 coupled to a computer system 320 that communicate and exchange information as described herein. The computer system 320 can communicate with the computer system 302 via a wired network, a wireless network, or both. The elements of the computer system 100 may be used within the computer(s) system 302 and the computer system 320, integrated into the computer system 302 and the computer system 320, or both. The software application 304 may be implemented as software 111 executed on one or more processors 101 as described in FIG. 1. Similarly, the compiler 322, the diagnostic tool 306, and the neural network model 200 may be implemented using software 111 configured to execute on one or more processors 101.

[0033] Compiler 322 is utilized to compile a neural network model, such as neural network model 200 or any other neural network model. Compiler 322 and neural network model 200 may be executed within computer system 320, within computer system 302, or within both. Compiler 322 is a multi-level intermediate representation (MLIR) compiler, uses the MLIR compiler framework, or both. MLIR is a reusable and extensible state-of-the-art compiler infrastructure. The MLIR compiler can define dialects and optimization passes. Dialects function as an abstraction level or intermediate representation, and optimization passes enable optimizations at the abstraction level or conversions between abstraction levels. There are immediately available dialects in MLIR, such as llvm, std, scf, and affine. The llvm dialect is a low-level dialect. The llvm dialect wraps LLVM intermediate representation (IR) types and instructions into MLIR types and operations. The std dialect includes standard operations such as load, store, addi, addf, absf, and call. The scf dialect defines control flow operations such as for and if. The affine dialect provides an abstraction of affine operations and analysis. Those skilled in the art understand a multi-level intermediate compiler.

[0034] Furthermore, compiler 322 may be additionally implemented to use Open Neural Network Exchange (ONNX) as a format for representing the input model (e.g., neural network model 200) of compiler 322 in combination with MLIR. ONNX is an open-source machine-independent format, as understood by those skilled in the art, and is widely used for exchanging neural network models. Compiler 322 is described using MLIR, which is the latest open-source compiler infrastructure using the LLVM project for multi-level intermediate representation. The LLVM project is a compiler infrastructure that is a collection of modular reusable compiler and tool chain technologies. Compiler 322 may be referred to synonymously as an ONNX-MLIR compiler, an MLIR compiler, or simply a compiler, or a combination thereof.

[0035] Neural network model 200 can be described or generated or both in ONNX or the ONNX format. Neural network model 200 can be an ONNX neural network model 200, sometimes referred to as an ONNX model. A neural network model may also sometimes be referred to as an artificial intelligence (AI) model. Neural network model 200 can be executed, including (initially) compiling the instructions of neural network model 200 into an executable form (e.g., executable file 334, etc.) using compiler 322 for execution by computer system 302 or computer system 320 or both during runtime. In one or more embodiments, software application 304 may be utilized to cause or initiate or both the compilation of neural network model 200 and the (subsequent) execution of the executable / compiled neural network model 200.

[0036] FIG. 4 is a block diagram of an exemplary architecture of an ONNX-MLIR compiler when compiling an ONNX model (e.g., neural network model 200) according to one or more embodiments of the present invention. In FIG. 4, names preceded by "--" are specified as paths. In FIG. 4, the input is an ONNX model (e.g., neural network model 200), and the output is a library containing the compiled code (e.g., executable file 334). The output library can include an entry function called "_dyn_entry_point_main_graph", and the inputs and outputs of this function are the same as the inputs and outputs of the ONNX model, respectively. To execute inference using the output library, the user writes their program to call the entry function by passing the input to the function to obtain the result. There can be fewer or more dialects, but onnx-mlir has five main dialects, which are onnx, krnl, affine, std, and llvm, organized into four levels of abstraction. The first level of abstraction is a high-level representation of ONNX operations. The first level of abstraction consists of operations in the onnx dialect and the std dialect, and the onnx dialect is automatically generated, for example, via an importer that is a Python (registered trademark) script or another script. The second level of abstraction includes the krnl dialect, the affine dialect, and the std dialect. The krnl dialect provides a representation suitable for loop optimizations that can easily perform affine transformations such as tiling, skewing, and permutation. The krnl dialect functions as an intermediate dialect for efficiently lowering the onnx dialect to lower-level dialects (e.g., affine, std, and llvm). The third level of abstraction includes the affine dialect and the std dialect, where existing optimization paths in MLIR can be freely applied. The fourth level of abstraction includes only the llvm dialect, which is ready to generate bitcode (i.e., binary code) (i.e., executable file 334).

[0037] There exists an MLIR pass for converting one dialect to another and for performing optimizations in a specific dialect. A multi-pass compiler is a type of compiler that processes the source code or abstract syntax tree of a program multiple times. A multi-pass compiler is in contrast to a one-pass compiler that traverses the program only once. Each pass takes the result of the previous pass as input and creates an intermediate output. In this way, the (intermediate) code is improved with each pass until the final pass generates the final code. The onnx dialect is converted to the krnl dialect via the pass --convert-onnx-to-krnl. Subsequently, the krnl dialect (except for some of its operations) is converted to the affine dialect and the std dialect via the pass --convert-krnl-to-affine. The remaining operations in the krnl dialect as well as the operations in the affine dialect and the std dialect are directly converted to instructions in llvm via the pass --convert-krnl-to-llvm. The right side of Figure 4 shows the optimization passes that can be performed at each level of abstraction in a shaded pattern. It should be understood that the list of optimization passes is not exhaustive. As described herein, MLIR provides an extensible framework for transformation of operations using the well-known concept of compiler passes. Since each transformation may have to consider the meaning of any operation, enabling any set of passes for any set of operations poses a significant scaling challenge. However, MLIR addresses this complexity by using Traits and Interfaces to enable the meaning of operations to be abstractly described and for transformations to operate more generically over operations. Traits often describe verification constraints on valid IRs, enabling complex invariants to be captured and checked.

[0038] As shown in FIG. 7, compiler 322 is configured to generate mapping file 330 at each stage of the compile / transformation process. For example, the mapping file 330 of FIG. 3 includes an onnx intermediate representation (IR) mapping file used to convert an ONNX model (e.g., neural network model 200 or artificial intelligence model or both) to onnx IR (code), a krnl IR mapping file used to convert onnx IR (code) to krnl IR (code), an affine IR mapping file used to convert krnl IR (code) to affine IR (code), and an LLVM IR mapping file used to convert affine IR (code) to llvm IR (code). As shown in FIG. 7, compiler 322 generates a list file 332 used to convert llvm IR (code) to an executable file 334, also referred to as an executable, library, binary code, assembly language, etc., and executable file 334 can be read and executed by the hardware / software of computer system 302 or computer system 320 or both. Thus, executable file 334 is run / executed by computer system 302 (or computer system 320 or both), which corresponds to running / executing neural network model 200 to perform inference, as also shown in FIG. 7. For example, neural network model 200 is designed to receive received data and classify the data as output. At runtime, computer system 302 generates a data file 340 containing various information, including instructions, instruction addresses, performance metric data, etc., and this information can be used to assist in improving neural network model 200, as further described herein.

[0039] Figure 5 is a flowchart of a computer-implemented process 500 for identifying performance hot spots within a neural network model, according to one or more embodiments of the present invention. The computer-implemented process 500 may be executed using the computer system 302 of FIG. 2. The computer-implemented process 500 of FIG. 5 is described with reference to FIG. 2.

[0040] At block 502, after the execution time of generating the data file 340, the software application 304 is configured to collect sample data 352 from the data file 340. The sample data 352 may be a part of the data file 340. In one or more embodiments, the sample data 352 may include all of the data file 340. The data file 340 is generated as a result of the execution of the previously compiled neural network model 200. The data file 340 includes instructions, instruction addresses of the instructions, performance metrics regarding each instruction or instruction address or both. Performance metrics include cache misses, information from hardware counters, information from software counters, CPU usage, memory usage, etc. The sample data 352 includes all types of information contained in the data file 340, along with the number of sample occurrences. Different computing architectures may include different numbers of counters. One architecture may include, for example, a counter number ranging from 10 to 1000. In one or more embodiments, as shown in the control flow of FIG. 7, the data file 340 can be an ETB file.

[0041] In block 504, software application 304 is configured to determine and identify performance hotspots and corresponding hot instructions within sample data 352. To identify performance hotspots, software application 304 may include, call, or execute one or more diagnostic tools 306 to identify instructions that meet, exceed, or both meet and exceed a predefined threshold regarding a metric or combination of metrics or both. A performance hotspot is an instruction, operation, or code, or combination thereof, that causes computer system 302 (or computer system 320 or both) to meet, exceed, or both meet and exceed a predefined threshold regarding a metric. For example, a performance hotspot can refer to determining that one or more instructions, operations, or code, or combination thereof, meet, exceed, or both meet and exceed one or a combination of predefined thresholds related to CPU usage (e.g., processor usage such as a percentage), execution time (e.g., ticks, clocks, etc.), cache misses, memory usage, etc. In some cases, a performance hotspot can exceed the average performance metric compared to other instructions within sample data 352, even if an instruction, operation, or code, or combination thereof, does not meet, exceed, or both meet and exceed any predefined threshold. For example, a performance hotspot can deviate from the statistical average of a performance metric or combination of performance metrics or both. Software application 304 or diagnostic tool 306 or both may parse sample data 352 and determine performance metrics that meet, exceed, or both meet and exceed one or a combination of predefined thresholds along with the associated instructions. The analysis of performance hotspots can be manual, automatic, semi - automatic, or a combination thereof.In one or more embodiments, experienced performance analysts and developers use analysis tools to pass through performance hotspots at different levels and find the parts that are most useful for optimization. Software application 304 or diagnostic tool 306 or both can employ known methods and techniques for software diagnosis and performance, as understood by those skilled in the art.

[0042] In block 506, software application 304 is configured to map the positions of LLVM IR to instructions within sample data 352 using a list file 332, where the instructions are, for example, at the instruction addresses within memory 308. The instructions are executable code such as the executable file 334 described in FIG. 4. As described in FIG. 4 and illustrated in FIG. 7, executable instructions are compiled from LLVM IR. Thus, there is a mapping relationship between LLVM IR and the instruction addresses. This relationship is represented by a file called list file 332. The list file 332 is generated when compiling LLVM IR into the executable file 334, and the executable file 334 can be executable code, binary code, machine language, etc. stored in memory 308. LLVM IR is the lowest level of code that is not binary code. The list file may include the exemplary columns of Table 1 below.

[0043]

Table 1

[0044] In block 508, the software application 304 is configured to map the positions and information of the LLVM IR to the related positions within the original neural network model 200 using the mapping file 330. The neural network model 200 (e.g., the original AI model) requires multiple conversion stages to convert the neural network model 200 into the final executable code (e.g., the executable file 334). Each conversion stage includes a mapping file, and the combination of these mapping files can be utilized to convert the LLVM IR back to the neural network model 200 (e.g., the original AI model) in the reverse direction. As shown in FIG. 7, each mapping file 330 includes the relationships and conversions from the current intermediate representation back to the previous higher-level intermediate representation, as well as the relationships and conversions from the highest-level intermediate representation back to the neural network model 200. Thus, the software application 304 can identify the positions of the LLVM IR within the original AI model, for example, by mapping the input code to the output code. For example, the software application 304 is configured to map the positions and information of the LLVM IR to the related positions of the affine IR in the affine IR (code) using the LLVM IR mapping file of the mapping file 330. Similarly, the software application 304 is configured to map the positions and information of the affine IR to the related positions of the krnl IR in the krnl IR (code) using the affine IR mapping file of the mapping file 330. Similarly, the software application 304 is configured to map the positions and information of the krnl IR to the related positions of the onnx IR in the onnx IR (code) using the krnl IR mapping file of the mapping file 330.The software application 304 is configured to map the position and information of the onnx IR to the ONNX model, which is a neural network model 200 (i.e., an AI model), in reverse using the onnx IR mapping file of the mapping file 330. As shown in the figure, the instructions are traced through the MLIR compilation stage in reverse order (e.g., the reverse order of Figure 4 and the reverse order of the conversion of Figure 7). The mapping relationships within the individual mapping files (such as the LLVM IR mapping file, the affine IR mapping file, the krnl IR mapping file, and the onnx IR mapping file) are used to trace each instruction address or operation or both in reverse through the MLIR compilation stage.

[0045] In block 510, software application 304 is configured to generate and display an MLIR performance matrix for each level with respect to hot instructions mapped to different intermediate instructions. FIG. 6 is a block diagram of an exemplary MLIR performance matrix showing performance hot spots that are interrelated at different levels of compilation, according to one or more embodiments of the present invention. The MLIR performance matrix can be displayed on a display such as display 119, thereby allowing the user to quickly visualize which operations can be optimized to modify one or more hot spots. Software application 304 generates the MLIR performance matrix using the instruction addresses of sample data 352 combined with a mapping file 330 that can include instruction addresses. The performance matrix shows performance hot spots displayed at different levels using a dot pattern so that the user or the software application 304 or both can select the best level (and operation) for optimization. For example, software application 304 can choose to optimize with LLVM IR, affine IR, krnl IR, or ONNX IR, or a combination thereof. With LLVM IR, affine IR, krnl IR, or ONNX IR, or a combination thereof, the selected optimization can include dialect operations or passes or both. LLVM IR is closest to the instructions of sample data 352, and the performance matrix progresses upward through ONNX IR. Although neural network model 200 is not shown in the performance matrix, it should be understood that higher levels including the nodes of neural network model 200 can be displayed, and some of those nodes are performance hot spots.In one or more embodiments, the software application 304 may be configured to optimize performance hotspots at any one or more of the levels shown in the performance matrix of FIG. 6, or to call an optimization tool (not shown) to optimize, or to perform both. The software application 304 may instruct the user to perform optimization at one or more levels. Some users may be accustomed to making changes at specific levels or may find it enjoyable to do so. Also, some levels may be specialized without causing unexpected problems and may modify performance hotspots or characteristics of performance hotspots or both. For example, optimizing or modifying smaller or less complex operations at an intermediate level, or both (while, for example, not optimizing other performance hotspots at the same level), may improve performance with the potential to cause unexpected problems. However, optimizing nodes in a neural network model may cause problems when the neural network model is compiled. For example, after MLIR compilation, there are thousands of hot instructions and hot basic blocks distributed among millions of instructions. When aggregating these instructions at a higher level, there may be only a few hot operations such as "std.add" and "matrix bias add". Therefore, the user needs to optimize only these hot operations through software tuning or hardware acceleration in order to obtain the maximum performance improvement with minimal optimization effort.

[0046] In Figure 6, the values of operations at various levels of compilation can be, for example, CPU usage, execution time, cache misses, etc., and the larger the value, the worse. In this example, the performance metric threshold can be, for example, 20%, such that each operation with a value of 20% or more is a performance hot spot. By displaying the values along with the names of the operations at each level (e.g., LLVM IR, affine IR, krnl IR, or ONNX IR, or combinations thereof), the user can easily and quickly visualize where improvements and optimizations can be made. For example, based on the mapping relationship between the instruction addresses and names provided by the mapping file 330, the ranges of individual instruction addresses 0x10f to 0x120, 0x204 to x228 belong to the "llvm.add" node in Figure 6. Next, samples at addresses 0x112, 0x120, 0x204, 0x208, etc. are aggregated at the llvm.add node. The aggregated number of samples for llvm.add is 40% of the total number of samples. Based on using the mapping file 330 to map nodes along the affine path to the llvm node, the llvm.add node is converted to the std.addf node. Thus, all the sample numbers for llvm.add are aggregated at the std.addf node. The std.addf node is the hottest node along the affine path as seen in Figure 6.

[0047] FIG. 7 is a block diagram of a computer-implemented control flow for identifying performance hot spots within a neural network model for optimization, according to one or more embodiments of the present invention. In one or more embodiments, the computer-implemented control flow may be executed by a computer system 302. In one or more embodiments, a portion of the computer-implemented control flow, such as the compilation and execution of the neural network model 200, may be executed by a computer system 320, and a portion of the computer-implemented control flow, such as determining and identifying performance hot spots, may be executed by a computer system 302. In the compilation process, an IA mapping file (such as, for example, mapping file 330) and a list file (such as, for example, list file 332) are generated. In the execution process, the computer system 302 uses a sampling tool to collect a data file, etc. (such as, for example, data file 340). After being executed by the computer system 302, the computer system 302 uses the IA mapping file, the list file, and the data file to identify performance hot spots.

[0048] FIG. 8 is a flowchart of a computer-implemented process 800 for identifying performance hot spots within a neural network model and optimizing one or more performance hot spots, according to one or more embodiments of the present invention. The computer-implemented process 800 of FIG. 8 can be executed by the system 300 of FIG. 3 and is described with reference to FIG. 3.

[0049] At block 802, software application 304 on computer system 302 is configured to collect sample data (e.g., sample data 352) having an instruction address (IA), and the sample data is related to neural network model 200. The sample data 352 can be from data file 340 generated during the execution of neural network model 200, specifically during the execution of the compiled neural network model 200.

[0050] At block 804, software application 304 on computer system 302 is configured to determine the instructions within the instruction address that are performance hotspots. For example, software application 304 may include the functionality of one or more known diagnostic tools 306 configured to determine performance hotspots based on performance metrics, or call / instruct one or more known diagnostic tools 306, or do both.

[0051] At block 806, software application 304 on computer system 302 is configured to use list file 332 to map the instructions of sample data 352 that are performance hotspots to positions within a lower-level intermediate representation. The lower-level intermediate representation can be LLVM IR (code).

[0052] At block 808, software application 304 on computer system 302 is configured to use a mapping file (e.g., mapping file 330) to map the locations of lower-level intermediate representations that are performance hotspots to operations in one or more higher-level (intermediate) representations (e.g., affine IR, krnl IR, ONNX IR, etc.), where one or more of these operations correspond to performance hotspots, and mapping file 330 is generated from compiling neural network model 200.

[0053] Software application 304 on computer system 302 is configured to use mapping file 330 to map operations within one or more higher-level representations to nodes within neural network model 200. Software application 304 on computer system 302 is configured to determine which nodes within neural network model 200 represent performance hotspots based on mapping from one or more higher-level (intermediate) representations mapped from a lower-level intermediate representation mapped from sample data 352.

[0054] The software application 304 on the computer system 302 is configured to classify performance hotspots at different levels of compilation of the neural network model 200 using the mapping file 330, where the different levels include from a lower-level intermediate representation to one or more higher-level representations. The different levels can include LLVM IR, affine IR, krnl IR, or ONNX IR, or combinations thereof. The sample data (e.g., sample data 352) further includes information from one or more counters. Information regarding the execution of the neural network model 200 from hardware counters or software counters or both can be stored in the data file 340 from which the sample data is collected. The performance hotspots include one or more (performance) metrics that meet or exceed or both one or more thresholds (e.g., thresholds for performance metrics). The software application 304 on the computer system 302 is configured to determine one or more operations within one or more higher-level representations for optimization and to address the performance hotspots.

[0055] Although this disclosure includes a detailed description regarding cloud computing, it should be understood that the implementations of the content shown herein are not limited to a cloud computing environment. Embodiments of the present invention can be implemented in combination with any other type of computing environment that is currently known or developed in the future.

[0056] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or service provider interaction. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0057] The characteristics are as follows.

[0058] On-demand self-service: Cloud users can provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider.

[0059] Broad network access: Capabilities are available over the network and can be accessed using standard mechanisms, facilitating use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0060] Resource pooling: Provider computing resources are pooled and provided to multiple users using a multi-tenant model. Various physical and virtual resources are dynamically assigned and reassigned according to demand. There is a sense of location independence, and users typically neither manage nor know the exact location of the resources provided, although at a higher level of abstraction, it may be possible to specify a location (e.g., country, state, or data center).

[0061] Rapid adaptability: The capabilities can be provisioned quickly, flexibly, and in some cases automatically, scale out rapidly, and be released quickly to scale in. The capabilities available for provisioning often appear to the user as if they can purchase any amount, without limit, at any time.

[0062] Measured services: Cloud systems automatically control and optimize the use of resources at an abstraction level suitable for the type of service (e.g., storage, processing, bandwidth, and active user accounts) by leveraging metering capabilities. The usage of resources can be monitored, controlled, and reported, providing transparency to both the provider and the user of the services being utilized.

[0063] The service model is as follows.

[0064] SaaS (Software as a Service): The capabilities provided to the user are the use of the provider's applications running on the cloud infrastructure. Those applications can be accessed from various client devices via a thin-client interface such as a web browser (e.g., web-based email). The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for limited user-specific application configuration settings.

[0065] PaaS (Platform as a Service): The capabilities provided to users are to deploy applications created or acquired by users, which are created using programming languages and tools supported by the provider, onto the cloud infrastructure. Users do not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, or storage, but can control the deployed applications and, in some cases, the configuration of the application hosting environment.

[0066] IaaS (Infrastructure as a Service): The capabilities provided to users are the provisioning of processing, storage, network, and other basic computing resources, and users can deploy and run any software that can include operating systems and applications. Users do not manage or control the underlying cloud infrastructure, but can control the operating systems, storage, deployed applications, and, in some cases, have limited control over selected network components (such as host firewalls).

[0067] The deployment models are as follows.

[0068] Private cloud: This cloud infrastructure is operated only for an organization. It can be managed by this organization or a third party and can exist on-premises or off-premises.

[0069] Community Cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community that shares concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by these organizations or a third party and can exist on-premises or off-premises.

[0070] Public Cloud: This cloud infrastructure is available for general users or large industry groups and is owned by an organization that sells cloud services.

[0071] Hybrid Cloud: This cloud infrastructure is a composite of two or more clouds (private, community, or public) that are joined together while leaving their distinct entities intact by means of standardized or proprietary technologies (e.g., cloud bursting to balance the load between clouds) that enable the migration of data and applications.

[0072] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the center of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0073] Referring now to FIG. 9, an exemplary cloud computing environment 50 is shown. As illustrated, cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers, such as, for example, personal digital assistants (PDAs) or cellular telephones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof, may communicate. Nodes 10 may communicate with one another. Nodes 10 may be physically or virtually grouped in one or more networks into private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described hereinabove (not shown). Thereby, cloud computing environment 50 may provide infrastructure, platforms, or SaaS, or combinations thereof, that cloud consumers do not need to maintain on local computing devices. The types of computing devices 54A - N shown in FIG. 9 are intended only as examples, and it will be understood that cloud computing nodes 10 and cloud computing environment 50 may communicate with any type of computer controlled device via any type of network or network addressable connection (e.g., connection using a web browser) or both.

[0074] Referring now to FIG. 10, a set of functional abstractions provided by cloud computing environment 50 (FIG. 9) is shown. It should be understood upfront that the components, layers, and functions shown in FIG. 10 are intended only as examples and that embodiments of the invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.

[0075] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.

[0076] The virtualization layer 70 includes an abstraction layer that can provide virtual entities such as virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual clients 75.

[0077] For example, the management layer 80 may provide the functions described below. Resource provisioning 81 dynamically procures computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 tracks the costs when resources are utilized within a cloud computing environment and sends bills or invoices for the use of those resources. For example, those resources may include application software licenses. Security verifies the identities of cloud users and tasks and protects data and other resources. The user portal 83 provides access to the cloud computing environment to users and system administrators. Service level management 84 allocates and manages the cloud's computing resources to meet the required service levels. Service level agreement (SLA) planning and execution 85 makes advance preparations and procurements of the cloud's computing resources for which future demands are anticipated, in accordance with the SLA.

[0078] The workload layer 90 shows examples of functions available in a cloud computing environment. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, delivery of virtual classroom education 93, data analysis processing 94, transaction processing 95, and software applications implemented in the workload and functions (e.g., software application 304, compiler 322, neural network model 200, etc.) 96. Also, the software application can function with resource provisioning 81, be integrated with resource provisioning 81, or both.

[0079] In this specification, various embodiments of the present invention are described with reference to the related drawings. Alternative embodiments of the present invention can be devised without departing from the scope of the present invention. In the following description and drawings, various connection and positional relationships between elements (e.g., above, below, adjacent, etc.) are shown. Those connections or positional relationships or both can be direct or indirect, and the present invention is not intended to be limited in this regard. Therefore, the coupling of each entity can refer to direct coupling or indirect coupling, and the positional relationship between each entity can be a direct positional relationship or an indirect positional relationship. Further, the various operations and process steps described herein can be incorporated into a more comprehensive procedure or process having additional steps or functions not described in detail herein.

[0080] One or more of the methods described herein can be implemented using any or a combination of techniques well known in the prior art, such as individual logic circuits having logic gates for implementing logical functions on data signals, application specific integrated circuits (ASICs) having appropriate combinational logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0081] For the purposes of brevity, the prior art relevant to the creation and use of aspects of the present invention may or may not be described in detail herein. Specifically, computing systems and particular aspects of specific computer programs for implementing the various technical features described herein are well known. Thus, for the sake of brevity, many details regarding conventional implementations are only briefly described herein or are omitted entirely without providing details of known systems or processes or both.

[0082] In some embodiments, the various functions or operations may be performed at a particular location, or in relation to the operation of one or more devices or systems, or both. In some embodiments, a portion of a particular function or operation can be performed at a first device or location, and the remaining portion of the function or operation can be performed at one or more additional devices or locations.

[0083] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprising", "comprises", and / or "comprising" when used herein, specify the presence of the stated function, integer, step, operation, element, or component, or a combination thereof, but do not preclude the presence or addition of one or more other functions, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.

[0084] All means or steps and corresponding structures, materials, acts, and equivalents within the scope of the following claims are intended to include any structure, material, or act for performing functions in combination with other claimed elements specifically claimed. This disclosure is presented for purposes of illustration and description but is not intended to be exhaustive or limited to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of this disclosure. Embodiments have been chosen and described in order to best explain the principles of this disclosure and its practical application, and to enable others of ordinary skill in the art to understand this disclosure with respect to various embodiments with various modifications as are suited to the particular use contemplated.

[0085] The figures shown in this specification are illustrative. Without departing from this disclosure, many modifications of the figures or steps (or acts) described herein are possible. For example, acts may be performed in a different order or acts may be added, deleted, or changed. Also, the term "coupled" represents that there is a signal path between two elements and does not mean a direct connection between elements without an element / connection intervening therebetween. All such modifications are considered to be part of this disclosure.

[0086] The following definitions and abbreviations are used in the interpretation of the claims and this specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," "containing," or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent to such composition, mixture, process, method, article, or apparatus.

[0087] Furthermore, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Embodiments or designs described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer greater than or equal to one, i.e., 1, 2, 3, 4, etc. The term "a plurality" is understood to include any integer greater than or equal to two, i.e., 2, 3, 4, 5, etc. The term "connected" can include both indirect "connection" and direct "connection."

[0088] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with the measurement of a particular quantity based on the skills available at the time of filing of the present application. For example, "about" can include a range of ±8% or 5%, or 2% of a particular value.

[0089] The present invention may be a system, method, or computer program product, or a combination thereof, at any possible technical detail level of integration. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0090] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction-executing device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable floppy (registered trademark) disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (registered trademark) disk, a mechanically encoded device such as a punched card or raised structures in grooves in which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a signal per se that is transient, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted via a wire.

[0091] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers them for storage on a computer-readable storage medium within each computing / processing device.

[0092] Computer-readable program instructions for carrying out the operation of the present invention may be source code or object code described in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk (registered trademark), C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially executed on the user's computer as a stand-alone software package, partially executed on the user's computer and a remote computer respectively, or executed entirely on a remote computer or a server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to carry out aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions for customizing the electronic circuit by utilizing the state information of the computer-readable program instructions.

[0093] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0094] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium that includes instructions for causing a computer, programmable data processing apparatus, or other device to function in a particular manner so that the product comprises an article of manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0095] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0096] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently depending on the functionality involved, or may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams or flowchart diagrams, or combinations of blocks in the block diagrams or flowchart diagrams or both, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or a combination of dedicated hardware and computer instructions.

[0097] The description of the various embodiments of the present invention has been presented for purposes of illustration but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen in order to best explain the principles of the embodiments, the practical application, or a technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments described herein.

Claims

1. Collecting sample data having an instruction address, wherein the sample data relates to a neural network model, the collecting; Determining an instruction within the instruction address that is a performance hot spot; Using a list file to map the instruction of the sample data that is a performance hot spot to a location within a lower-level intermediate representation; Using a mapping file to map the location of the lower-level intermediate representation that is a performance hot spot to an operation within one or more higher-level representations, wherein one or more of the operations correspond to the performance hot spot and the mapping file is generated from compiling the neural network model, the mapping; A method executed by a computer, comprising.

2. The method of claim 1, further comprising using the mapping file to map the operation within the one or more higher-level representations to a node within the neural network model.

3. The method of claim 1 or 2, further comprising determining which node within the neural network model represents the performance hot spot based on mapping from the one or more higher-level representations mapped from the lower-level intermediate representation mapped from the sample data.

4. The method according to any one of claims 1 to 3, further comprising using the mapping file to classify the performance hot spot at different levels of compilation of the neural network model, the different levels including from the lower-level intermediate representation to the one or more higher-level representations.

5. The method according to any one of claims 1 to 4, wherein the sample data further includes information from one or more counters.

6. The method according to any one of claims 1 to 5, wherein the performance hot spot includes one or more metrics that meet or exceed one or more thresholds.

7. The method according to any one of claims 1 to 6, further comprising determining one or more of the operations within the one or more higher levels of representation that optimize the performance hot spot and thereby address the performance hot spot.

8. A memory having computer-readable instructions, One or more processors for executing the computer-readable instructions A system comprising: the computer-readable instructions controlling the one or more processors to Collect sample data having instruction addresses, wherein the sample data relates to a neural network model, the collecting; Determine instructions within the instruction addresses that are performance hot spots; Using a list file, map the instructions of the sample data that are performance hot spots to positions within a lower level of intermediate representation; Using a mapping file, map the positions of the lower level of intermediate representation to operations within one or more higher levels of representation, wherein the mapping file is generated from compiling the neural network model, the mapping; A system that executes a process including.

9. To a processor, Collecting sample data having a command address, wherein the sample data relates to a neural network model, the collecting; Determining an instruction within the command address that is a performance hot spot; Using a list file to map the instruction of the sample data that is a performance hot spot to a location within a lower-level intermediate representation; Using a mapping file to map the location of the lower-level intermediate representation to an operation within one or more higher-level representations, wherein the mapping file is generated from compiling the neural network model, the mapping; A computer program for causing a process including the above to be executed. A computer-readable storage medium recording the computer program according to claim 9.

Citation Information

Patent Citations

  • Logic diagram debug processing system

    JP1988223928A

  • Method and device for compilation, executing method, and program executing device

    JP2000047879A

  • Source code profiling through enhanced mapping

    US20190042395A1