System and method for automatic parallelization of processing code for a multiprocessor system having optimized latency
The compiler system optimizes multiprocessor systems by dividing code into computational block nodes and minimizing latency through matrix optimization, enhancing parallelization efficiency and throughput in complex computing environments.
Patent Information
- Application Number
- JP2024520843
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-27
AI Technical Summary
Existing compiler systems struggle to efficiently manage multiprocessor systems by optimizing latency and data dependencies for higher throughput, particularly in computationally rich applications, due to complex program analysis and hardware complexity, leading to suboptimal parallelization and increased latency times.
A compiler system that translates source code into machine code for multiprocessor systems, utilizing a parser module to divide code into computational block nodes, a matrix builder to generate computational chains, and an optimizer module to minimize latency times through numerical matrix optimization techniques, enabling efficient parallelization across multiple processing units.
The system achieves optimized latency and improved throughput by automatically parallelizing code across multiple processing units, adapting to various hardware infrastructures, and addressing latency issues in complex computing environments.
Smart Images

Figure 0007708369000008 
Figure 0007708369000009 
Figure 0007708369000010
Abstract
Description
Technical Field
[0001] The present invention generally relates to a multi-processor system, a multi-core system, and a parallel computing system that allow multi-processing in which two or more processors cooperate to simultaneously execute a plurality of program codes and / or processor instruction sets. The multi-processor system or parallel computing system has several central processing units linked to each other to enable the multi-processing or parallel processing. In particular, the present invention relates to a compiler system that optimizes a computer program for execution by a parallel processing system having a plurality of processing units. More specifically, the described embodiments relate to systems and methods for automatic parallelization of code for execution in the described multi-processor system. In the technical field of multi-processor systems, important characteristics and classifications arise, among other things, from how processor-memory access is handled and whether the system processors are of a single type or various types in the system architecture.
Background Art
[0002] The increasing need for compute-intensive applications is being driven, among other things, by the emergence of Industry 4.0 technologies, the ongoing automation of traditional manufacturing with an ever-growing number of industrial implementations using modern smart technologies, large-scale machine-to-machine communication (M2M), the Internet of Things (IoT) with improved communication and monitoring means having large-scale data access and aggregation, and the increasing importance of industrial simulation technologies such as, for example, digital twin technology, thereby causing a paradigm shift from centralized computing to parallel and distributed computing. Parallel computing involves distributing computing jobs across various computing resources. These resources generally include several central processing units (CPUs), graphics processing units (GPUs), memory, storage, and networking support. In addition to this growing demand, over the past 50 years, there have been significant developments in the performance and capabilities of computer systems at the hardware level. This has been made possible with the help of very large scale integration (VLSI) technology. VLSI technology enables the accommodation of a large number of components on a single chip and increases the clock speed. Thus, more operations can be executed in parallel at once. Parallel processing is also related to data locality and data communication. Therefore, the field of parallel computer architecture typically refers to systems and methods that orchestrate all the resources of a parallel computer system to maximize performance and programmability within the limits given by technology and cost at any given point in time. Parallel computer architecture adds dimensions to the development of computer systems by using an ever-increasing number of processors. In principle, the performance achieved by utilizing a large number of processors is higher than that of a single processor at a given point in time. However, parallelization of processor code is complex and difficult to automate for truly optimized parallelization.
[0003] Centrally centralized computing functions well in many applications, but may be insufficient in computationally rich applications and in the execution of large amounts of data processing. Programs may be executed serially or distributed to be executed on multiple processors. When a program is executed serially, only one processor can be utilized, and thus the throughput is limited by the speed of the processor. Such a system with one processor is sufficient for many applications, but not sufficient for computationally intensive applications such as modern computer-based simulation technologies. The processing code can be executed in parallel in a multiprocessor system, which can bring higher throughput. A multiprocessor system needs to divide the code into smaller code blocks and efficiently manage the execution of the code. For processors to execute in parallel, the data to each processor must be independent. To improve throughput, instances of the same code block can be executed simultaneously on several processors. When a processor needs data from a previous execution or another process currently performing a calculation, the parallel processing efficiency may decrease due to the latency caused by data exchange between processor units and / or within processor units. Generally, when a processor enters a state where the execution of a program is interrupted or not executed for some reason, and the instructions belonging to it are not fetched or executed from memory, those states cause the processor to be in an idle state and affect the parallel processing efficiency. When scheduling processors, it is necessary to consider data dependencies. It is difficult to efficiently manage multiple processors and data dependencies for higher throughput. It is desirable to have a method and system for efficiently managing code blocks in computationally rich applications. Note that the latency problem also exists in a single-processor system. In that case, for example, a latency-oriented processor architecture is used to minimize the problem.It is a microarchitecture of a microprocessor designed to provide serial computing threads with low latency. These architectures generally aim to execute as many instructions as possible belonging to a single serial thread within a given time window, and the time to fully execute a single instruction from the fetch stage to the retire stage can vary from several cycles to hundreds of cycles in some cases. However, these techniques are not automatically applicable to the latency problems of (large-scale) parallel computing systems.
[0004] In general, parallel computing systems require efficient parallel coding or programming, and parallel programming becomes a programming paradigm. This includes, on the one hand, ways to divide a computer program into individual sections that can be executed concurrently, and on the other hand, ways to synchronize the concurrently executed code sections. This is in contrast to classical sequential (or serial) programming and coding. Parallel execution of a program can be supported on the hardware side, in which case the programming language typically conforms to this. For example, parallel programming can be done explicitly by having the programmer execute parts of the program in separate processes or threads, or it can be done automatically such that causally independent (parallelizable) sequences of instructions are lined up and thus executed in parallel. This parallelization can be done automatically by a compiler system if a computer with a multi-core processor or a parallel computer is available as the target platform. Some modern CPUs can recognize such independence (in the machine code or microcode of the program) and distribute instructions to different parts of the processor so that the instructions are executed simultaneously (out-of-order execution). However, as soon as individual processes or threads communicate with each other, they affect each other, and in that sense, they are no longer concurrently parallel as a whole. Only the individual sub-processes are still concurrently parallel with each other. If the execution order of the communication points of the individual processes or threads cannot be properly defined, collisions can occur. In particular, the so-called deadlock when two processes wait for each other (or block each other), or a race condition when two processes overwrite each other's results. In the prior art, synchronization techniques such as, for example, mutual exclusion (mutex) techniques are used to solve this problem. Such techniques can prevent race conditions, but do not allow optimized parallel processing of processes or threads with the minimum latency of the processor unit.
[0005] (Micro)processors are based on integrated circuits that enable arithmetic and logical operations to be performed based on two binary values (the simplest 1 / 0). For this purpose, the binary values must be available to the processor's computational unit. The processor unit needs to obtain two binary values in order to calculate the result of the expression a = b operand c. The time required to obtain data for these operations is known as latency. There is a wide hierarchical range of these latency times, from registers, L1 caches, memory access, I / O operations, or network transfers, as well as the processor configuration (e.g., CPU or GPU). Since every single component has a latency time, the overall latency time for calculations is, in modern computing infrastructure, mainly a combination of the hardware components required to obtain data from one location to another. In modern architectures, different software layers (e.g., of the operating system) also have a major impact. The difference between the fastest and the slowest locations for the CPU (or GPU) to obtain data can be huge (>10 9 times the range). Figure 1 shows the formation of latency times in modern computing infrastructure. As shown in Figure 1, parallel computing machines have been developed using different distinct architectures. It is important to note that parallel architectures improve the traditional concept of computer architecture using a communication architecture. Computer architecture defines key abstractions (such as the user-system boundary and the hardware-software boundary) and organizational structures, while communication architecture defines basic communication and synchronization operations. It also deals with organizational structures.
[0006] Computer applications are typically written at the highest level, i.e., in a high-level language, based on the corresponding programming model. For example, various parallel programming models are known, such as (i) shared address space, (ii) message passing, or (iii) data parallel programming that refers to the corresponding multiprocessor system architecture. Shared memory multiprocessors are such a class of parallel machines. A shared memory multiprocessor system provides better throughput for multiprogramming workloads and supports parallel programs. In this case, the computer system allows a set of processors and I / O controllers to access a collection of memory modules through some hardware interconnection. The memory capacity can be increased by adding memory modules, and the I / O capacity can be increased by adding devices to the I / O controller or by adding additional I / O controllers. The processing power can be increased by implementing faster processors or by adding more processors. As shown in Figure 2, the resources are organized around a central memory bus. Through the bus access mechanism, any processor can access any physical address within the system. Since all processors are assumed to be equidistant from all memory locations or are actually equidistant, the access time or latency for all processors is the same for memory locations. This is called a symmetric multiprocessor system.
[0007] Message passing architecture is another class of parallel machines and programming models. It provides communication between processors as an explicit I / O operation. The communication is combined at the I / O level instead of the memory system. In message passing architecture, user communication is performed by using operating system or library calls that execute lower-level actions including the actual communication operation. As a result, there is a gap between the programming model and the communication operation at the physical hardware level. Sending and receiving are the most common user-level communication operations in a message passing system. Sending specifies a local data buffer (to be sent) and a receiving-side remote processor. Receiving specifies the sending process and a local data buffer where the sent data is to be placed. In the sending operation, an identifier or tag is attached to the message, and the receiving operation specifies matching rules such as a specific tag from a specific processor or any tag from any processor. The combination of a send and a matching receive completes the memory-to-memory copy. Each end specifies its local data address and a pairwise synchronization event. Message passing and shared address space have traditionally represented two separate programming models, each with its own paradigm for sharing, synchronization, and communication, but today, the basic machine structures are converging towards a common fabric.
[0008] Finally, data parallel processing is a further class of parallel machines and programming models, also known as processor arrays, data parallel architectures, or single instruction multiple data machines. The main feature of this programming model is that operations can be performed in parallel on each element of a large regular data structure (such as an array or matrix). Data parallel programming languages are typically implemented by looking at the local address spaces of a group of processes (one per processor) that form an explicit global space. Since all processors communicate together and there is a global view of all operations, either a shared address space or message passing can be used. However, improving computer efficiency cannot be achieved by developing the programming model alone, nor can it be achieved by developing the hardware alone. Furthermore, the top-level programming model necessarily introduces boundary conditions, such as model-specific architectures, given by the requirements of the programming model. Since parallel programs consist of one or more threads that operate on data, the underlying parallel programming model defines what data the threads need, what operations can be performed on the required data, and in what order the operations follow. Thus, there are limitations to the optimization of machine code for multiprocessor systems due to the boundaries of the underlying programming model. Parallel programs must always coordinate the activities of their threads to ensure that dependencies between programs are enforced.
[0009] As shown in FIG. 1, parallel computing machines are developed using different distinct architectures, each causing different latency times in their computing infrastructure. One of the most common multiprocessor systems is the shared memory multiprocessor system. Essentially, for shared memory multiprocessor systems, three basic architectures are known: (i) Uniform Memory Access (UMA), (ii) Non-uniform Memory Access (NUMA), and (iii) Cache Only Memory Architecture (COMA). In the UMA architecture (see FIG. 3), all processors uniformly share the physical memory. All processors have equal access times to all memory words. Each processor may have a private cache memory. Peripheral devices also follow the same rule. If all processors have equal access to all peripheral devices, the system is called a symmetric multiprocessor. If only one or a few processors can access the peripheral devices, the system is called an asymmetric multiprocessor. In the NUMA multiprocessor architecture (see FIG. 4), the access time varies depending on the position of the memory word. The shared memory is physically distributed among all processors, called local memory. The set of all local memories forms a global address space that can be accessed by all processors. Finally, the COMA multiprocessor architecture (see FIG. 5) is a special case of the NUMA multiprocessor architecture. In the COMA multiprocessor architecture, all distributed main memories are converted into cache memories. The COMA architecture can also be applied to distributed memory multicomputers. A distributed memory multicomputer system consists of a plurality of computers, typically called nodes, interconnected by a message passing network.Each node functions as an autonomous computer having a processor, local memory, and sometimes I / O devices. In this case, all local memory is private and accessible only to the local processor, and thus such a machine is also called a no-remote-memory-access (NORMA) machine. Other known multiprocessor architectures are, for example, multi-vector computers, and single instruction multiple data (SIMD) parallel computers, parallel random access machines (PRAM), and parallel computers based on very large scale integration (VLSI) chips, etc., all having different multiprocessor architectures and infrastructure characteristics. In summary, different multiprocessor architectures result in different latency times in their computational infrastructure, so that improving the efficiency of a computer cannot be achieved by developing only the programming model, nor can it be done by developing only the hardware.
[0010] As described above, when a computer having a multi-core processor or a parallel computer is available as a target platform, code parallelization can also be automatically performed by a compiler system. Such automatic parallelization, also called auto-parallelization, means converting sequential code into multi-threaded and / or vectorized code in order to use multiple processors simultaneously in, for example, a shared-memory multiprocessor (SMP) machine. In prior art systems, fully automatic parallelization of sequential programs is technically difficult because it requires complex program analysis and may depend on parameter values that are not known at compile time and for which the best approach is not known. The programming control structure on which compiler systems focus most for automatic parallelization is the loop, because typically most of the execution time of a program takes place inside some form of loop. There are two main techniques for loop parallelization: pipelined multi-threading and cyclic multi-threading. Compiler structures for automatic parallelization typically include a parser, an analyzer, a scheduler, and a code generator. The parser of the compiler system covers the first processing stage where, for example, a scanner reads the input source file and identifies all static and external uses. Each line in the file is checked against a predefined pattern and separated into tokens. These tokens are stored in a file for later use by the grammar engine. The grammar engine checks for patterns of tokens that match predefined rules to identify variables, loops, control statements, functions, etc. in the code. In the second stage, the analyzer identifies sections of code that can be executed concurrently. The analyzer uses the static data information provided by the scanner-parser. The analyzer first detects all fully independent functions and marks them as individual tasks. Then, the analyzer finds out which tasks have dependencies.In the third stage, the scheduler lists all tasks and their dependencies with respect to execution time and start time. The scheduler generates an optimal schedule with respect to the number of processors used or the total execution time for the application. In the fourth and final stage, the scheduler generates a list of all tasks and details of the cores on which they are executed, along with the length of time for which they are executed. The code generator then inserts special structures into the code read during execution by the scheduler. These structures instruct the scheduler as to which core a particular task is executed on, along with the start time and end time.
[0011] When a loop multi-threaded parallelizing compiler is used, the compiler attempts to split each loop so that each iteration of the loop can be executed concurrently on separate processors. During auto-parallelization, the compiler typically performs two automated passes of evaluation prior to actual parallelization to determine the following two basic preconditions for parallelization. (i) In the first pass, based on dependency analysis and alias analysis, it is evaluated whether it is safe to parallelize the loop. (ii) In the second pass, based on an estimate (modeling) of the program workload and the capacity of the parallel system, it is evaluated whether it is worthwhile to parallelize it. The first pass of the compiler performs a data dependency analysis of the loop to determine whether each iteration of the loop can be executed independently of other iterations. Data dependencies can sometimes be addressed, but can introduce additional overhead in the form of message passing, shared memory synchronization, or some other method of processor communication. The second pass attempts to justify the effort of parallelization by comparing the theoretical execution time of the parallelized code to the sequential execution time of the code. It is important to understand that the code does not necessarily benefit from parallel execution. The extra overhead associated with using multiple processors can eat into the potential speedup of the parallelized code.
[0012] When a pipelined multi-threaded parallelizing compiler is used for automatic parallelization, the compiler attempts to split the sequence of operations within a loop into a series of code blocks such that each code block can be executed concurrently on a separate processor.
[0013] In a specific system using pipes and filters, there are many parallel problems with such relatively independent code blocks. For example, when creating a live broadcast, many different tasks must be performed many times per second.
[0014] A pipelined multi-threaded parallelizing compiler attempts to assign each of these operations to different processors, usually arranged in a systolic array, and insert appropriate code to transfer the output of one processor to the next. For example, in modern computer systems, one focus is to use the capabilities of GPUs and multi-core systems to compute such independent code blocks (or independent iteration steps of a loop) at runtime. Then, the memory that is (directly or indirectly) accessed is marked for different iteration steps of the loop and can be compared for dependency detection. Using this information, the iteration steps are grouped into levels such that iteration steps belonging to the same level are independent of each other and can be executed in parallel.
[0015] In the prior art, there are many compilers for automatic parallelization. However, the latest prior art compilers for automatic parallelization rely on the use of Fortran as a high-level language, that is, they are only applicable to Fortran programs. This is because Fortran provides stronger guarantees regarding aliasing than languages such as C. Typical examples of such prior art compilers are (i) the Paradigm compiler, (ii) the Polaris compiler, (iii) the Rice Fortran D compiler, (iv) the SUIF compiler, and (v) the Vienna Fortran compiler. A further drawback of automatic parallelization by prior art compilers is that it is often difficult to achieve a high degree of code optimization. This is because (a) for code using indirect addressing, pointers, recursion, or indirect function calls, it is difficult to detect such dependencies at compile time, making dependency analysis difficult, (b) loops often have an unknown number of iterations, (c) it is difficult to coordinate access to global resources with respect to memory allocation, I / O, and shared variables, and (d) due to the fact that irregular algorithms using input-dependent indirection interfere with compile-time analysis and optimization.
[0016] One important task of a compiler is to efficiently handle latency times. Compilation is the transcription from a so-called high-level language (such as C, Python, Java (registered trademark), etc.) that can be read by humans to assembler / processor code, and the assembler / processor code consists only of instructions available on a given processor. As already mentioned, modern applications with a large demand for data or calculations must target an appropriate infrastructure, introducing many different latency times, and currently only some of them can be solved by prior art compiler optimization techniques.
[0017] For every level of complexity (hardware component), solutions have been historically developed and evolved, from compiler optimization techniques, to multithreaded libraries for concurrent data structures to prevent race conditions, vectorization of code, GPU systems with corresponding programming languages (e.g., OpenCL (Open Computing Language)), frameworks such as "TensorFlow" for programmers to distribute computations, to big data algorithms such as "MapReduce". Here, MapReduce is a programming technique and related implementation for processing and generating big data sets using parallel distributed algorithms on a cluster. In the field of high-performance computing, theoretical mathematical-based techniques have been developed and defined to split large matrices into special gridding techniques for the finite difference method or the element method. This includes, for example, protocols in cluster infrastructure. The Message Passing Interface (MPI) supports the transfer of data to different processes through the infrastructure.
[0018] As described above, the list of prior art optimization techniques is long. However, from a systems theory perspective, the problems are more or less always the same, namely, how can code (any code) interact most efficiently with latency time in a complex hardware infrastructure. The compiler functions well when using a single CPU. As soon as the hardware complexity increases, the compiler is no longer able to actually parallelize the code. Parallelization simply becomes CPU emulation, for example, by introducing microprocessors. The hardware industry for CPUs, GPUs, and their clusters mainly focuses on their specific domains, and developers and research mainly focus on implementation techniques and framework development, and so far have not migrated to the field of more general (cross-industry) approaches. Furthermore, Non-Patent Document 1 revealed splitting an algorithm for processing a data stream among different processor cores. The author's model is a simple communication network with a simple additive communication model between processor cores. This does not allow realistic conclusions about the actual communication load caused by splitting across multiple cores.
[0019] Generally, known processor manufacturers focus on their processors and related hardware components, while other developers, such as research groups in high-performance computing (HPC), focus on the use of numerical methods and libraries. Currently, there is no attempt to solve problems regarding compiler system optimization from a systems theory perspective by accessing latency dynamics resulting from a given source code. Prior art source code simply consists of a series of statements that result in direct reads and writes for a given target infrastructure.
Prior Art Documents
Non-Patent Documents
[0020]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0021] One object of the present invention is to compile program code into machine code having an optimized latency for the processing units of a multiprocessor system, thereby efficiently managing the multiprocessor and data dependencies for higher throughput and not having the drawbacks of the prior art systems as described above. A compiler system for multiprocessor systems and multicomputer systems is provided. In particular, an object of the present invention is to achieve the highest performance in a multiprocessor machine through automatic parallelization, and thereby provide a system and technique that can be used to optimize the utilization of low-level parallelism (temporal and spatial) at the level of machine instruction processing. A further object of the present invention is to overcome the drawbacks of the prior art and typically overcome the limitations of dealing with parallelized sections, such as loops or specific sections of code, which are restricted to a particular system that is assumed.
[0022] An automatic parallelization system should be able to optimize the identification of parallelization opportunities as a crucial step while generating a multi-threaded application. The prior art document M. Kandemir et al., "Slicing Based Code Parallelization for Minimizing Inter-processor Communication", 2009 International conference on compilers, architecture, and synthesis for embedded systems (cases '09), Grenoble, France October 11-16, 2009, p. 87-96 discloses a system for automatic parallelization aimed at minimizing inter-processor communication in a distributed memory multi-core architecture by applying the concept of sequential iteration space slicing. That is, this prior art system is based on a sequential iterative approach. The disclosed system does this by determining the partitioning of the output array by sequentially iteratively determining the partitioning of other arrays in the application code. That is, it is by sequential iterative determination of the parts of the array. The information is taken sequentially iteratively from the previously sliced array parts. In code parallelization, slicing represents the process of extracting from a program the statements that may affect a particular statement of interest that is the slicing criterion (see, for example, J. Krinke, "Advanced slicing of sequential and concurrent programs," 20th IEEE International Conference on Software Maintenance, 2004. Proceedings., 2004, pp. 464-468).These slicing techniques exhibit similar effects as point data / control dependencies (see, e.g., J.L. Hennessy, D.A. Patterson "Computer Architecture", Fifth Edition: A Quantitative Approach, The Morgan Kaufmann Series in Computer Architecture and Design, fifth edition, p. 150ff) and data flow analysis (see, e.g., Gary A. Kildall "A unified approach to global program optimization", proceedings of the 1st annual ACM SIGACT-SIGPLAN symposium on Principles of programming languages (POPL '73). Association for Computing Machinery, New York, 1973, USA, 194-206). Using sequential iteration space slicing, these systems can evaluate which sequential iterations of which statements affect the values of a given set of elements from a particular array A. Thus, sequentially, the system can return a set of loop operations assigned to processor p from, for example, loop nest s, by relying on a particular set of data elements accessed by processor p from array A. Furthermore, the prior art document Fonseca A. et al, "Automatic Parallelization: Execution Sequential Programs on a Task-Based Parallel Runtime", International Journal of Parallel Programming, April 2016 discloses another system for automatically parallelizing sequential code for use in a multi-core architecture. This system discloses using data groups and memory layout and then relying on task parallelism to check for dependencies. Thus, in order to automatically parallelize a program, the system needs to analyze the accessed memory to evaluate possible dependencies between parts of the program. For example, in an example of automatic parallelization of code for generating the Fibonacci sequence, the disclosed system evaluates the cost of creating a new task to be higher than the cost of executing the method for low input numbers. This evaluation is to be used as a major requirement regarding the placement of tasks during the automatic generation of parallelized code. Here, the evaluation is done by a specific function that relies on a set of seven requirements to find the best placement. Finally, this function outputs what are called hard dependencies, where a task can be introduced later, and what are called soft dependencies, a set of already defined tasks that the current task has to wait for its execution. Parallelization is complete when all tasks with the specified placement are instantiated. Here, tasks are marked for execution by waiting for the execution of current tasks and reading their results. Finally, U.S. Patent Application Publication No. 2008 / 0263530 discloses a system for converting application code into optimized application code or execution code suitable for execution on a computing architecture having at least first and second level data memory units. When scheduling instructions, the principle of locality, also known as the locality of reference, is used. This relates to the phenomenon that the same value or related memory locations are frequently accessed. Different types of locality of reference are distinguished. Temporal locality means that a resource referenced at one point in time will soon be referenced again. Spatial locality means that the likelihood of referencing a particular memory location increases if nearby memory locations have recently been referenced. Programs and systems exhibiting locality show predictable behavior, thus providing the code designer with opportunities to improve performance by prefetching, precomputing, and caching code and data for future use. For this type of data evaluation optimization, the disclosed system accesses locality prior to layout locality. Thereby, for data that is accessed frequently, accesses are grouped in time if possible when data transfer operations occur, and data that is accessed subsequently is grouped spatially if possible. This selection of one option may be made based on a cost function. As a variation of the embodiment, the system also addresses the issues of parallel data transfer and memory utilization by focusing on portions of the application code having data parallel loops. The conversion structure addresses both data level aspects of memory units such as background memory, foreground memory, registers, and various levels of functional units.
Means for Solving the Problem
[0023] According to the present invention, these objects are achieved in particular by the features of the independent claims. Further, further advantageous embodiments can be derived from the dependent claims and the associated description.
[0024] According to the present invention, a compiler system and a corresponding method for optimizing and compiling program code for execution by a parallel processing system having a plurality of processing units that simultaneously process data in the parallel processing system by executing the program code are achieved. In particular, the compiler system translates the source programming language of the source machine code of the computer program into machine code as the target programming language, and has means for generating the machine code as parallel processing code that is executable by the plurality of processing units of the parallel processing system or includes some instructions for controlling the operation of the plurality of processing units; the parallel processing system has a memory unit, and the memory unit includes at least a main execution memory unit including a plurality of memory banks for holding at least a part of the data of the processing code, and a transition buffer unit storing the start position of the processing code and the data segment, and including at least a high-speed memory including branch or jump instructions and / or memory references and data values used, and the main execution memory unit provides an access time slower than that of the transition buffer unit; the execution of the processing code by the parallel processing system includes the occurrence of a latency time, and the latency time is given by the idle time of the processing unit between sending back data to the parallel processing system after a particular block of instructions of the processing code has been processed on the data by a processing unit and receiving the data required for the execution of consecutive blocks of instructions of the processing code by the processing unit; the compiler system has a parser module for translating the source programming language into code having a flow of basic instructions executable by the processing unit, the basic instructions being selectable from a set specific to the processing unit of the basic instructions, and the basic instructions including basic arithmetic operations and / or logical operations and / or control operations and / or memory operations for the plurality of processing units;The parser module has means for dividing the code of the basic instructions into computational block nodes, each computational block node consisting of the smallest possible segmentation of the non-further decomposable sequence of the basic instructions of the code that can be processed by a single processing unit, the smallest possible segmentation of the basic instructions being characterized by a sequence of basic instructions consisting of consecutive read instructions and write instructions, the sequence not being further decomposable by a smaller sequence of basic instructions between consecutive read and write instructions, the read and write instructions being required to receive the data necessary for processing the sequence of basic instructions by the processing unit and to send back the data after processing by the sequence; the compiler system has a matrix builder for generating a numerical matrix from the computational chains divided from the code depending on the latency time, the numerical matrix including a computational matrix and a transfer matrix, the computational chain being formed by one or more computational block nodes creating an ordered flow of computational block nodes, each computational chain being executed by one processing unit; the computational matrix includes the computational chains of the processing units in each row, and each column has the sequence of basic instructions of the computational block nodes in the computational chains of the rows, the transfer matrix including the transfer characteristics associated with the data transfer from one computational block node to the consecutive computational block node; the compiler system has an optimizer module using a numerical matrix optimization technique for minimizing the overall occurring latency time by integrating all occurring latency times by providing an optimized structure of the computational chains processed by the plurality of processing units, and the code generator generates optimized machine code with an optimized overall latency time for the plurality of processing units.; When matrices or tensors are used, the optimization by the optimization module can be based on, for example, numerical matrix optimization techniques (or more general numerical tensor optimization techniques). Technically, this optimization problem is formulated by using tensors and / or matrices, and in this way, a matrix / tensor field optimization problem can be obtained. For linear optimization, for example, matrices and linear programming can be used by the optimizer module. In certain applications of the present invention, the concept of tensors can be technically beneficial, for example. In optimization, tensor techniques can solve systems of non-linear relationships and equations, for unconstrained optimization using second-order derivatives. The tensor method can be used as a general-purpose method especially for problems where the Jacobian matrix at the solution is singular or ill-conditioned. The tensor method can also be used for linear optimization problems. An important feature of tensors is that their values do not change when they cause regular non-linear coordinate transformations, and thus, this concept can be technically useful for characterizing structural properties that do not depend on regular non-linear coordinate transformations. Thus, tensor optimization can also be applied within the framework of non-linear optimization. However, one of the technical advantages of the present invention is that all matrices known to date are linear with respect to optimization, in contrast to the prior art optimization techniques in the field of automatic parallelization of source code, while prior art systems have to rely mainly on non-linear optimization. Here, optimization means the problem of finding a set of inputs to an objective function that results in a maximum or minimum function evaluation. For this technically difficult problem, various machine learning algorithms can also be used together with the optimizer module, from fitting a logistic regression model to training an artificial neural network. When the optimizer module is implemented by the implemented machine learning structure, it can be formulated to be provided, for example, by using continuous function optimization, and the input arguments to the function are real-valued numerical values, for example, floating-point values. The output from this function is also a real-valued evaluation of the input values.However, as a variation of the embodiment, an optimization function that takes discrete variables, i.e., provides a combinatorial optimization problem, can also be used. To technically select the best optimization structure, for example, one approach could be to group the selectable optimization structures based on the amount of information available about the target function being optimized, and this could be used and utilized by the optimization algorithm. Clearly, the more information is available about the target function, the easier it becomes to optimize the function by machine learning. Of course, this depends on the fact whether the available information can be effectively used in the optimization. Thus, one selection criterion can be associated with differentiable target functions, for example, by the question of whether the first derivative (gradient or slope) of the function can be calculated for a given candidate solution. This criterion divides the available machine learning structures into those that can utilize the calculated gradient information and those that cannot, i.e., machine learning structures that use differential information and those that do not. For applications where the differential objective function can be used, it should be noted herein that a differentiable function refers to a function that can generate a derivative for any given point in the input space. The derivative of a function with respect to a value is the rate of change or amount of change of the function at that point, which is also called the slope. The first derivative is defined as the slope or rate of change of the objective function at a given point, and the derivative of a function with two or more input variables (e.g., multivariate input) is called the gradient. Thus, the gradient can be defined as the derivative of a multivariate continuous objective function. The derivative of a multivariate objective function is a vector, and each element of the vector can be called a partial derivative or the rate of change for a given variable at that point assuming that all other variables are kept constant. Further, the partial derivative can be defined as an element of the derivative of the multivariate objective function. Then, the derivative of the derivative of the objective function, i.e., the rate of change of the rate of change of the objective function, can be generated. This is called the second derivative. Thus, the second derivative can be defined as the rate at which the derivative of the objective function changes.In the case of a function taking multiple input variables, this is a matrix, called the Hessian matrix, which is defined as the second derivative of a function having two or more input variables. A simple differentiable function can be optimized analytically using known analysis. However, the objective function may not be solvable analytically. If the gradient of the objective function can be generated, the optimization used becomes significantly easier. Some machine learning structures that can use gradient information and can be used for the present application include bracketing algorithms, local descent algorithms, first-order algorithms, and second-order algorithms.
[0025] The present invention has the advantage, among other things, of providing and achieving large-scale optimizations based on the lowest possible code structure, reducing high-level programming language code to a small number of basic instructions that cannot be further reduced with respect to data input and data output points for the limited set of machine instructions executed on a CPU / microprocessor. The basic instructions include, for example, (i) arithmetic operations: +, -, *, / in applied numerical applications ->, i.e., mathematical operations such as integration or differential analysis are reduced to these basic instructions, (ii) logical operations: AND, OR, etc., (iii) variable and array declarations, (iv) comparison operations: same, greater than, less than, etc., (v) code flow: jump, call, etc., (vi) if (condition) {codeA} else {codeB}, (vii) loop (condition). The interaction between today's state-of-the-art high-level languages (e.g., python, C, java (registered trademark), etc.) and the limited resources of processor instructions is analyzed and made accessible by creating a "mapping" of the reading and writing of "data points" by their operations. In other words, by mapping the read and write interactions of a single instruction using an appropriate representation (which can also be represented more graphically), they can be made available for numerical optimization techniques, thereby automatically parallelizing the source code and always resulting in executable parallel code. There are several ways to access these interactions, but ultimately it proceeds to map the read and write patterns of the source code to the data introduced by the programmer's variable definitions, then extract the required sequential chains and introduce potential communication patterns, so that there is nothing like "mapping" the code to a wide range of hardware infrastructures. This method discloses a new way of "adapting" the source code to a given hardware infrastructure across all levels (CPU, GPU, cluster).
[0026] The present invention further has the advantage that the disclosed method and system can solve known technical problems such as solving nested loops with an array for solving PDEs (partial differential equations) from a new perspective, or well-known problems that occur in the optimization steps of SOTA compilers. The method gives a new perspective to the code, and this new scope is based on the physical effects that occur in all classical computing infrastructures. As a result, a general method for mapping the computation to a given hardware structure and deriving a simultaneous parallel representation of the code on a given hardware or the ideal hardware for a given code is obtained. This is based on the result of maintaining all dependencies of the "reads" and "writes" of the introduced data nodes and constructing a chain of instructions depending on these dependencies. The resulting computational block nodes and their flow graph return a well-formed base of matrices, and as a result, a general method for returning code applicable to different computing units (e.g., CPU, GPU, FPGA, microcontroller, etc.) is obtained. This method has a new perspective on the interaction between ICT software and hardware according to the principles of system theory. This leads to a method that can bring new solutions in a wide range of fields as follows.
[0027] (a) Adaptive Hardware - FPGA (Field Programmable Gate Array [Field Programmable Gate Array]) / ACID (Atomicity, Consistency, Isolation, Durability [Atomicity, Consistency, Isolation, Durability]): The present invention decomposes the code into chains of instructions that clearly represent the logic elements in the integrated circuit. Since this method gives a general form for optimizing the combination of computation and communication, it optimizes groups of instructions based on the same "bit pattern" / "signal" and can be used, for example, to automatically optimize the transfer of software to an FPGA respectively, or to bring a new approach for filling the gap from the code to chip floor planning.
[0028] (b) ML (Machine Learning) / AI (Artificial Intelligence): Machine learning and artificial intelligence codes require a lot of resources, especially in the training phase. This method can be used, for example, (i) to optimize known codes, (ii) to support code development that adapts complexity at runtime and is thus difficult to parallelize in advance (since this method always results in optimized code), and (iii) to support future methods (such as genetic algorithms, see, for example, Inside HPC Special Report by R. Farber, AI-HPC is Happening Now) rather than neural network-based approaches.
[0029] (c) HPC (High-Performance Computing) applications: This method can convert code from, for example, Python to C code using an MPI support library, thus filling the gap that exists, for example, between different research fields (from HPC to AI development). Another application can be an adaptive mesh refinement implementation used in numerical model software packages for engineering applications, weather prediction models, etc. Alternatively, it can be used to combine models with different spatial and temporal resolutions (such as computational fluid dynamics models and agent-based models) and improve existing software packages in different fields such as modeling and analysis software packages.
[0030] (d) Automated business processes: The present invention can also be used in process management. The decision of whether a unit should work on a task or send the task to another unit is a well-known problem. This method presents one approach to this problem.
[0031] (e) Clouds, desktop operating systems, virtual machines, deployment in general: Enabling a general approach to "shrinking" code with respect to basic required operations and possible parallel options supports a wide range of solutions at the interface between software and hardware. This interface clearly occurs, in particular, for any form of operating system, software virtualization and / or deployment, or more specifically, for example, virtualization solutions for cloud infrastructure, operating systems (with multi-core systems), virtual machines supporting a mix of different operating systems, or the like.
[0032] (f) Heterogeneous platforms, (i) IoT and edge computing: Heterogeneous platforms that govern a given situation in different fields, such as IoT projects, autonomous driving, combined mobile and cloud applications, and other forms of applications executed with and / or on hybrid hardware infrastructure. The method can adapt the code to determine how to optimally distribute data, computing, and / or data communication on the platform. Further, it can be incorporated into the process of deploying / developing software for different characteristics of the hardware components of a given network of computing units to optimize the code to meet target characteristics, such as reducing latency for some parts of a software system.
[0033] (g) Embedded systems: Embedded systems have high requirements for a given code, for example, for power consumption or other specific adaptations. For example, only a reduced instruction set on some microprocessors, or a similar issue. The method can optimize for given physical characteristics and thus directly support this mapping, resulting in the most efficient code representation for any given code.
[0034] (h) Self-optimization algorithm: The present invention allows for a completely autonomous cycle. This means that the algorithm can optimize itself on a given platform without any manual interaction. This enables new applications and fields that were not previously known.
Brief Description of the Drawings
[0035] The present invention will be described in more detail by way of example with reference to the drawings.
[0036]
Figure 1
[0037]
Figure 2
[0038]
Figure 3
[0039]
Figure 4
[0040]
Figure 5
[0041]
Figure 6
[0042]
Figure 7
[0043]
Figure 8
[0044]
Figure 9
[0045]
Figure 10
[0046]
Figure 11
[0047]
Figure 12
[0048]
Figure 13
[0049]
Figure 14
[0050]
Figure 15
[0051]
Figure 16
[0052]
Figure 17
[0053]
Figure 18
Figure 19
[0054]
Figure 20
[0055]
Figure 21
[0056]
Figure 22
[0057]
Figure 23
[0058]
Figure 24
[0059]
Figure 25
[0060]
Figure 26
[0061]
Figure 27
[0062]
Figure 28
[0063]
Figure 29
[0064]
Figure 30
[0065]
Figure 31
[0066]
Figure 32
[0067]
Figure 33
[0068]
Figure 34
[0069] Regarding assigning a runtime number to a block graph or a tree structure, the numbering can be implemented, for example, as a recursive function that parses a graph of computational block nodes and their edges for each branch node. At each branch node 3341, the computational block node 333 starts the numbering. Then, it is stepped recursively through the computational block node chain 341 of the branch node 3341. After the branch node, the maximum block number is known given the computational block nodes within the branch node. This is used for the next branch node. The rule is that the computational block node 333 steps to the computational block node 333. If there is only one previous computational block or no previous computational block -> set that block number to the actual block number and increment the block number (for the next computational block node). If there are two preceding computational block nodes -> this is a merge situation -> if this is the first visit to this node -> append to the local list of block numbers. Otherwise, if this is the second visit -> use the highest block number (actual from the function call or saved in the list of nodes). If there is one subsequent computational block node -> call the function for the computational block node (recursive approach). If there are two subsequent computational block nodes -> split situation -> call both computational block node numbering functions recursively. Otherwise (thus, if there is no next computational block node) -> give the next branch node and, if any, end the actual recursive call. By appropriately adjusting the block numbers in the branch node transmission, the call graph is numbered in the form of a discrete-time graph based on the computational block node connections, and each computational block node has a finite number that must be computed during the same time period.
[0070]
Figure 35
[0071]
Figure 36
Figure 37
[0072]
Figure 38
[0073]
Figure 39
[0074] Regarding optimization, there are a number of currently applicable automatic optimization techniques: combining rows to reduce parallelism, moving a single operation chain to the earliest point (the send command in a cell is like a barrier), reducing communication with the best combination of rows, etc. The compiler system 1 can be used to obtain an estimate of the runtime of the cell entry, or a method such as that by Agne Fog of the Technical University of Denmark can be used to extract CPU and cache interactions, or a table from the CPU manufacturer can be used, or it can be compiled with openCL, etc. For example, in the perspective of having one tensor for computation and one tensor for transfer, different parallel code versions are obtained by combining the same rows in each tensor, and thus new combinations of "computation and transfer" can be generated for each block number. This leads to a reduction in the number of parallel / simultaneous parallel units. By combining rows, it is possible to reduce or group transfers (in the final code communication), and it is possible to consolidate computations. If there is no send() or read() at the beginning of a block, the operation nodes within the block can be moved to the previous block. Each cell knows the amount of data (such as memory) it needs. Different given Δt on the target platform latencyUsing 35, a determination can be made as to which sequential parts (cell entries within the computational matrix) must be computed on which hardware units. The communication types in the infrastructure can be implemented as needed, ranging from asynchronous or non-blocking to explicit blocking of send and receive commands in the MPI framework, prevention of race conditions by ensuring that the correct barriers are set and released, bulk copy transfers in the GPU infrastructure, etc. Thus, the selection of optimization techniques can be easily done by selecting appropriate prior art optimization techniques, for example, by using a SOTA compiler for each piece of code (computation and transfer) per unit. The optimization techniques used introduce a much broader perspective in order to obtain a more "perfectly parallel" code than others (perfect in the sense of a linear dependence of speedup on the number of processes with a slope of 1). Thus, it is possible to numerically optimize the matrix for the new parallelized / optimized / simultaneously parallel code for the target hardware. This can be done automatically and is thus a large step compared to other methods. Since this is done by software, the software can now parallelize its own code, which is new and brings new possibilities. For example, adaptive models in machine learning (ML) or artificial intelligence (AI) applications, meshes in computational fluid dynamics (CFD) calculations, or particle sources in vortex methods, or combinations of different model methods with different spatial and temporal resolutions (finite volume method (FVM) with agent-based models and statistical models), etc.
[0075]
Figure 40
[0076]
Figure 41
[0077]
Figure 42
[0078]
Figure 43
Figure 44
Figure 45
[0079]
Figure 46
[0080]
Figure 47
[0081]
Figure 48
[0082]
Figure 49
[0083]
Figure 50
[0084]
Figure 51
[0085]
Figure 52
[0086]
Figure 53
[0087]
Figure 54
[0088]
Figure 55
Figure 56
[0089]
Figure 57
[0090]
Figure 58
[0091]
Figure 59
[0092]
Figure 60
[0093]
Figure 61
[0094]
Figure 62
[0095]
Figure 63
[0096]
Figure 64
[0097]
Figure 65
[0098]
Figure 66
[0099]
Figure 67
[0100]
Figure 68
[0101]
Figure 69
[0102]
Figure 70
[0103]
Figure 71
[0104]
Figure 72
[0105]
Figure 73
[0106]
Figure 74
Mode for Carrying Out the Invention
[0107] Definition (i) "Computing block node" Grouping instructions, the definition of the term "computing block node" related to communication / transfer data to other computing block nodes is important for this application. The term "computing block node" used in this specification has no generally recognized meaning and is different from similar terms used in the current state of the art.
[0108] Well-known basic blocks (see, e.g., Proceedings of a symposium on Compiler optimization; July 1970, pages 1-19, https: / / doi.org / 10.1145 / 800028.808479) are central definitions in classical control flow graphs (CFGs). Simplifying, they group statements that do not have jumps or jump targets internally. Thus, for a given input, they can execute operations without interruption until their respective outputs, until the end. This is a fundamental concept in today's compilers. The definition of basic blocks has historically been targeted at a single computational unit and is very well established. There are optimization methods for a wide range of problems, showing how they solve different technical problems. However, looking at the code with the aim of splitting statements into different dependent units (connected, e.g., by a shared cache, via a bus or network, etc.), this definition lacks granularity, and the classical scope hinders a broader perspective. By determining the scope of a block based on any unique information given in the code (viewing information as bit patterns), combining this scope with relevant times, and computing and transferring information within the system, a different but physically well-founded perspective is created for a given code. This alternative scope enables new options and solves some well-known technical problems of today's SOTA compilers (see the following examples for PDE, Fibonacci, or pointer ambiguity resolution).
[0109] The term "computing block node" as used in the present application is based on the mutual relationship between transfer and computing time for a given set of sentences, which are different and not thus applied, in a SOTA compiler. These newly defined "computing block nodes" group together instructions that use the same information that is not changed by any other instruction (sentence) within any other computing block node during a specific time within the complete code. In this way, they group instructions that can be processed or computed independently of any other sentence having a scope with respect to information (information as a clearly distinguishable bit pattern), which is called a "non-further-dividable instruction chain" in the present application. Each of these instruction chains in a computing block node has a physically based "time" related to how much time it takes for a unit to process or compute them on a given hardware. Since hardware characteristics have a fundamental influence on the time required to process or compute instructions (as well as software components such as the OS, drivers, etc.), the "computing block nodes" are also correlated with the time required for possible transfer of any other information within the complete code to another "computing block node" during a specific time - if required during a specific program step. Each "computing block node" knows which information for its instructions needs to be exchanged (communicated / transferred, "received" or "sent" respectively) with other "computing block nodes" at what time. Therefore, the "computing block node" as used in the present application brings new scopes and decision criteria, which means, in the context, one of the central aspects of code parallelism: computing information (bit pattern) on a unit or transferring this information to another unit for parallel computing. This decision can only be made if it is guaranteed that the information used is not changed during a specific program step (or time) in any other part of the program.Furthermore, the building block nodes having this scope not only bring the advantage of parallelizing a given code, but this scope also shows several advantages for problems that are not well handled by SOTA compiler optimization techniques. Such problems and various solutions are well documented, for example, in the publication "Modern compiler design" by D. Grune. To show the technical benefits, several advantages resulting from using the new scope for some of these known technical problems are shown below, such as pointer ambiguity resolution and different performance for Fibonacci sequence code, and further, how the present method solves problems that have not been solvable until now, such as PDE parallelization, is shown below.
[0110] First, the different scopes are illustrated using a schematic example in FIG. 6. This figure has been somewhat adapted to more easily show what the system and method of the present invention do. For example, the resulting "compute" -> "communicate" model, in which the statement "y > a" does not occur twice within the same computational chain, results in a corresponding arrangement of loops, jumps, or flow control instructions. Nevertheless, this example shows what the different scopes of the computational block nodes give compared to basic blocks: "a" and "b" do not change in block 1, the information of the condition "y > a" is already known after the evaluation of the statement "y := a * b", and the alternative scope of the computational block nodes takes these characteristics into account. The proposed view of the computational block nodes groups statements together in a new way. This change in view results from a technique of grouping operations based on information that does not change at the same time step within the code. In this example, it is shown that two independent computational chains have developed, both of which are independent with respect to the information of "y".
[0111] In a simplified but more realistic example of FIG. 7, the method of the present invention utilizes the fact that logarithmic instructions (e.g., FYL2X in modern CPUs) take much longer than floating-point addition and / or multiplication instructions (e.g., FADD, FMUL). The method of the present invention adds the statements "x:=a+b" and "y:=a*b" to two different computational block nodes (cbn), because both are based on the same information "a" and "b". The "log2(x)" statement is appended to "x=a+b" because it uses the information "x". The information "y" is transferred and the information is available for both independent computational chains.
[0112] (ii) "Computing matrix" and "transfer matrix" The method of the present invention forms two technically defined matrices, called "computation matrix" and "transfer matrix" in this application, from the flow graph of computational block nodes as described above. Both are numerical matrices. The "computation matrix" contains instruction chains, and the "transfer matrix" contains possible transfer characteristics (from other computational block nodes and to other computational block nodes). Therefore, the code extracted from these matrices always forms the pattern "compute->communicate" [compute -> communicate], as will be explained in more detail in the following section. When the code is mapped to one unit, the communication part disappears and the method of the present invention is reduced to an approach with basic blocks, which can each be processed by a SOTA compiler. Referring to this form of the example of FIG. 7, it can be seen that unit 2 can compute the long-running execution instructions of the statement "log2(y)", and unit 1 can process the loop using the increment of a ("a=a+1") (see FIG. 8). To see how this works, the matrix builder is shown in detail in the following section.
[0113] The name "matrix" is used to name a structure of the form (m×n×p), where m, n, p ∈ N0. Since m, n, p depend on the code, this can include different forms of mathematical objects, particularly with respect to the dimensions of points, vectors, matrices, tensors, etc. m is the number of the maximum number of computational blocks, each being a block number and a segment number as in Figure 43. n is the number that, although independent, indicates the chain number of computational block nodes with the same segment number as in Figure 43. Depending on the maximum level of conditions in the code (or a series of branch nodes as in Figure 26), p is defined and is shown as the path number in Figure 43. Thus, whenever the term "matrix" is used, it can be said that it can be a vector, a matrix, or a tensor, or any other object having a representation in the form of (m×n×p), or in the form of a graph or a tree. Therefore, optimization can also be carried out using one tensor, for example, by combining a computational matrix and a transfer matrix into one structure, or in a code having only one block / segment number, the transfer matrix and the computational matrix may each be a tensor, or both may be vectors. The dimensions of these "matrices" depend on the form of the code and the way information is handled / represented in the way the method is applied.
[0114] So-called "numerical matrices" can also include text forms such as the transfer "1->2". Depending on the optimization / mapping technology used, the text in the matrix can be numerical (depending on the character encoding used) or can be reduced to numbers, and thus they can be searched or compared, for example. Alternatively, the text transfer "1->2" can be represented / encoded by numbers from the start and directly compared with other transfers, thus omitting the character encoding.
[0115] Since a "numerical matrix" can be used as another appropriately formatted form representing a graph / tree-like structure, it is possible to function without a "numerical matrix" and perform all optimizations / mappings in the form of a graph / tree. Whatever the mathematical or computational form representing a group of instructions (here named computational block nodes) and their transfer dynamics, it is represented here, for example, by the "transmission package" mentioned in Figure 23, which is based on the physical underlying dependencies of transfer and computational latency occurring in all binary-based electronic computing systems. This is based on the rule of placing an instruction to "read" information A in the same group (here named computational block node) where the instruction to "write" information A is located, and grouping instructions (any form of instruction / operation, for example, representing the form of an electronic circuit) in the code (any form of code (high-level, assembly, etc.)). This is significantly different from the grouping of instructions with well-known basic blocks. When the dependencies are not clear (shown in Figure 21 in the case of a graph element with two "read" nodes and one "write" node per instruction), a "transfer" (from the one holding the "write" instruction to the one holding the "read" instruction) is introduced between two computational block nodes. In this form, the code can be represented in the form of computational block nodes connected by the required transfers. This information can be included in a graph / tree or an array or any appropriate structure. To execute these instructions on a given hardware, the well-known methods in state-of-the-art compilers, which optimize the chain of computational block nodes grouping the instructions and then can be used (in this case called transpiling) for execution on the involved computing units for each unit, shall not be used.
[0116] (iii) "Matrix builder" and how to return to the code After parsing the code and adding all instructions to the computational block nodes (cbn), each computational block node is numbered depending on its position in the flow graph. This leads to a similar form of the control flow graph given by the edges between the cbn and the connections of the defined branch nodes. These positioning numbers can be used to place information about the unique position in the flow of the code, what to compute and what to transfer, in two matrices, the "computation matrix" and the "transfer matrix", i.e., the "computation matrix" for computation and the "transfer matrix" for transfer. It is clear that metadata such as the size of the data required for each cbn, the size of the transfers between cbn, etc. can be easily derived.
[0117] Matrices represent a form of information access that is far more scalable than graphs, and each exhibits the well - formed properties of a control - flow graph. These are not absolutely essential for the present method, and this step can also be performed directly on the graph / tree structure. However, the definition of block nodes also exhibits the general properties of the matrices of the present invention. Each row has an independent flow of instructions (= calculations) (dependency by transfer) and the required transfers (= communications) to other cbn. It is guaranteed that a) no additional information is required to compute all the instructions within a computing block node (which is similar to a basic block, but with a completely different scope of independence), b) the information used is not changed anywhere else during the same computational step throughout the code, and c) only information that is not affected by the calculations during this time step can be transferred. This fact enables each cell within the computational matrix to have all the instructions that can be computed independently and in parallel with all the instructions in other cells within the same column. In the transfer matrix within each cell, the required transfers at the beginning and end of each computational step (the corresponding cell of the computational matrix) are now known. Returning the code for execution on different units results in an expression in the form of "communication -> calculation -> communication -> calculation", etc. Each row represents a chain of computational and communication characteristics that form a series of calculations combined by communication with other rows = chains of calculations. The information required to communicate with other rows = chains is within the transfer matrix. Each row in the computational matrix (and the same combination in the transfer matrix) can be combined with any other row in the matrix (computing all the instructions in both combined cells and performing the necessary transfers for both cells based on the transfer matrix) to create a new combination of the compute <-> transfer behavior of the given code. In this step, the calculations are grouped together, and (by combination) the transfers on the same unit disappear. This leads to the fact that later, simple forms of optimization, as well as the need for sequential iterative or similar solutions, are not required, resulting in code for which the optimization / mapping step can be reliably executed.Debugging parallel code is a very complex problem, which is also code parallelization of a well-known problem in the art (see, for example, "ParaVis: A Library for Visualizing and Debugging Parallel Applications", A. Danner et al.).
[0118] Each row of the computational matrix defined here represents a chain of instructions for one unit. The unit depends on the level of implementation (e.g., bare assembly, thread, process, computing node, etc.). Obviously, empty blocks (empty cells within the computational matrix) or unused communication entries (empty cells within the transfer matrix or transfers on the same unit) disappear. As seen in Figure 10, the start and end communications that link together are also like this. This occurs when there are no transfers when the code is executed on a single unit and all transfers disappear by variable re-assignment / name change (applying well-known optimization methods in SOTA compilers to obtain optimized code for a given unit respectively).
[0119] Depending on how communication is implemented, non-blocking or blocking mechanisms can be used. This is because it is guaranteed that information used between computational block nodes, simultaneously in the said instruction, or in another cbn with the same number, is not transferred. Depending on the level, the computational and communication parts can be implemented, and as a result, this method can be used as a compiler or transpiler. The transfer back to code in the form of a "computation -> communication" approach also facilitates using the most suitable language for the application (e.g., C, C++, python), depending on the target infrastructure such as IPC (InterProcess Communication) methods (e.g., queues / pipes, shared memory, MPI (Message Passing Interface), etc.), libraries (e.g., eventlibs, multiprocessinglibs, etc.), and the use of available SOTA compilers.
[0120] (iv) "Optimization" The general well - defined properties of the matrices are a unique basis for a wide range of possibilities for mapping / optimizing code to a given hardware or for evaluating the optimal hardware configuration for a given code. This structure, as a result, guarantees executable code. Each hardware infrastructure has its own performance characteristics, and in combination with the available software layers, modern ICT (Information and Communication Technology) infrastructure is very complex. The method of the present invention allows for optimizing code for the hardware or for providing the ideal hardware. Most obviously, it is by constructing different combinations of rows from the computation matrix and the transfer matrix, and it is important that the same combination is constructed in both matrices. Each combination (for example, combining row 1 and row 2 in the transfer matrix) is a new version of the parallel / simultaneous code for a given input code, and its characteristics (see Figure 11 for a simple example) on the target infrastructure can be evaluated. and By different combinations of rows in the matrices (for example, combining row 1 and 2, or row 2 and 5 in the computation and transfer matrices), different combinations of the computation <-> communication ratio are obtained. Then, each of the combinations including other known metadata for a given hardware / software infrastructure can be examined. For example, the data type of the data node () can be used to evaluate hardware characteristics such as cache length, available memory, or other characteristics of the target platform. This form of exploring by combining the optimal computation / communication ratio for a given code and a given hardware does not require a sequential iterative solution or a similar approach to find a solution in the optimization step, and for example, no solution involving a race condition, deadlock, etc. will occur, or a deadlock can be detected, so it always results in executable simultaneous code.
[0121]
[0122] The grouping of instructions is based on an inherent physical constraint that transfers are typically larger than would be computed at the same "location". This method produces the form of the optimal solution space for the most divisible form of the code, resulting in a well - defined way and search space for finding the optimal map for a given hardware or the ideal hardware for a given code. This makes the method very general - purpose and solves the technical problem of automatically adapting a given code to a target platform. When this method is used as a transpiler, SOTA compilers can be used to optimize the code for the target hardware / unit.
[0123] Other methods do not utilize the inherent temporal characteristics and dependencies given in the code and do not generate this form of the unique solution space for optimization in the form defined by this method of computational block nodes (by grouping based on the scope of unchanging information), which directly solves the technical problem in many ways. Examples in the detailed description show this in more detail.
[0124] (v) "Computing block node" according to the present invention and "basic block or slice" of the prior art Computational block nodes are not the same as basic blocks or slices. They have different scopes. They do not follow the definition of basic blocks, for example, by connecting independent program parts based on jumps / branches. The instruction chains in computational block nodes have basic data dependencies, meaning that the information used in these chains is not changed anywhere else in the given code and is not transported at the same time / program step.
[0125] Therefore, the time or location of the instructions within the program is subject to the dependencies in the computational block nodes. The computational block nodes consist of chains of instructions based on the same information, and the information is defined in any form of bit pattern (e.g., data variables, pointer addresses, etc.). By introducing transfers / communications that not only use the information at time steps within the code but also combine with the instructions, the time correlation of these two physical bases is achieved within the computational block nodes. This method arranges the computational blocks so that it is clear at all times which information can be transferred where and which information can be computed in parallel. This gives different perspectives, especially for optimization techniques. There is a wide range of technical problems that can be solved by this change in perspective. The computational block nodes associate the position of the information (bit pattern) within the computational framework with the time when this information is used within the program. This is supported by the basic physical principles of computing in classical infrastructure. The following example with nested loops and arrays shows that this effect is good. The branches in the loop definition can be transformed into reading and writing data points according to the array when they are distributed across parallel computational block nodes.
[0126] The method of the present invention divides the code into segments of computation, and the resulting information has to be transferred. Thus, by the definition of the computation block nodes, the method of the present invention is given the constraint that no other place within the infrastructure requires the same information during any given point in time where instructions based on the same information are grouped, to generate a matrix system. The grouped instructions cannot be further divided as they cannot reach faster computation for a particular instruction group. This is because any form of transport is longer than computing this given chain at the various computation block nodes. As shown, the general properties of the computation and transfer matrices enable the split code to be optimized to the most parallelizable solution possible for a given hardware. The ratio of computation to transport depends on the target hardware, and the method provides different solutions for splitting at different ratios for a given code. By converting the dependency graph into a matrix, these can be used in a more effective way to map / optimize the split code to the target platform, which includes the specific characteristics of this infrastructure (for example, GPUs require a different handling of transfer / computation distribution than CPUs). However, using matrices is not essential, and the optimization / mapping can be done directly on the graph / tree structure.
[0127] (vi) "Idle time" - "latency time" The latency time as defined herein is given by the idle time of processing unit 21 between sending data back to parallel processing system 2 after processing a particular block of instructions of processing code 32 on the data by processing unit 21 and receiving (i.e., after retrieving and / or fetching) the data required for execution of successive blocks of instructions of processing code 32 by the same processing unit 21. In contrast, the idle time of a processing unit can be defined herein as the amount of time the processing unit is not busy between two computational block nodes or, otherwise, the amount of time to execute an idle process of the system. Thus, the idle time allows for measuring the unused capacity of the processing units of a parallel processing system. Maximum speedup, efficiency, and throughput are the ideal case for parallel processing, but these are not achieved in practice because the speedup is limited due to various factors contributing to the idle time of the processing units. The idle time of the processing unit as used herein can find its origin in various causes including, inter alia: (A) Dependency between successive computing block nodes : There may be a dependency between the instructions of two computational block nodes. For example, an instruction may not be able to start until a previous instruction returns a result. This is because both are independent. Another example of data dependency is when both instructions attempt to modify the same data object, also called a data hazard; (B) Resource constraint conditions : Delays occur in pipelining when resources are not available at execution time. For example, if one common memory is used for both data and instructions and there is a need to read / write and fetch instructions simultaneously, only one can execute while the other has to wait. Another example is of limited resources such as execution units that may be busy at the required time; (C) Branch instructions and interrupts in the program: Programs are not a straight - line flow of sequential instructions. There may be branch instructions that change the normal flow of a program, which can delay execution and potentially affect performance. Similarly, interrupts may occur, deferring the execution of the next instruction until the interrupt is serviced. Branches and interrupts can potentially have a detrimental effect on minimizing idle time.
[0128] The task of minimizing idle time is sometimes referred to as "load balancing", which is noted to mean the goal of distributing work among processing units so that, ideally, all processing units are always kept busy.
[0129] (vii) "Basic instruction" [elementary instruction] As used herein, the term "elementary operation" refers to machine operations that do not include simpler operations. The execution of an instruction typically consists of the successive execution of several operations, including resetting of registers, resetting of memory storage, shifting a character in a register by one digit to the left or right, transferring data between registers, and comparing and performing logical OR and AND on data items. A set of elementary operations can provide the structure for executing a particular instruction. Elementary operations include the basic logical functions of logic gates including AND, OR, XOR, NOT, NAND, NOR, and XNOR. Such elementary operations can be assumed to require a certain amount of time on a given processing unit and can vary only by a certain factor when executed on different processing units 21 or in a parallel - processing system 2.
[0130] In a first step, which will be described in more detail below, the system of the present invention converts source code into a sequence of structured elementary operations 321,..., 325 or code 32 structured in loops, branches, and sequences. It is independent of the platform and compiler optimization level, and thus the same conversion can be used to optimize execution time on any platform.
[0131] This approach is based on decomposing source code 31 written in a programming language into basic operations 32 / 321, …, 325, i.e., distinct transformed parts of the source code 31. The set of basic operations is finite for each processing unit and has several subsets such as integer, floating-point, logical, and memory operations. These sets correlate with part of the processor's architecture and the memory data path. The basic operations used in this specification can be classified into different levels as follows, for example. The top level includes four operation classes: integer, floating-point, logical, and memory. The second level of classification can be based on the origin of the operands (i.e., the location in the memory space), i.e., local, global, or procedure parameter. Each group can exhibit different timing behaviors: local variables are frequently used and are almost always in the cache, while global and parameter operands have to be loaded from any address and can cause cache misses. The third level of classification is by operand type, i.e., (1) scalar variables and (2) arrays of one or more dimensions. Pointers are treated as scalar variables when the pointer value is given using a single variable, or as arrays when the pointer value is given using multiple variables. Operations belonging to the integer (INTEGER) and floating-point (FLOATING POINT) classes are addition (ADD), multiplication (MUL), and division (DIV). The logical (LOGIC) class includes logical operations (LOG): (i.e., AND, OR, XOR, and NOT), and shift operations (SHIFT): operations that perform bit-by-bit shifts (e.g., rotation, shift, etc.). The operations of the memory (MEMORY) class are as follows: single memory assign (ASSIGN), block transaction (BLOCK), procedure call (PROC). A memory block (MEMORY BLOCK) represents a transaction of a block of size 1000 and can have only array operands.A memory procedure (MEMORY PROC) represents a function call with one argument and one return value. The argument can be a variable or an array that is declared locally or given as a parameter of the calling function rather than globally.
[0132] Detailed description of the preferred embodiment FIG. 12 and FIG. 13 schematically show an architecture for a possible implementation of an embodiment of a compiler system 1 for optimized compilation of program code 3 for execution by a parallel processing system 2 having a plurality of processing units 21. Note that in this specification, when an element is referred to as being “connected” or “coupled” to another element, that element may be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being “directly connected” or “directly coupled” to another element, there are no intervening elements. Further, some portions of the following detailed description are presented using procedures, logic blocks, processes, and other symbolic representations of operations on data bits within a computer memory. These descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. In this application, procedures, methods, logic blocks, processes, etc. are considered to be a self-consistent sequence of steps or instructions that result in the desired outcome. Variations of the embodiments described herein may be discussed in the general context of processor-executable instructions, such as program code or code blocks, that are present on some form of non-transitory processor-readable medium and executed by one or more processors or other devices. Generally, program code includes routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functions of the program code may be combined or distributed as desired in various embodiments. The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof, unless specifically described as being implemented in a particular manner. Any features described as modules or components may also be implemented together in an integrated logic device or separately as discrete but interoperable logic devices.When implemented in software, the techniques can be realized at least in part by a non-transitory processor-readable storage medium having instructions that, when executed, perform one or more of the methods described above. The non-transitory processor-readable data storage medium may form part of a computer program product. For firmware or software implementations, the methods may be implemented using modules (e.g., procedures, functions, etc.) that have instructions to perform the functions described herein. Any machine-readable medium that tangibly embodies the instructions may be used in implementing the methods described herein. For example, the software code may be stored in a memory and executed by one or more processors. The memory may be implemented, for example, within the processor as a register, or external to the processor.
[0133] The various illustrative logical blocks, modules, circuits, and instructions described in connection with the embodiments disclosed herein may be executed by one or more processors or processor units 21. The 21 may include, for example, a processor 2102 having a control unit 2101, registers 21021 and combinational logic 21022, and / or a graphics processing unit (GPU) 211, and / or a sound chip 212 and / or a vision processing unit (VPU) 213, and / or a tensor processing unit (TPU) 214, and / or a neural processing unit (NPU) 215, and / or a physical processing unit (PPU) 216, and / or a digital signal processor (DSP) 217, and / or a synergistic processing unit (SPU) 218, and / or a field programmable gate array (FPGA) 219, or, for example, a motion processing unit (MPU) and / or a general-purpose microprocessor and / or an application specific integrated circuit (ASIC) and / or an application specific instruction set processor (ASIP) and any other processor unit 21 known in the art, or other equivalent integrated or discrete logic circuits, etc. The term "processor" or "processor unit" 21 as used in the specification may refer to either the above structure or any other structure suitable for implementing the techniques described herein. Further, in some aspects, the functions described herein may be provided within dedicated software modules or hardware modules configured as described herein. Also, the techniques may be implemented entirely in one or more circuits or logic elements. The general-purpose processor 21 may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. In the described embodiments, the processing elements refer to multiple processors 21 and related resources such as memory or memory unit 22.Some of the exemplary methods and apparatuses disclosed herein may be implemented, in whole or in part, to facilitate or support one or more operations or techniques for processing code in a plurality of processors. The multiprocessing system 2 may also include a processor array that includes a plurality of processors 21. Each processor 21 of the processor array may be implemented in hardware, or a combination of hardware and software. The processor array may represent one or more circuits capable of performing at least a portion of an information computing technology or process. By way of example and not limitation, each processor of the processing array may include one or more processors, controllers, microprocessors, microcontrollers, application specific integrated circuits, digital signal processors, programmable logic devices, field programmable gate arrays, etc., or any combination thereof. As described above, the processor 21 may be any of a general-purpose central processing unit (CPU), or a dedicated processor such as a graphics processing unit (GPU), a digital signal processor (DSP), a video processor, or any other dedicated processor.
[0134] The present invention has a compiler system 1 having subsystems 11, …, 16. In a non-limiting embodiment, the subsystems include at least a lexer / parser 11, and / or an analyzer 12, and / or a scheduler 13, and / or a matrix module 14, and / or an optimizer module, and / or a code generator 16. Further, they can also comprise a processor array and / or a memory. The compiler 1 segments the code into code blocks. For the described embodiments, a block or code block refers to a section or portion of code grouped together. The grouping enables groups of statements / instructions to be treated as if they were one statement, and also enables the scoping of variables, procedures, and functions declared within the block to be restricted so as not to collide with variables having the same name used elsewhere in the program for different purposes.
[0135] The above-mentioned memory or memory unit 22 of the parallel processing system 2 can comprise any memory for storing code blocks and data. The memory 22 can represent any suitable or desired information storage medium. The memory can be coupled to the processor unit 21 and / or the processing array. As used herein, the term "memory" 2 refers to any type of long-term, short-term, volatile, non-volatile, or other memory, and should not be limited to any particular type of memory or a particular number of memories, or to the particular type of medium on which the memory is stored. The memory 2 can include, for example, a primary storage unit 211 as a processor register 2111, and / or a processor cache 2112 including a multi-level cache such as an L1 cache 21221, an L2 cache 21222, etc., and / or a random access memory (RAM) unit 2113. Regarding this application, it should be noted that the problem of the multi-level cache lies in the trade-off between cache latency and hit rate. The larger the cache, the better the hit rate, but the latency becomes longer. To address this trade-off, multiple levels of cache can be used, with a small, fast cache being backed up by a larger, slower cache. The multi-level cache generally operates by first checking the fastest level 1 (L1) cache. If a hit occurs, the processor can proceed with processing at a higher speed. If the smaller cache misses, the next fastest cache (level 2, L2) is checked, and so on until access to the external memory. As the latency difference between the main memory and the fastest cache (see Figure 1) increases, some processors are beginning to utilize three levels of on-chip cache.Memory 2 can further include, for example, a secondary storage unit 212 and / or a third storage unit 213 (e.g., tape backup, etc.). The secondary storage unit 212 can include, for example, a hard disk drive (HDD) 2121 and / or a solid state drive (SSD) 2122 and / or a universal serial bus (USB) memory 2123 and / or a flash drive 2124 and / or an optical storage device (CD or DVD drive) 2125 and / or a floppy disk drive (FDD) 2126 and / or a RAM disk 2127 and / or a magnetic tape 2128, etc. In at least some implementations, one or more portions of the storage media described herein can store signals that represent information represented by a particular state of the storage media. For example, an electronic signal representing information can be “stored” in a portion of the storage media (e.g., memory) by affecting or changing the state of a portion of the storage media to represent the information. Thus, in a particular implementation, such a change in the state of a portion of the storage media for storing a signal representing information constitutes a conversion to a different state or thing of the storage media. As described above, memory 2 can include, for example, a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a first in first out (FIFO) memory, or other known storage media.
[0136] (Micro)processors are based on integrated circuits that enable arithmetic and logical operations to be performed based on (two) binary values (in the simplest case 1 / 0). For this purpose, the binary values must be available for the calculation unit of the processor. The processor unit has to obtain two binary values in order to calculate the result of the expression a = b operand c. The time required to obtain data for these operations is known as latency time. There is a wide range of latency times for registers, L1 caches, memory access, I / O operations, or network transfers, as well as for these latency times from the processor configuration (e.g., CPU or GPU). Since every single component has a latency time, the overall latency for a calculation is mainly the combination of hardware components required to obtain data from one location to another in a modern computing infrastructure. The difference between the fastest and the slowest location for the CPU (or GPU) to obtain data can be large (in the range of more than 10^9 times).
[0137] Latency from a general perspective is the time delay between the cause and effect of some physical change in the system being observed or measured. The latency used in this specification is directly related to the physical structure of the multiprocessing system 2. The multiprocessing system 2 includes an integrated circuit-based processor unit 21, which enables arithmetic and logical operations based on (two) binary values (in the simplest case 1 / 0). These binary values must be available for the computing units of the processor. The processor unit needs to obtain two binary values to calculate the result of the expression a = b operand c. The time required to obtain data for these operations is known as the latency time. There is a wide range of hierarchical values for these latency times from registers, L1 cache, memory access, I / O operations, or network transfers, as well as from the processor configuration (e.g., CPU or GPU). Since every single component has a latency time, the overall latency for a calculation is mainly a combination of the hardware components necessary to obtain data from one location to another within the infrastructure of the multiprocessing system. It is notable that while the speed of microprocessors has increased by more than a factor of 10 in 10 years, the speed of general-purpose memory (DRAM) has only doubled, i.e., the access time has been halved. Thus, the latency of memory access with respect to the processor clock cycle increases by a factor of 6 in 10 years. The multiprocessing system 2 exacerbates this problem. In a bus-based system, establishing a high-bandwidth bus between the processor and memory tends to increase the latency of obtaining data from memory. When the memory is physically distributed, the latency of the network and network interface is added to the latency of accessing local memory on the node. The more nodes there are, the more communication there is for a calculation, the more jumps there are within the network for general communication, and the more contention there is, so the latency typically increases with the size of the multiprocessing machine 1.The main goal of parallel computing hardware design is to reduce the overall usage latency of data access by maintaining a high scalable bandwidth, while the main goal of parallel processing coding design is to reduce the overall idle time of the processor unit 21. Generally, the idle time of the processor unit 21 may have several causes such as memory access, deadlock or race condition (for example, when the sequence or timing of code blocks or threads processed by the processor unit 21 depends on each other, that is, depends on the relative timing between interfering threads). As used herein, deadlock refers to a state in which a member of the processor unit 21 is waiting for the output of another member, for example, the output of an instruction block processed by another processor unit 21. Deadlock is a common problem in multiprocessing systems 1, parallel computing, and distributed systems where software and hardware locks are used to arbitrate shared resources and enforce process synchronization. Thus, deadlock as used herein occurs when a process or thread enters a waiting state because the requested system or data resource is held by another waiting process or has not yet been achieved by that process (the process may be waiting for another resource or data held by another waiting process). When the processor unit 21 cannot be further processed because the resource requested thereby is being used by another waiting process (data access or output of an unfinished process of another processor unit 21), this is shown herein as a deadlock leading to the idle time of each processor unit 21.
[0138] The compiler system 1 includes means for translating the source programming language 31 of the computer program 3 into machine code 32 as the target programming language, and generates processing codes 3.1, …, 3.n that are executable by a plurality of processing units 21 of the parallel processing system 2 or include some instructions for controlling the operations of the plurality of processing units 21. The source programming language may be, for example, a high-level programming language 31. The high-level programming language 31 may include, for example, C and / or C++ 311 and / or python 312 and / or Java (registered trademark) 313, Fortran 314, OpenCL (Open Computing Language [Open Computing Language]) 315 or any other high-level programming language 31. It is important to note that the automatic parallelization compiler system 1 can also be applied to the machine code 31 or assembly code 31 as the source code in order to achieve the parallelization of the code. In this case, the translation of the high-level language into machine code instructions does not need to be executed by the compiler system 10.
[0139] The parallel processing system 2 includes a memory unit 22, and the memory unit 22 includes at least a main execution memory unit 221 / 2212 including a plurality of memory banks for holding at least a part of the data of the processing code 32, and a transition buffer unit 221 / 2211 including a high-speed memory for storing the start position of the processing code 32 and a data segment including at least branch or jump instructions and / or used memory references and data values. Here, the main execution memory unit 2212 provides an access time slower than that of the transition buffer unit 2211. The transition buffer unit 2211 may include, for example, a cache memory module 2211 and / or an L1 cache 22121.
[0140] The execution of the processing code 32 by the parallel processing system 2 involves the occurrence of a latency time 26, which is given by the idle time of the processing unit 21 for obtaining and / or storing the data necessary for the execution of a specific block of instructions of the processing code 32 by the processing unit 21. The latency time can include, for example, the access time of the register 2211, and / or the access time of the L1 cache 22121, and / or the access time of the memory 2213, and / or the I / O operation time, and / or the data network transfer time, and / or the processor configuration time.
[0141] The compiler system 1 includes a parser module 11 for translating the source programming language 31 into the code 32 of basic instructions that can be directly executed by the processing units that can be executed by the several processing units 21. The basic instructions can be selected from a set specific to the processing units of basic instructions including arithmetic operations 321 and / or logical operations 322 and / or control operations and / or I / O operations, in particular variable and array declaration instructions 323, comparison operation instructions 324, and code flow instructions 325. The arithmetic operations 321 can include, for example, operations of addition, subtraction, multiplication, and division. The logical operations 322 can include, for example, several logical expressions such as equal, not equal, greater than, less than, greater than or equal to, less than or equal to. The control operations can include, for example, "branch expressions" and / or "loop expressions".
[0142] As a modification of the embodiment, at least two of the processing units can have, for example, different sets of basic instructions. Different processing units having different sets of basic instructions can include, for example, a central processing unit (CPU) 210, a graphics processing unit (GPU) 211, a sound chip 212, a vision processing unit (VPU) 213, a tensor processing unit (TPU) 214, a neural processing unit (NPU) 215, a physical processing unit (PPU) 216, a digital signal processor (DSP) 217, a synergistic processing unit (SPU) 218, a field programmable gate array (FPGA) 219, etc.
[0143] The parser module 11 includes means for splitting the code of the basic instructions into computational block nodes 333, and each computational block node consists of the smallest possible segmentation of a non - further decomposable unit, each containing a sequence of basic instructions that require the same input data. Two or more computational block nodes 333, each having a chain of basic instructions, form a computational chain 34 that creates an ordered flow of operations / instructions on the input data. The chain (sequence of basic instructions) in the computational block node 333 is constructed according to fixed rules. The instructions are placed in positions within the chain 34 such that a new basic instruction reads after a basic instruction that writes a data point. This automatically forms a data - centric chain 34 and maps the necessary physically - limited read (READ) and write (WRITE) operations to the computational registers 2211, L1 cache 2212, network I / O, etc. of the CPU 21.
[0144] The compiler system 1 comprises a matrix builder 15 for generating several numerical matrices 151, …, 15i from a computation chain 34 depending on the latency time 26. In the graph from the computation chain 34, the dependencies of the system can be evaluated, but they cannot be simply decomposed into individual independent chains. Within the chain 34, there are situations of fusion and splitting of computation block 333 chains resulting from data dependencies and / or code branches. The latency for distributing information in a hardware system is introduced as physical time. By assigning at least this time interval length to each computation block node 333, each computation block node 333 in the graph or tree structure is numbered according to its position in the graph, given a block number, and thus, the same number can be given to the computation block nodes 333 having the same “time position” in the graph. The computation block node 333 must be at least as long as the time required to distribute information in the system, and as many instructions as possible must be computed during this time. If each computation block node 333 has a number based on its position in the graph depending on the program flow, a set of matrices 151, …, 15i can be constructed.
[0145] The matrices 151, …, 15i also show that they can be mathematically captured (graph models based on CB, CC, etc. are not simply captured as tables / matrices). Thus, data manipulation and communication are given to construct a numerical matrix, which can be optimized (changed) by using, for example, ML or AI depending on the target platform and / or hardware setup.
[0146] The compiler system comprises a numerical matrix optimization module 16 that uses a numerical matrix optimization technique to minimize the overall latency time as the total latency time 26 by providing an optimized structure of a calculation chain 34 processed by a plurality of processing units 21. Here, the code generator 17 generates optimized machine code for a plurality of processing units of a parallel processing system having an optimized overall latency time 26. The optimization can now be applied to the hardware infrastructure. The quantities important for the optimization are known in numerical form in matrices 151, …, 15i for each time unit and for each independent chain and branch. For example, from the content of the matrices, (i) the number of basic instructions that must be sequential, (ii) the size of the data transfer from a calculation block x within a calculation chain u, and when this transfer is again required at a calculation block y within a calculation chain v (where y > x) (for example, whether it is possible via a network or the calculation blocks are combined to keep the data on the same cache line). As an example, for a GPU, the data should be copied from memory to GPU memory in one process step, and then all basic instructions with the same characteristics should be executed at once, but on the CPU, operations with the same data should be on the same cache line (CPU-dependent), or operations of a specific data type can be calculated on the corresponding CPU with a better instruction set than necessary. In the system of the present invention, since the basic elements are grouped into sequential groups within the CB, this always results in parallel code even without optimization.
[0147] To summarize, since the (micro)processor understands only basic instructions, the source code 31 is split into these basic instructions in order to achieve the most basic level of parallelization. (Note that the present invention is also applicable to the technical problem of optimizing a (micro)processor based on the principle of an integrated circuit (IC), which is a set of electronic circuits. The above instructions are linked to the configuration of the electronic circuits on the (micro)processor, and thus the following topics are applicable to any form of integrated circuit or, conversely, can be used to derive an optimized integrated circuit (or the configuration of electronic circuits or direct electronic circuits) for a given code. This is because the instructions can be regarded as the form of the configuration of electronic circuits representing computational operations (such as +, -, data manipulation, etc.).) The following key points are the key to the system of the present invention: (1) The basic instructions of the processor system are combined according to their respective "READ" and "WRITE" behaviors, that is, in the chain of nodes and the link to the calculation block 333, a chain of instructions is formed according to the rule that "the instruction writes to X1 and the new instruction that reads from X1 is appended after the last instruction that writes to X1". (2) When it is necessary to propagate information / data in the multiprocessor system 2, a new calculation block 333 starts. (3) Each calculation block 333 has a minimum time length. This is proportional to the length of the time (latency) required to propagate information (data or signal) to / from the block 333 in the hardware system. (4) In the case of a graph model having two read data nodes, instruction nodes, and one write data node connected by a link, the chain from the calculation block 333 has two chains at the places where (a) they merge (for example, the instruction reads from two data points written in two different calculation blocks 333 or there is a branch), (b) they occur (for example, when two calculation blocks 333 can be started by reading simultaneously).If necessary, the graph model can be based on, for example, more than two read nodes and some write nodes, or can combine several instructions in one operation / instruction node. (5) These chains 34 can be decomposed, and the instructions and necessary information transfers for discrete time intervals can be captured in a matrix. For example, each row is divided into columns (one column for each time interval), contains independent instruction chains, and the necessary information is transferred to other instruction chains. Thus, these are tangible for an automatic optimization process, especially a numerical optimization process. This provides a basis for the fully automatic parallelization of source code. In this automatic parallelization system 1, the graph model not only is based on a simple matrix or table representation, but rather, provides a multi-dimensional nested tree structure of a computational block graph as a task graph associated with the computational parallel chains 34, allowing the system 1 to evaluate the characteristics that can be used for automatic parallelization, code optimization, computational block 333 scheduling, and even automatic cost estimation or automatic mapping to different architectures of a multiprocessing system 2.
[0148] As described above, in the computation matrix, a cell contains a chain of instructions 34 given by a sequence of computation block nodes 333 forming a chain 34, and in the transfer matrix, a cell contains transfer characteristics, i.e., the necessary transfers to and from other computation block nodes 333. Note that the transfer characteristics include what information is required at another computation block node 333. Depending on the level of the target infrastructure, this can be resolved by classical compilation / controlled by the processor = transferred in the cache (classical compilers allocate data to registers and caches), or explicitly shared by shared memory, guarded by locks, or explicitly transmitted and received between different nodes in a cluster using, for example, the Message Passing Interface (MPI) protocol, or transmitted and received by sockets in multi-core Inter-Process Communication (IPC), etc. The transfer characteristics can have any form of communication, from that processed by the processor (cache) to the explicit ones by communication patterns (e.g., IPC via queues, MPI, etc.). This makes the method of the present invention scalable for a wide range of platforms and / or infrastructures. Thus, the transfer characteristics can include, for example, information as transmission data %1 (integer) to cell (1,7). Transfer and / or communication typically require a certain time, which is directly dependent on the (transfer / communication) latency time in a particular system. Finally, as described above, in a particular row of the computation matrix, column cells contain a flow or sequence of instructions, and each column cell of a row contains one of the computation block nodes 333 of the sequence of computation block nodes 333, and note that they form a chain 34 for a particular row. However, in a variant of a particular embodiment, each cell of a row does not necessarily have to contain a sequence or instruction of instructions. One or more of the cells of one or more particular rows of the computation matrix may be empty. This is also the case for the transfer matrix. The computation matrix and the transfer matrix are usually equal in their size, i.e., the number of rows and columns.Computation matrices, matrices, and transfer matrices provide possible technical structures for automatic parallelization.
[0149] Technically, a matrix builder connects computational block nodes 333 according to a program flow (e.g., FIG. 35) and / or code 32, or from a parser module 11. The program flow is represented by connecting computational block nodes 333 (similar to basic blocks), for example, when adding a new instruction (= operation node), a new computational block node 333 can be generated, where generally there is no explicit dependency (see FIG. 21), and the new computational block node 333 is placed after another computational block node 333 or under a branch node. After adding all the instructions, it is possible to number each computational block node 333 along the program flow that forms the sequence of chains 34 thus generated. These numbers (block numbers) are equal to the columns of the matrix. Computational block nodes 333 having the same block number are distributed in different rows in the same column of the computation matrix.
[0150] Application of the present invention to known technical problems The following shows how the system and method of the present invention derive a technical solution to a well-known technical problem in the field of parallel processing.
[0151] (i) Calculation of Fibonacci sequence To generate and calculate a Fibonacci series, different parallel processing implementations with different performance characteristics are known. The following shows how the present invention is applied to implementations using a) recursive function calls and b) loops. Both implementations exhibit different performance. To understand the processing problem, for example, https: / / www.geeksforgeeks.org / program-for-nth-fibonacci-number / may be referred to.
[0152] The following shows an example of processing code for generating a Fibonacci sequence using recursion. / / Fibonacci sequence using recursion #include<studio.h> int fib(int n) { if (n<=1) return n; return fib(n-1)+fib(n-2); } int main () { int n = 9; printf("%d", fib(n)); getchar(); return 0; }
[0153] Furthermore, the following shows how to give a pseudo-token language by parsing the function code of the fib(n) declaration. The function code of fib(n) is given as follows. int fib(n) { if (n<=1) return n; return fib(n-1) + fib(n-2); } The pseudo-token language is given by Table 1 below.
Table 1
[0154] Figures 41 and 42 schematically show a computational block node (CB) having operations and corresponding data nodes (similar to the tokens in Table 1 above). The icon (outer 1) JPEG0007708369000002.jpg5115 indicates a transfer to a position (open or end point of the computational block node) within the computational block node. As seen in Figure 42, the recursive call of the function results in additional transfers.
[0155] The next step is to number the computational block nodes depending on the call positions within the code. This results in a pseudo-graph as schematically represented in FIG. 43, which in turn results in computational and transfer matrices as shown in FIGS. 44 and 45.
[0156] By linking the start communication cells and end communication cells in the transfer matrix, removing the empty cells in the computational matrix, and returning them to different code segments, the result of FIG. 46 is obtained. Based on this code segment, the code can be directly generated (as a compiler) or transferred back to the code and then machine code can be generated using a SOTA compiler (as a transpiler).
[0157] To show how the system and method of the present invention provide advantages in the compilation of a recursive implementation according to the Fibonacci series, the following section explains how the method of the present invention maps and / or optimizes the code with a more parallel solution than the input code. As shown for optimizing the code, it is a combination between rows, and the graph is shown together with the combination of cbn at the branch nodes marked as branch 2b (see FIG. 43), and they are combined (see FIG. 48). Calling a function, in the present method, is to place the computational block nodes at the correct positions within the matrix in order to apply the corresponding transfers as shown in FIG. 47. Keeping this in mind, a recursive call of a function, in the method of the present invention, can be seen as the transfer of result variables and function parameters in the return statement as well as "reads" and "writes" (see FIG. 48).
[0158] According to the example of fib(4), FIG. 49 shows step by step how the reduction of additional computational block nodes and transfers (since all are on one computational chain) results in a more optimized source code.
[0159] The depth of this chain directly depends on the number n in fib(n). Since recursive calls can be very easily detected in the code, it is easy not to perform recursive calls in full dimension in the final application. For ease of understanding, Figure 50 shows this in a somewhat simplified manner. Step 4 shows the cbn for the call at "n = 4".
[0160] The next step in Figure 51 shows the transfer to be made (transport the information of the last "write" to the data node to the place where the "read" of the data node is performed and remember this information in the corresponding computational block node). Since all computations are performed on one chain, as seen in Figure 52, the transfer disappears. When all steps are resolved, a program in the form of Figure 53 is obtained.
[0161] Converting this back to code using a row-column representation, it can be seen that the resulting code is obtained. This results in more efficient code than the original fib(n = 4) implemented by Table 1 above. By compiling this code with a SOTA compiler respectively, better optimized code is obtained than when this method is not applied. f(n) { if (n <= 1): return n; t1 = n-1 t2 = f(t1) t3 = n-2 t4 = f(t3) t5 = t2+t4 return t5 } main() { r1w = f(1) r0w = f(0) r2w = r1w+r0w r3w = r2w+r1w r4w = r3w+r2w }
[0162] Recursive calls transfer information (function parameter param[i]) to the corresponding cbn within the branch node of the function declaration and then transfer the return value(s) back to a[i], so it can also be interpreted as an array operation to which the method of the present invention is applied (see Fig. 54). This leads to a perspective also seen in partial differential equations. This form of implementation of the Fibonacci series is further referred to below. However, as a next step, the following paragraphs show how to handle partial differential equations that result in mostly nested loops and extensive use of array operations.
[0163] (ii) Partial Differential Equation (PDE) According to the present application, by arranging the operation nodes in the dependencies of the "read" and "write" patterns and resolving ambiguous dependencies by transfer, it is possible to derive a calculation and communication model for a discretized implementation of a PDE given freely. The PDE of the 2D heat equation is used here, and the 2D heat equation is [Number] given by
[0164] Using the finite difference method, [Number] discretize it to
[0165] Fig. 55 shows a part of the implementation in Python. Each entry in the array is a data node. Reading from the array index is, in this method, an operation node having the data node of the index and the base address of the array. Array operations can be seen as operation nodes having the corresponding data nodes (see Fig. 56).
[0166] For an example of a loop array with the array operation a[i+Δiw]=a[i], Figure 57 shows the corresponding computational block nodes. Keeping this in mind, the initial block (Figure 55) can be represented in detail in a graph like Figure 58. All calculations in the computational block (see Figure 55) are performed in the j loop. By using 1D array notation and applying the basic rules of the method, a schematic form for the computational block nodes within the loop is derived as shown in Figure 59.
[0167] Array reads (e.g., u[k][i+1][j]) create computational block nodes, and metadata for transferring values to this index is added to both cbn. One is the "read" node, which is what this data node was last written to (e.g., at a[k+1][i][j]). This leads to the fact that statements like the array operation in the j loop (Figure 55) result in five computational block nodes representing the "read" operation of the array, then a computational block node for calculating the arithmetic solution, and then a computational block node for writing to the array at position [k+1][i][j]. This representation is a schematic view to show more clearly how the method takes into account such array operations. This leads to the situation shown in Figure 60.
[0168] One of the most fundamental principles of the proposed method is to find the last operation node A that "writes" to data node B, and then place the new operation node C that "reads" from B after operation node A. If it is not an explicit 0-dependency or 1-dependency, add transfers to the computational block nodes that include operation nodes A and C respectively. Thus, a[i1]=a[i2] is a "read" of the data node with base address "a" and index "i2", as well as a "write" to the data node with base address "a" and index "i1" (see Fig. 56). Thus, each loop creates a new computational block node with a "read" or "write" operation node and the corresponding transfer. It can be derived from the following scheme (Fig. 61).
[0169] The loop can be seen as in Fig. 61, showing that each loop pass is a new computational block node on a new computational chain (a row in a matrix). These computational block nodes are concurrent (resulting in the same block or segment number in the numbering process step) as long as there are no transfers, as the empty computational block nodes and the entries in the transfer matrix disappear. If a transfer occurs, the cbn numbers are different and they cannot be computed in the same step. The offsets at each index of the "read" and "write" operations can result in transfers between the computational block nodes that include each of the "read" and "write" operation nodes, as seen in Fig. 62.
[0170] The dependencies in the "reading" and "writing" of the index within the loop lead to the conclusions of Figure 63. If ΔI is less than 0, it indicates that the loop can be registered in as many computational block nodes as the length of the loop. The length means the number of sequential iterations defined by the loop definition, for example, the start value "i0", the maximum value "Ai", and the increment value "Bi" in "for i = i; i <> Ai; Bi". If ΔI is greater than 0, only ΔI computational block nodes can be executed in parallel. If a number less than this is used, a transfer occurs. If the loop is resolved on more than ΔI units, the computational block nodes are not parallel (in the sense of having the same number in the flow graph). Thus, ΔI is a unique number for determining how the "reading", which uses the information within the array in the index but does not change it, can be distributed across different computational units.
[0171] This has several important consequences. It is possible to extract a single one-dimensionality from nested loops and, for each "read" to an array, determine which of these reads leads to a "write" within the loop and which values are "reads" from the data nodes prior to the loop. Thus, for each "read", it can be derived whether there is a transfer to a computational block node (= computational chain) of the "write" part. This can be implemented as a model, bringing a new and fairly general perspective on arrays within nested loops. Further, according to the concept of FIG. 63, since there is no transfer between the "read" data node and the "write" data node, it can be derived which of the nested loops can be resolved, i.e., computed in parallel in this method. For the 2D heat equation using the central difference method, this results in the dependencies shown in FIG. 64. This means that the i and j loops can be resolved. This is because ΔI is always less than or equal to the number of sequential iterations in the i and j loops, but not in the k loop. The fact that the i and j loops can be resolved in this method means that the "reads" and "writes" in the array are distributed among computational block nodes that can be executed in parallel, while the k loop must be sequentially iterative. And after each k loop, a transfer between units is performed. At this point, it should be noted that the method eliminates transfers in the mapping / optimization step if they are on the same unit.
[0172] In most implementations for solving PDEs, the computation is only done for a subset of the array (e.g., by treating boundary conditions differentially), which makes this step a bit cumbersome. In the example of two nested loops, as shown in FIG. 65, for each loop nest, rules can be derived for handling the results for the gaps and "transfers" in the proposed model. Constant values are either appropriate for the operations prior to the loop within the "computational block" (see FIG. 55) or values from the computation.
[0173] By incorporating gaps that occur by looping over subsets of an array, a transfer pattern of the discretized equations can be derived, and depending on the gap size and the size of the loop (= mesh in this case), a model for transfer is obtained, which represents the computational block nodes of the method as shown in FIG. 66.
[0174] Shown using a very small example of nX = 5 and nY = 4, this results in a computation and transfer matrix as shown in FIG. 67. In the transfer matrix, the arrow “->” indicates obtaining information, and “<-” indicates sending data to other computational block nodes.
[0175] The method of the present invention extracts all necessary transfers / communications between grid elements within the grid required for the discretized PDE. FIG. 68 shows the transfers occurring between units. Each unit is represented by one row in the matrix. In FIG. 68, the extraction of the transfer and computation matrices can be seen, with dark gray being the “receiving” side of the transfer and light gray being the “sending” part of the transfer. It is important to note in this regard that this does not mean that it has to be sent and received. It can be protected by barriers or locks or disappearances since, for example, it is shared by a shared memory segment, shared by a cache, and thus processed by the CPU. It depends on the transfer / communication mechanism used. This can each result in code in a form as shown in FIG. 69 being transferred to the model. It must be noted that the computation between p0 and p1 evolves. This situation occurs because this method is defined to add transfers from the “written” data nodes to the “read” data nodes, and the array operations within the nested loops start from the “read” operation nodes.
[0176] Figure 70 shows how transfer can be used to create a time model. The gray areas are meta-values (known, for example, from variable type definitions or the like), which can also be used to map / optimize the matrix to a given hardware infrastructure. These values are not used in this example. To show possible optimization steps, a very simple model is fully decomposed and shown in Figure 71 below, where the combinations of rows (1,2), (3,4), and (5,6) result in three units that calculate the following:
Number
Number
Number
[0177] These steps illustratively show how the method generates different ratios between computation and communication, and how, depending on the combination, the transfer is reduced and the computation per unit increases. To illustrate this further, a very simple model of computation and communication can be used, and some "real values" for time can be assumed. Following the performance values for floating-point capabilities (FP) per cycle and its frequency, two generations of Intel CPUs provide power values for these to compute floating-point arithmetic (such as T in the model). We performed addition and multiplication of 4-cycle floating values using P5 FP32:0.5 and Haswell FP32:32, with frequencies for 66 MHz and 3.5 GHz cycles. To obtain values for Δtransfer representing communication latency in the network (the author acknowledges that many assumptions are made and latency is not the only key value for the network), two types of latency were used. That is, 1 ns for the InfiniBand type and 50 ns for the cache latency. The behavior in Figure 72 can be derived by Δt = combined cbn*cpu + number of transfers * network latency for a grid with nX = 2048 and nY = 1024.
[0178] It is always easy to obtain the code returned from this method or two matrices that is adapted to the available IPC options. This general form of code optimization brings innovation to the field of automatic code parallelization. For example, the code can be optimally mapped to, for example, an MPI cluster infrastructure, or it may be optimally mapped using some finer-grained tuning that applies a hybrid approach by combining MPI (shared memory) and threading local to the node. Different combinations (for example, combining the first x cbn in the i direction with the y cbn in the j direction) make it possible to obtain an optimal ratio between "computation" and "communication" for any given discretized PDE, but depending on the available hardware infrastructure. Since the required characteristics can also be tested or calculated for the available hardware infrastructure, this can be done without any manual interaction, solving the practical technical problems of the prior art.
[0179] (iii) Fibonacci using loops The Fibonacci source can also be implemented using loops. / / Fibonacci sequence using dynamic programming #include<studio.h> int fib(n) { / * Declare an array to store Fibonacci numbers * / int f[n+2]; / / One extra to handle the case of n=0 int 1; / * The 0th and 1st numbers in the sequence are 0 and 1 * / f[0] = 0; f[1] = 1; for (i=2; i<=n; i++) { / * Add and store the two preceding numbers in the sequence * / f[i] = f[i-1] + f[i-2]; } return f[n]; } int main() { int n=9; printf("%d", fib(n)); getchar(); return 0; }
[0180] According to the example using the 2D heat equation, this results in the results of FIG. 73, and as a result, source code of the following form is obtained: main() { a1 = f(1); b1 = f(0) for (i=0; i<=(n-1) / 2; i++) { c1 = a1+b1 a2= c1 b2 = b1 c2 = a2+b2 a1 = c2 b1 = c1 } } This has performance characteristics similar to those of the code of FIG. 53 shown above.
[0181] (iv) Resolution of pointer ambiguity Pointer ambiguity resolution is also a technical problem that has not been completely solved (see, for example, Runtime pointer disambiguation by P. Alves). Applying the method of the present invention, it can be shown that the method of the present invention solves the ambiguity resolution that occurs by passing a function parameter as a pointer. This is because the pointer is taken as information and this ambiguity resolution is solved in the step of transfer elimination as shown in FIG. 74.
[0182] Differences between the system of the present invention and an exemplary prior art system The various approaches known by the prior art are essentially different by virtue of their technical approaches. In the following, in order to further illustrate the system and method of the present invention, the essential differences from three of the prior art systems will be described in detail.
[0183] Sclicing based code parallelization for minimizing inter-process communication by M. Kandemir et al. (hereinafter referred to as Kandemir) discloses a method for scalable parallelization to minimize inter-processor communication in a distributed memory multi-core architecture. Using the concept of sequential iteration space slicing, Kandemir discloses a code parallelization method for data-intensive applications. This method targets a distributed memory multi-core architecture and formulates the problem of data calculation distribution (division) across parallel processors using slicing. Starting from the division of the output array, it sequentially determines not only the sequential iteration space of loop nests in the application code but also the partitions of other arrays. The goal is to minimize inter-processor data communication based on this problem formulation based on sequential iteration space slicing. However, Kandemir does this by sequentially determining the partitions of other arrays in the application code using the partition of the output array (see page 87). The method of the present invention disclosed herein is a different approach as it does not involve sequential iterative determination of the combined array portions. The method of the present invention obtains this information directly from the code. As disclosed on page 88 of Kandemir, program slicing was originally introduced by Weiser in a pioneering paper. Slicing means extracting from the program the statements that potentially affect a particular statement of interest that serves as the slicing criterion (see Advanced slicing of sequential and concurrent programs by J. Krinke). Slicing techniques exhibit similar effects to well-known point data / control dependencies (see Computer architecture by D. Patterson, pages 150 and below) and data flow analysis (e.g., [8]). Our method appears similar to these methods at first glance.This stems from the fact that data dependency is one of the most central points in programming and compilation, but the method of the present invention disclosed herein has a different perspective. Novelty, as shown by the report, lies in the fact that this new approach provides a new perspective, namely the definition of computational block nodes. The method of the present invention explicitly takes each variable as unique information (represented by a specific bit pattern), and since computational instructions (= statements) can only be executed using available bit patterns (in an accessible storage device, such as a CPU register), the method of the present invention extracts the time between the modification and use of a variable at a certain "location" (such as a register). Known state-of-the-art compilers neither extract this time from the (source) code nor focus on a single data entity (bit pattern) within a statement, and all known prior art systems always focus on using the result of the complete statement. As disclosed by Kandemir (page 89), sequential iteration space slicing can be used to answer questions such as "which sequential iteration of which statement can affect the values of a given set of elements from array A". This shows a different perspective. The method of the present invention derives this very effect in a general form by finding the "reads" and "writes" for each element of an array and correlating this latency of "reads" and "writes" with the time required to transfer the elements. The method of the present invention discovers this "effect" and does not attempt to discover it by "simulating" its interaction with linear algebra methods or by sequential iterative methods. On page 91, Kandemir summarizes a function that returns a set of loop sequential iterations assigned to processor p from loop nest s. Here, Zp,r is the set of data elements accessed by processor p from array Ar. In contrast, the system and method of the present invention are not based on this approach.The method and system of the present invention first derive all dependencies of the code (even within nested loops), then create a matrix for all code in the program, and finally map / optimize all operations of the program by combining the matrices. The method of Kandemir derives by creating matrices for array loop-index dependencies and assigning them to processors, then takes the Presbuerger set, then finds out how sequential and iterative this approach needs to be (see Kandemir page 92) by generating code as output "(a series of potentially nested loops)", and then iterates sequentially over the unknowns. The approach of Kandemir does not consider the hardware specifications either.
[0184] (b) Automatic Parallelization by A. Fonseca: Executing Sequential Programs on a Task - Based Parallel Runtime (hereinafter referred to as Fonseca) discloses another prior - art system for automatically parallelizing sequential code in modern multi - core architectures. Fonseca discloses a parallelizing compiler that analyzes read and write instructions in the program, as well as control - flow modifications, in order to identify the set of dependencies between instructions in the program. Subsequently, the compiler rewrites the program and organizes it into a task - oriented structure based on the generated dependency graph. Parallel tasks are composed of instructions that cannot be executed in parallel. A work - stealing - based parallel runtime is responsible for scheduling and managing the granularity of the generated tasks. Additionally, a compile - time granularity control mechanism also avoids the creation of unnecessary data structures. Fonseca focuses on the Java language, but this technology may also be applicable to other programming languages. However, in contrast to the method of the present invention disclosed in this application, in Fonseca's approach, in order to automatically parallelize a program, it is necessary to analyze the memory accessed to understand the dependencies between parts of the program (see Fonseca, page 6). The method of the present invention disclosed herein is different. Fonseca uses data groups and memory layouts and then checks for dependencies. Since Fonseca explicitly maintains task parallelism, this is not in the same scope as the method of the present invention. For example, in the Fibonacci example, as described above, the cost of creating a new task is higher than the cost of executing the method for low input numbers (see Fonseca, page 7). This indicates that this is not the same approach. The method of the present invention handles this example in a completely different way. On the other hand, this also directly demonstrates the technical problems that can be solved by the method of the present invention. Additionally, Fonseca (see page 9) has to define the main requirements for the location of future creations.In contrast, the method of the present invention knows where to place each instruction or statement depending on the data dependency of a single piece of information / variable. Fonseca (see page 9) also discloses that algorithm 18 must be used to find the best position for generating the future. In contrast, the method of the present invention accurately places the instructions in that position based on the new scope. In Fonseca (see page 10), commutative and associative operations are used. The method of the present invention excludes, for example, division (used in many mathematical models), so its approach is not based on this form. This limitation is also typically for the map-reduce approach (described herein).
[0185] (c) U.S. Patent Application Publication No. 2008 / 0263530 discloses a system for converting application code into optimized application code or execution code suitable for execution on an architecture having at least first and second levels of data memory units. This method involves obtaining application code, which includes data transfer operations between levels of memory units. This method further includes converting at least a portion of the application code. The conversion of the application code includes scheduling data transfer operations from a first level of memory unit to a second level of memory unit such that accesses to data accessed multiple times are brought closer in time than in the original code. The conversion of the application code further includes determining the layout of data in the second level of memory unit to improve data layout locality such that data accessed closer in time is also brought closer in layout than in the original code. US2008 / 0263530A1 allows for improving layout locality (see column 4, paragraph 0078 of US2008 / 0263530A1). In contrast, the method of the present invention disclosed herein has a different scope. The method of the present invention finds this form of "locality" and orders the instructions such that instructions that must be "local" are grouped. In that case, this inherent form of the code is arranged in a matrix, in which case general non-sequential iterative ones can each exist by combining elements, in which case an optimal mapping to a given hardware can be set, which always results in a simultaneous parallel form of the code. Sequential iterative solutions may miss the solution and end with ambiguous results. Further, US2008 / 0263530A1 (page 6, paragraph 0092) discloses that it can be regarded as a complex non-linear problem for which a reasonable, near-optimal, and scalable solution is provided.In contrast, the method of the present invention does not result in a non-linear problem, but rather prevents obtaining a complex non-linear problem that requires a sequential iterative numerical solution method / optimization approach / algorithm. US2008 / 0263530A1 (page 6, paragraph 0093) believes that its access locality is improved by calculating reuse vectors and applying them to find an appropriate transformation matrix T. In contrast, the method of the present invention disclosed herein does not require a transformation matrix nor the calculation of such a transformation matrix. The method of the present invention reads the inherent logical connection between array operations based on a given piece of code (in the form of a loop definition or loop block in the compiler IR language, or a jump definition in the assembly code). Finally, US2008 / 0263530A1 (page 6, paragraph 0093) claims that after fixing T, the layout M for the arrays accessed within the loop nest is fixed, and that layout has not yet been fixed. This discloses a sequential iterative numerical solver-based approach. The method of the present invention disclosed herein reads this dependency within the code without using a sequential iterative solution method that uses linear algebra methods to find the solution to a system of simultaneous equations.
Explanation of Signs
[0186] List of Reference Signs 1 Compiler System 11 Lexer / Parser 12 Analyzer 13 Scheduler 14 Computational Block Chain Module 15 Matrix Builder 151,…,15i Matrices 16 Optimization Module 17 Code Generator 2 Parallel Processing System / Multiprocessor System 21 Processor Unit 210 Central Processing Unit (CPU) 2101 Control Unit 2102 Processor 21021 Register 21022 Combinatorial Logic 211 Graphics Processing Unit (GPU) 212 Sound Chip 213 Vision Processing Unit (VPU) 214 Tensor Processing Unit (TPU) 215 Neural Processing Unit (NPU) 216 Physical Processing Unit (PPU) 217 Digital Signal Processor (DSP) 218 Synergistic Processing Unit (SPU) 219 Field Programmable Gate Array (FPGA) 22 Memory Unit 221 Primary Memory Unit 2211 Processor Register 2212 Processor Cache 22121 L1 Cache 22122 L2 Cache 2213 Random Access Memory (RAM) Unit 222 Secondary Memory Unit 2221 Hard Disk Drive (HDD) 2222 Solid State Drive (SSD) 2223 Universal Serial Bus (USB) Memory 2224 Flash Drive 2225 Optical Storage Device (CD or DVD Drive) 2226 Floppy Disk Drive (FDD) 2227 RAM Disk 2228 Magnetic Tape 223 Tertiary Memory Unit (Tape Backup, etc.) 23 Memory Bus 231 Address Bus 232 Data Bus 24 Memory Management Unit (MMU) 25 Input / Output (I / O) Interface 251 Memory-Mapped I / O (MMIO) or Port-Mapped I / O (PMIO) Interface 252 Input / Output (I / O) Channel (Processor) 3 Program Code 31 High-Level Language 311 C / C++ 312 Python 313 Java 314 Fortran 315 OpenCL (Open Computing Language) 32 Low-Level Language (Machine Code / Assembly Language) 321 Arithmetic Operation Instructions 322 Logical Operation Instructions 323 Variable and Array Declaration Operations 324 Comparison Operation Instructions 325 Code Flow Instructions / Memory Operations / I / O Operations 33 Nodes 331 Data Nodes (Holding Certain Data Values) 3311 Input Data Nodes for Reading (Datanode in (Read Access)) 3312 Output Data Nodes for Writing (Datanode out for Write Access) 332 Operation Nodes (Executing Operations [Calculations, Operations]) [Calculation Nodes, Operation Nodes] 333 Calculation Block Nodes (CB1, CB2,…, CBx) 334 Control Flow Nodes 3341 Branch Nodes 3342 Hidden Branch Nodes 3343 Loop Branch Nodes 335 Conditional Nodes 34 Chains 341 Calculation Chains (Chains of Calculation Block Nodes) 342 Operation Chains (Chains of Operation Nodes) 35 Latency time Δt (time to transfer data from one process to another) 351 Δt readwrite = Time between write access and read access to the data node 352 Δt computation = Time to compute all operation nodes in the computing block node 353 Total (overall) latency time Δt total 4 Network 41 Network controller instruction Instruction latency Latency Processor Processor Core Core Cache Cache non-blocking Non-blocking unit Unit compute Compute 2 units Two units computechain Computation chain computation Computation communication Communication minimal runtime Minimum runtime resultat Result operation Operation datanode Data node operation node Operation node source Source target Target variable Variable direction Direction start Start end End forward Forward direction backward Backward direction returnvalue Return value value value base base index index read read write write compare compare loopstate loop state computations computations target unit target unit True True False False Datenknoten data node Operationsknoten operation node type: branch transmission type: branch transmission type: fusion transmission type: fusion transmission type: split transmission type: split transmission type: b2b transmission type: b2b transmission
Claims
Claim 1 A compiler system (1) for optimizing compilation and automatic parallelization of program code (3) for execution by a parallel processing system (2) having a plurality of processing units (21) that simultaneously process instructions on data by executing the program code (3) in the parallel processing system (2), wherein the compiler system (1) has means for translating source program code written in a programming language into machine code (3 / 32) as target programming language by generating machine code (32) as parallel processing code that is executable by the plurality of processing units (21) of the parallel processing system (2) or that includes some instructions for controlling the operation of the plurality of processing units (21), execution of the processing code (32) by the parallel processing system (2) involves the occurrence of latency times (26), which are given by the idle time of the processing units (21) between sending data back to the parallel processing system (2) after a particular block of instructions of the processing code (32) has been processed on the data by the processing units (21) and receiving the data required for execution of successive blocks of instructions of the processing code (32) by the processing units (21), the compiler system (1) has a parser module (11) for translating the source programming language (31) into code (32) having a flow of basic instructions executable by the processing units (21), the basic instructions being selectable from a set specific to the processing units of the basic instructions, the basic instructions including basic arithmetic operations (321) and / or logical operations (322) and / or control operations (325) and / or memory operations (325) for the plurality of processing units (21), The parser module (11) has means for dividing the code (32) of the basic instruction into computational block nodes (333), each computational block node (333) consisting of the smallest possible segmentation of the non - further decomposable sequence of basic instructions of the code (32) that can be processed by a single processing unit (21). The smallest possible segmentation of the basic instruction is characterized by a sequence of basic instructions composed of consecutive read instructions and write instructions. The sequence cannot be further decomposed by a smaller sequence of basic instructions between consecutive read and write instructions. The read and write instructions receive the data necessary for the processing unit (21) to process the sequence of basic instructions and are required to send back data after processing by the sequence. The compiler system (1) has a matrix builder (15) for generating a computational matrix and transfer matrices (151,…,15i) from a computational chain (34) divided from the code (32). The computational chain (34) is formed by one or more computational block nodes (333) that create an ordered flow of computational block nodes (333). Each computational chain (34) is executed by one processing unit (21). The computational matrix includes the computational chains (34) of the processing units (21) in each row, and each column has the sequence of basic instructions of the computational block nodes (333) within the computational chain (34) of the row. Each cell of the transfer matrix includes transfer characteristics associated with the data transfer from one computational block node to a consecutive computational block node (333). The transfer characteristics include what transfers are required at the start and end of each corresponding cell in the computational matrix. The compiler system (1) has an optimizer module (16), and the optimizer module constructs different combinations of rows from the calculation matrix and the transfer matrix, each of the different combinations representing possible machine code (32) as parallel processing code, and also minimizes the overall latency time (26) that occurs as the combined latency time, using optimization techniques to provide an optimized combination as the structure of a calculation chain (34) to be processed by the plurality of processing units (21), and the code generator (17) generates optimized machine code with a minimized overall latency time (26) for the plurality of processing units (21) of the parallel processing system. Compiler system. **Claim 2** The compiler system according to claim 1, characterized in that the source programming language is a high-level programming language. **Claim 3** The compiler system according to claim 2, characterized in that the high-level programming language includes C and / or C++ and / or python and / or java (registered trademark). **Claim 4** The compiler system according to any one of claims 1 to 3, characterized in that the transition buffer unit includes a cache memory module and / or an L1 cache. **Claim 5** The compiler system according to any one of claims 1 to 4, characterized in that the latency time includes register access time, and / or L1 cache access time, and / or memory access time, and / or I / O operation time, and / or data network transfer time, and / or processor configuration time. **Claim 6** The compiler system according to any one of claims 1 to 5, characterized in that at least two of the processing units have different sets of basic instructions. **Claim 7** The compiler system according to any one of claims 1 to 6, characterized in that the arithmetic operations include addition, subtraction, multiplication, and division operations. **Claim 8** The compiler system according to any one of claims 1 to 7, characterized in that the logical operation includes several logical expressions such as equal, not equal, greater than, less than, greater than or equal to, less than or equal to, etc.
9. The compiler system according to any one of claims 1 to 8, characterized in that the control operation includes a "branch expression" and / or a "loop expression".
10. The compiler system according to any one of claims 1 to 9, characterized in that the calculation block node (333) consists of connected operation nodes (332), and the operation nodes (332) are configured in one chain (34) having one or more input data nodes (331 / 3311) and one or more output data nodes (331 / 3312).
11. The compiler system according to claim 10, characterized in that the calculation block node (333) is connected to the next calculation block node (333), or to the control flow node (334), or to one or more input data nodes (331 / 3311) and one or more output data nodes (331 / 3312) that construct the chain (34).
12. The compiler system according to any one of claims 1 to 11, characterized in that each pair of the calculation matrix and the transfer matrix is realized so as to form a single numerical matrix and / or a multi-dimensional numerical tensor respectively.
13. The compiler system according to claim 12, characterized in that the numerical matrix optimization is realized as tensor optimization.
14. The compiler system according to claim 12 or 13, characterized in that the numerical matrix optimization or tensor optimization by the optimizer module is based on a dedicated applied machine learning structure.
15. Each calculation block node (333) of the calculation chain (34) is numbered along the ordered flow of the calculation block nodes (333) within the calculation chain (34) by assigning a block number to the calculation block node (333). In the calculation matrix and the transfer matrix, the block number of the calculation block node (333) is equal to the column of the matrix, and the calculation block nodes (333) having the same block number are distributed in different rows of the same column. The compiler system according to any one of claims 1 to 14, characterized in that.
16. In the calculation matrix, each row includes a calculation chain (34) of the processing unit (21). The cells of the calculation matrix include the calculation block nodes (333) of the calculation chain (34). In the transfer matrix, the cells include the necessary transfer characteristics to and / or from other calculation block nodes (333). The compiler system according to any one of claims 1 to 15, characterized in that.
17. The transfer matrix having cells with the transfer characteristics includes at least what information is required at another calculation block node (333). The compiler system according to claim 16, characterized in that.
Citation Information
Patent Citations
Scheduler, processor system, program generating method, and program generating program
JP2010108153A
Program, code generation method and information processing apparatus
JP2013206289A
System and method for job scheduling optimization
US20130191843A1