Simulation program generation methods, simulation testing methods and related equipment

By acquiring the compilation information of the target hardware and using heterogeneous and simulation dialect operations to automatically generate simulation programs, the problem of low simulation program development efficiency caused by hardware adaptation is solved, and efficient simulation program generation and testing are achieved.

CN120315689BActive Publication Date: 2025-10-28INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510798633.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-28
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In existing technologies, due to the different compatibility between hardware and simulation programs, programmers need to manually compile and optimize the code, resulting in low efficiency in simulation program development.

Method used

By acquiring the compilation information of the target hardware, the preset MLIR code is encapsulated into heterogeneous and simulation dialect code blocks using start and stop operations, mapped to MLIR code for the target hardware, and converted into a simulation program for the simulation platform, thus achieving automated generation.

Benefits of technology

It improves the development efficiency of simulation programs, simplifies the code adaptation process, reduces development complexity and error rate, and enhances code flexibility and portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315689B_ABST
    Figure CN120315689B_ABST
Patent Text Reader

Abstract

This invention provides a simulation program generation method, a simulation testing method, and related equipment, relating to the field of code generation. The simulation program generation method includes: obtaining compilation information associated with the target hardware in a configuration file; using start and stop operations, encapsulating the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks to obtain MLIR code containing heterogeneous and simulation dialects; mapping the MLIR code containing heterogeneous and simulation dialects to MLIR code oriented towards the target hardware; and converting the MLIR code oriented towards the target hardware into a simulation program oriented towards the simulation platform. The simulation program generation method provided by this invention, based on the compilation information associated with the target hardware, converts the preset MLIR code into a simulation program oriented towards the simulation platform through dialect injection and mapping operations, thereby realizing the automated generation of simulation programs corresponding to the target hardware and improving the efficiency of simulation program development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of code generation technology, and in particular to a simulation program generation method, a simulation testing method, and related equipment. Background Technology

[0002] With the development of technology, simulation programs developed using programming languages ​​need to run on various hardware, especially in scenarios requiring high performance and near-hardware-level operation. Programming languages ​​are widely used in system software, embedded systems, and other fields due to their powerful performance and flexibility.

[0003] However, the degree of compatibility between hardware and simulation programs varies. Programmers typically need to perform heterogeneous code translation for each piece of hardware to obtain a simulation program compatible with that hardware. For example, in wafer-level chip (WLS) applications, multiple computing clusters are included, and each cluster may contain various computing units such as CPUs (Central Processing Units) and GPUs (Graphics Processing Units). Therefore, in code compilation applications, traditional compilation methods usually require programmers to manually compile and optimize code for the target hardware, resulting in low efficiency in simulation program development. Summary of the Invention

[0004] This invention provides a simulation program generation method, a simulation testing method, and related equipment to solve the problem of low simulation program development efficiency caused by requiring programmers to manually compile and optimize code for the target hardware in the prior art. It realizes the automated generation of simulation programs corresponding to the target hardware, effectively improving the simulation program development efficiency.

[0005] This invention provides a method for generating simulation programs, comprising the following steps.

[0006] Retrieve compilation information associated with the target hardware from the configuration file;

[0007] By using start and stop operations, the code blocks corresponding to the compilation information in the preset MLIR code are encapsulated into heterogeneous and simulation dialect code blocks, resulting in MLIR code containing heterogeneous and simulation dialects;

[0008] Map MLIR code containing heterogeneous dialects and emulated dialects to MLIR code oriented towards the target hardware;

[0009] Convert MLIR code for target hardware into simulation programs for the simulation platform.

[0010] According to a simulation program generation method provided by the present invention, the compilation information includes kernel function names, computing clusters contained in the target hardware, and the task load of computing units contained in the computing clusters.

[0011] The code blocks corresponding to the compilation information in the preset MLIR code are encapsulated into heterogeneous and simulation dialect code blocks, resulting in MLIR code containing heterogeneous and simulation dialects, including:

[0012] Traverse the preset MLIR code and identify the kernel function code block in the preset MLIR code that matches the kernel function name;

[0013] By using start and stop operations, the kernel function code blocks in the preset MLIR code are encapsulated into heterogeneous and simulation dialect code blocks;

[0014] The corresponding parameters in the heterogeneous and simulated dialect code blocks are assigned values ​​using the task load of the computing cluster and computing units, resulting in MLIR code containing both heterogeneous and simulated dialects.

[0015] According to a simulation program generation method provided by the present invention, MLIR code containing heterogeneous dialects and simulation dialects is mapped to MLIR code oriented towards target hardware, including:

[0016] Identify parallel task operation sequences in heterogeneous and simulated dialect code blocks;

[0017] Based on the values ​​of the corresponding parameters in the heterogeneous and simulation dialect code blocks, the parallel task operation sequence is converted into a cyclic execution operation sequence for different computing clusters;

[0018] The cyclic execution sequence of operations for different computing clusters is converted into specific implementations on different computing units of the corresponding computing clusters, resulting in MLIR code for the target hardware.

[0019] According to a simulation program generation method provided by the present invention, the computing units include a CPU and a GPU; the method converts MLIR code for the target hardware into a simulation program for the simulation platform, including:

[0020] Extract the MLIR code for the GPU in the target hardware from the MLIR code for the target hardware, and generate a PTX file based on the MLIR code for the GPU in the target hardware.

[0021] and,

[0022] Extract the MLIR code for the CPU and GPU in the target hardware from the MLIR code for the target hardware, identify the operation sequences in the MLIR code for the CPU and GPU in the target hardware, convert the operation sequences in the MLIR code for the CPU and GPU in the target hardware into interface functions for the simulation platform, and obtain the simulation program for the simulation platform.

[0023] According to a simulation program generation method provided by the present invention, the operation sequences in the MLIR code for the CPU and GPU in the target hardware are converted into interface functions for the simulation platform to obtain a simulation program for the simulation platform, including:

[0024] By utilizing simulation dialects and operations, MLIR code for the target hardware (CPU and GPU) is converted into MLIR code for the simulation platform.

[0025] The MLIR code for the simulation platform is subjected to a first descent transform to obtain the LLVM intermediate representation;

[0026] Perform a second descent transformation on the LLVM intermediate representation to obtain the simulation program for the simulation platform.

[0027] The present invention also provides a simulation testing method, comprising:

[0028] A simulation platform is constructed based on the structural information of the target hardware.

[0029] The simulation program is run on the simulation platform to obtain the simulation test results of the target hardware; the simulation program is generated using the simulation program generation method provided by this invention.

[0030] The present invention also provides a simulation program generation apparatus, comprising:

[0031] The acquisition module is used to obtain the compilation information associated with the target hardware from the configuration file;

[0032] The generation module is used to encapsulate the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks by using the start and stop operations, so as to obtain MLIR code containing heterogeneous dialect and simulation dialect.

[0033] The mapping module is used to map MLIR code containing heterogeneous dialects and simulation dialects to MLIR code oriented towards the target hardware;

[0034] The conversion module is used to convert MLIR code for the target hardware into simulation programs for the simulation platform.

[0035] The present invention also provides a simulation testing device, comprising:

[0036] The building block is used to construct a simulation platform based on the structural information of the target hardware.

[0037] The simulation test module is used to run a simulation program on a simulation platform to obtain simulation test results of the target hardware; the simulation program is generated using the simulation program generation method provided by this invention.

[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the simulation program generation methods or simulation testing methods described above.

[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the simulation program generation methods or simulation testing methods described above.

[0040] The simulation program generation method provided by this invention converts preset MLIR code into a simulation program for the simulation platform through heterogeneous dialects and operations, and simulation dialects and operations, based on the compilation information associated with the target hardware. This achieves automated generation of simulation programs corresponding to the target hardware and effectively improves the efficiency of simulation program development. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0042] Figure 1 This is one of the flowcharts illustrating the simulation program generation method provided by the present invention.

[0043] Figure 2 This is the second flowchart of the simulation program generation method provided by the present invention.

[0044] Figure 3 This is the third flowchart of the simulation program generation method provided by the present invention.

[0045] Figure 4 This is the fourth flowchart of the simulation program generation method provided by the present invention.

[0046] Figure 5 This is the fifth flowchart of the simulation program generation method provided by the present invention.

[0047] Figure 6This is one of the flowcharts of the simulation testing method provided by the present invention.

[0048] Figure 7 This is the sixth flowchart of the simulation program generation method provided by the present invention.

[0049] Figure 8 This is the second flowchart of the simulation testing method provided by the present invention.

[0050] Figure 9 This is a schematic diagram of the simulation program generation device provided by the present invention.

[0051] Figure 10 This is a schematic diagram of the simulation testing device provided by the present invention.

[0052] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0054] With the development of hardware, especially the hardware of heterogeneous computing systems whose architecture has shifted from a single-function mode to a highly complex one, it is common to find that not only multiple computing clusters are included, but each computing cluster may also contain multiple computing units (such as CPUs, GPUs, etc.). However, due to the different degrees of compatibility between different hardware architectures and code, in code compilation applications, programmers are usually required to manually compile and optimize code for the target hardware, resulting in low efficiency in simulation program development.

[0055] To address the aforementioned shortcomings, this invention provides a simulation program generation method, a simulation testing method, and related equipment, which enables the automated generation of simulation programs corresponding to the target hardware, effectively improving the efficiency of simulation program development.

[0056] Figure 1 This is one of the flowcharts illustrating the simulation program generation method provided by the present invention, such as... Figure 1 As shown, the method includes the following:

[0057] Step 101: Obtain the compilation information associated with the target hardware from the configuration file;

[0058] In one embodiment, the target hardware can be a traditional CPU, GPU, etc., or a multi-core processor, edge computing device, etc., or a wafer-level chip or heterogeneous computing platform.

[0059] In this invention, compilation information refers to a set of metadata used to guide the compiler in converting high-level code into a hardware-specific executable form. Compilation information typically describes the characteristic parameters, constraints, and optimization directions of the target hardware, enabling the compiler to generate efficient code adapted to the target hardware. Taking a CPU as an example, the compilation information associated with the CPU may include supported instruction sets (such as AVX-512, SSE), cache hierarchy (L1 / L2 / L3 cache size), number of cores, and frequency. Taking a wafer-level chip as an example, the compilation information for the wafer-level chip may include kernel function names, the computing clusters contained in the target hardware, and the workload of the computing units contained in the computing cluster.

[0060] Step 102: Using the start and stop operations, encapsulate the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks to obtain MLIR code containing heterogeneous dialects and simulation dialects.

[0061] In one embodiment, the launch and termination operations are used to encapsulate code blocks; in practice, the launch and termination operations can be customized by the user. For example, a launch operation can be added to the beginning of the code that the user needs to encapsulate, and a termination operation can be added when the identification ends.

[0062] In one embodiment, MLIR (Multi-Level Intermediate Representation) code is a multi-level intermediate representation framework that enables the conversion from high-level algorithm descriptions to low-level hardware. For example, by modifying MLIR code according to the computational tasks of the target hardware, the modified MLIR code can run in the target hardware's runtime environment. Specifically, MLIR can not only serve as an intermediate representation for front-end languages ​​such as Python, C++, and TensorFlow, but also generate code for different hardware such as CPUs, GPUs, and wafer-level chips. In this invention, the preset MLIR code can be pre-integrated into the compiler or development platform, or it can be obtained by developers by modifying existing MLIR code according to the target hardware.

[0063] In MLIR, developers are allowed to extend and modify language features, adapting them to specific domain requirements through user-defined dialects (CustomDialect). MLIR achieves modularity through a dialect mechanism, with each dialect defining domain-specific operations (Ops), types, and transformations. In this invention, to automatically generate a simulation program for the target hardware based on compilation information associated with the target hardware, kernel function blocks in the pre-defined MLIR code that match the computational tasks of the target hardware can be encapsulated into heterogeneous and simulation dialect code blocks, so that the final generated simulation program can meet the simulation requirements of the target hardware.

[0064] Heterogeneous dialects abstract the characteristics of different hardware into a unified representation, enabling MLIR code to adapt to different hardware without modifying the core logic. Taking wafer-level chip code generation as an example, the kernel function in the preset MLIR code is determined based on the kernel function name in the compilation information; the kernel function is the core of the target hardware's computation process. Then, the computing cluster, computing units in the computing cluster, and task load of the computing units are defined in the compilation information, providing a basis for the subsequent allocation of computing tasks on different computing units. Simulation dialects are used to further generate interface functions for the simulation platform, enabling MLIR code containing heterogeneous and simulation dialects to be converted into MLIR code oriented towards the simulation platform.

[0065] Therefore, when the target hardware is a heterogeneous architecture such as a wafer-level chip, in order to accurately simulate the computing tasks of the target hardware, the compilation information may include the kernel function name, the computing cluster contained in the target hardware, and the task load of the computing units contained in the computing cluster.

[0066] Correspondingly, such as Figure 2 As shown, step 102 may include the following steps:

[0067] Step 201: Traverse the preset MLIR code and identify the kernel function code block in the preset MLIR code that matches the kernel function name;

[0068] In one embodiment, when mapping MLIR code containing heterogeneous dialects and simulated dialects to MLIR code oriented towards the target hardware, it is necessary to divide the computational tasks of the target hardware according to the compilation information in the configuration file. Before this, the kernel functions (i.e., the core computational functions that perform the computational tasks) must be identified. In this invention, the preset MLIR code can be traversed, and kernel function code blocks in the preset MLIR code that match the kernel function names can be identified through keyword matching.

[0069] Step 202: Using the start and stop operations, encapsulate the kernel function code block in the preset MLIR code into heterogeneous and simulation dialect code blocks;

[0070] In one embodiment, a kernel function code block in a preset MLIR code can be encapsulated through a launch operation and a terminator operation; for example, a launch instruction can be generated at the initial position of the kernel function code block in the preset MLIR code, and a terminator instruction can be generated at the end position of the kernel function code block in the preset MLIR code.

[0071] Step 203: Use the task load of the computing cluster and the computing unit to assign values ​​to the corresponding parameters in the heterogeneous and simulated dialect code blocks to obtain MLIR code containing heterogeneous dialects and simulated dialects.

[0072] The scheme shown in steps 201 to 203 hides the underlying implementation details by identifying kernel function code blocks and encapsulating them into heterogeneous and simulation dialect code blocks. Developers do not need to delve into the internal logic of kernel functions; they only need to focus on the input and output of the code blocks and their overall functionality, thus improving the abstraction level of the code. At the same time, by assigning parameters based on the workload of the computing cluster and computing units, the code blocks can adapt to different resource conditions, enhancing the flexibility of the code. When resource conditions change, there is no need to significantly modify the code structure; only the parameters need to be adjusted, which facilitates code maintenance and expansion.

[0073] Since the subsequent task allocation process will utilize information such as the workload of computing clusters and computing units in the compilation information, it is also necessary to use the workload of computing clusters and computing units to assign values ​​to the corresponding parameters in heterogeneous and simulation dialect code blocks, so as to provide a basis for the allocation of subsequent computing tasks.

[0074] Step 103: Map the MLIR code containing heterogeneous dialects and simulation dialects to MLIR code oriented towards the target hardware;

[0075] In one embodiment, step 102 has already converted the preset MLIR code into MLIR code containing heterogeneous dialects and simulation dialects. At this point, it is only necessary to convert the task proportion logic (task load of computing units) abstracted in the MLIR code containing heterogeneous dialects and simulation dialects into resource allocation logic that the target hardware can understand, so that the MLIR code containing heterogeneous dialects and simulation dialects can be mapped into MLIR code oriented towards the target hardware.

[0076] Based on this, such as Figure 3 As shown, step 103 may include the following steps:

[0077] Step 301: Identify the parallel task operation sequences in heterogeneous and simulated dialect code blocks;

[0078] Here, `hyper.launch` is a high-level abstraction for initiating parallel tasks. The sequence of parallel task operations in heterogeneous and simulation dialect code blocks is typically identified by `hyper.luanch`. Recognizing the `hyper.luanch` operation allows identification of the operation sequence within the heterogeneous and simulation dialect code block. An operation sequence refers to a series of operations executed sequentially in the dialect; these operations collectively constitute the abstract task allocation logic.

[0079] For example, the parallel task operation sequence of the target hardware may include:

[0080] hyper.launch %i = %c0 to %c100 { compute_kernel() hyper.terminator};

[0081] Here, hyper.launch indicates that the parallel task is launched on the target hardware, %i = %c0 to %c100 defines the execution range of the parallel task, which is assumed to be 0-100% here; then the compute_kernel() function is executed in sequence, and finally the hyper.terminator operation is executed, which is the parallel task termination operation.

[0082] Step 302: Based on the values ​​of the corresponding parameters in the heterogeneous and simulation dialect code blocks, convert the parallel task operation sequence into a cyclic execution operation sequence for different computing clusters;

[0083] In one embodiment, the values ​​of the corresponding parameters in the heterogeneous and simulated dialect code blocks are usually represented by dutyRatio. DutyRatio defines the task load of each computing unit in the target hardware according to the assignment in step 203. Based on the computing cluster to which different computing units belong and the task load of different computing units, the task allocation of each computing cluster can be obtained. Then, based on the predefined conversion pass, the parallel task operation sequence can be converted into a cyclic execution operation sequence for different computing clusters.

[0084] For example, if the target hardware contains two computing clusters, the sequence of cyclic execution operations targeting different computing clusters could include:

[0085] hyper.for %i = %c0 to %c60 { compute_kernel()},

[0086] and,

[0087] hyper.for %i = %c60 to %c100 { compute_kernel()};

[0088] In this example, `hyper.for %i = %c0 to %c60 { compute_kernel()}` means starting the `compute_kernel()` function (the target kernel function) on the first compute cluster and executing iteratively within the range of 0-60%. `hyper.for%i = %c60 to %c100 { compute_kernel()}` means starting the `compute_kernel()` function on the second compute cluster and executing iteratively within the range of 60%-100%. It's important to note that 0-60% and 60%-100% are just examples; their actual ranges can be determined based on the values ​​assigned in `dutyRatio`.

[0089] Step 303: Convert the cyclic execution operation sequence for different computing clusters into specific implementations on different computing units of the corresponding computing clusters to obtain MLIR code for the target hardware.

[0090] Taking the target hardware's computing units as including a CPU and a GPU, and each of the first and second computing clusters containing one CPU and one GPU respectively, as an example, the specific implementation of the CPU in the first computing cluster may include:

[0091] scf.for %i = %c0 to %c30 { compute_kernel()};

[0092] In this context, `scf.for %i = %c0 to %c30 { compute_kernel()}` means starting on the CPU of the first computing cluster and executing the `compute_kernel()` function in a loop within the range of 0-30%.

[0093] The specific implementation of GPUs within the first computing cluster may include:

[0094] gpu.launch %c30 to %c60 { compute_kernel() gpu.terminator};

[0095] `gpu.launch %c30 to %c60 { compute_kernel() gpu.terminator}` means launching on the GPUs of the first computing cluster, first executing the `compute_kernel()` function within the range of 30%-60%, and then executing `gpu.terminator`.

[0096] The specific implementation of CPUs within the second computing cluster may include:

[0097] scf.for %i = %c60 to %c80 { compute_kernel()};

[0098] The statement `scf.for %i = %c60 to %c80 { compute_kernel()}` means starting the `compute_kernel()` function on the CPU of the second computing cluster and executing it in a loop within the range of 60%-80%.

[0099] The specific implementation of GPUs within the second computing cluster may include:

[0100] gpu.launch %c80 to %c100 { compute_kernel() gpu.terminator};

[0101] `gpu.launch %c80 to %c100 { compute_kernel() gpu.terminator}` means launching on the GPU of the second computing cluster, first executing the `compute_kernel()` function within the range of 80%-100%, and then executing `gpu.terminator`.

[0102] The scheme shown in steps 301 to 303 can map MLIR code containing heterogeneous and simulation dialects into MLIR code oriented towards the target hardware based on the values ​​of corresponding parameters in the heterogeneous and simulation dialect code blocks. In this process, by identifying and reasonably converting the parallel task operation sequence, tasks can be accurately allocated to different computing clusters and computing units (such as CPUs and GPUs), ensuring the reliability of the target hardware simulation. Furthermore, converting the code containing heterogeneous and simulation dialects into MLIR code oriented towards the target hardware allows the code to better adapt to different target hardware, eliminating the need for extensive code rewriting when running on different hardware platforms, thus improving code portability. At the same time, developers can perform more abstract and convenient programming based on heterogeneous and simulation dialects, with the conversion scheme responsible for mapping it to target hardware code, reducing the difficulty of direct hardware-oriented programming and improving the development efficiency of simulation programs.

[0103] Step 104: Convert the MLIR code for the target hardware into a simulation program for the simulation platform.

[0104] In one embodiment, the MLIR code for the target hardware can be first converted into MLIR code for the simulation platform; then, the MLIR code for the simulation platform can be converted into a simulation program for the simulation platform through the LLVM intermediate representation.

[0105] In summary, this invention provides a simulation program generation method, comprising: obtaining compilation information associated with the target hardware in a configuration file; using start and stop operations to encapsulate code blocks corresponding to the compilation information in a preset MLIR code into heterogeneous and simulation dialect code blocks, thereby obtaining MLIR code containing heterogeneous and simulation dialects; mapping the MLIR code containing heterogeneous and simulation dialects to MLIR code oriented towards the target hardware; and converting the MLIR code oriented towards the target hardware into a simulation program oriented towards the simulation platform.

[0106] The simulation program generation method provided by this invention converts preset MLIR code into a simulation program for the simulation platform through dialect injection and mapping operations based on the compilation information associated with the target hardware, thereby realizing the automated generation of simulation programs corresponding to the target hardware and improving the efficiency of simulation program development.

[0107] Furthermore, the simulation program generation method provided by this invention, through heterogeneous dialects, can automatically calculate the task allocation process in the preset MLIR code, realize efficient task scheduling in different computing clusters and computing units, and simplify the task management process for developers, providing efficient and reliable support for simulation tasks of complex hardware.

[0108] Furthermore, the simulation program generation method provided by this invention utilizes a mapping mechanism to automatically generate MLIR code for target hardware into simulation programs for the simulation platform, reducing the complexity and error rate for developers and further improving development efficiency.

[0109] In one embodiment, taking a wafer-level chip as an example, if the computing units of the target hardware include both a CPU and a GPU, the code for the GPU can be converted into a PTX file designed to express parallel computing. PTX files are better adapted to thread hierarchies and memory models. After converting the MLIR code for the GPU in the target hardware into a PTX file, it can ensure that the simulation platform accurately reflects the GPU behavior during simulation.

[0110] Based on this, in one embodiment, the computing unit includes a CPU and a GPU; step 104 includes:

[0111] Extract the MLIR code for the GPU in the target hardware from the MLIR code for the target hardware, and generate a PTX file based on the MLIR code for the GPU in the target hardware.

[0112] and,

[0113] Extract the MLIR code for the CPU and GPU in the target hardware from the MLIR code for the target hardware, identify the operation sequences in the MLIR code for the CPU and GPU in the target hardware, convert the operation sequences in the MLIR code for the CPU and GPU in the target hardware into interface functions for the simulation platform, and obtain the simulation program for the simulation platform.

[0114] In one embodiment, since the subsequent compilation of the code related to the computational tasks executed on the GPU is completed by the GPU manufacturer's own toolchain (such as the CUDA suite), the actual computational instructions executed in the GPU are generated in the PTX file. Therefore, at the intermediate representation level, only the calls to GPU kernel functions and the addresses of the kernel functions are visible. The same applies to the emulator, so it is necessary to convert the MLIR code for the GPU in the target hardware into a PTX file.

[0115] In one embodiment, more simulation platforms are able to support PTX files as native input and simulate the GPU's execution flow by parsing PTX files. If GPU-oriented MLIR code is used directly, an additional complex conversion layer needs to be developed, leading to reduced simulation program development efficiency.

[0116] Furthermore, such as Figure 4 As shown, the operation sequences in the MLIR code for the CPU and GPU in the target hardware are converted into interface functions for the simulation platform to obtain a simulation program for the simulation platform, including the following steps:

[0117] Step 401: Convert the MLIR code for the target hardware (CPU and GPU) into MLIR code for the simulation platform;

[0118] In one embodiment, simulation dialects and operations can be used to convert the operations in the simulation dialects and operations into interface functions in the simulation platform, thereby realizing the conversion of MLIR code for the target hardware CPU and GPU into MLIR code for the simulation platform.

[0119] Step 402: Perform a first descent transform on the MLIR code for the simulation platform to obtain the LLVM intermediate representation;

[0120] Step 403: Perform a second descent transformation on the LLVM intermediate representation to obtain the simulation program for the simulation platform.

[0121] The first descent transformation and the second descent transformation are conversion processes from high-level code to low-level code. The terms "first" and "second" only indicate that the specific processes of the descent transformation are different and have no actual physical meaning.

[0122] It is important to note that although the descent transformation of the LLVM intermediate representation into a simulation program for the simulation platform utilizes existing functionalities of the LLVM tool, an additional runtime library needs to be linked during the descent transformation process. This runtime library is encapsulated by the call interface functions in the simulator, thereby mapping the functions in the LLVM intermediate representation generated by the simulation dialect and the descent transformation operation to the call interface functions in the simulator.

[0123] By using the LLVM intermediate representation to convert MLIR code for CPUs and GPUs targeting the target hardware into MLIR code for the simulation platform, the advantages of the LLVM intermediate representation can be fully utilized. It can accept multiple front-end languages ​​(such as C, C++, Fortran, Rust, etc.), does not depend on specific compilers (such as GCC, Clang, etc.), and is supported by multiple simulation platforms (such as QEMU, SystemC, etc.). The simulation program is more adaptable and has a wider range of applications.

[0124] Based on this, the present invention also provides a simulation testing method. For example... Figure 5 As shown, the simulation test method includes the following steps:

[0125] Step 501: Based on the structural information of the target hardware, construct a simulation platform;

[0126] In one embodiment, taking a wafer-level chip as an example, the structural information of the target hardware should typically include GPU component information, CPU component information, network connectivity component information, and memory component information.

[0127] In one embodiment, the structural information of the target hardware may also be included in the configuration file.

[0128] Step 502: Run the simulation program on the simulation platform to obtain the simulation test results of the target hardware; the simulation program is generated using the simulation program generation method provided by this invention.

[0129] Taking the target hardware as a wafer-level chip as an example, such as Figure 6 As shown, step 502 may include the following steps:

[0130] Step 601: Pass the simulation program to the CPU component in the simulation platform, and pass the PTX file to the GPU component in the simulation platform;

[0131] Step 602: After the CPU component recognizes the function call to the GPU in the simulation program, it starts the GPU component by calling it. The network connection component connects to the various components on the simulation platform. The memory component manages the memory usage on the wafer-level chip simulation platform in a unified manner, so that the various components on the simulation platform can work together to complete the simulation test of the simulation program.

[0132] Step 603: Output simulation results.

[0133] Furthermore, the simulation program and target hardware design can be optimized based on the output simulation results, thereby shortening the development cycle of both the simulation program and the target hardware.

[0134] The simulation testing method provided by this invention, by constructing a simulation platform based on the structural information of the target hardware, enables the testing and optimization of simulation programs to be independent of the hardware entity, thereby significantly accelerating the code development cycle and enabling early detection and optimization of defects in the code development process, thus reducing the development cost of simulation programs and improving the development efficiency of simulation programs.

[0135] Furthermore, since the structural information of the target hardware can also be included in the configuration file, the unified management of the generation of simulation programs and simulation platforms through configuration files can greatly simplify the development process and avoid repetitive operations; at the same time, it can also ensure a high degree of consistency between the simulation program and the simulation platform, improving the convenience and adaptability of simulation program development.

[0136] To facilitate a better understanding of the simulation program generation method and simulation testing method provided by this invention, the technical solution of this invention will be further explained below using a wafer-level chip as the target hardware.

[0137] With the development of wafer-level chip packaging technology, wafer-level chips not only contain multiple computing clusters, but each computing cluster may also integrate multiple computing units, such as CPUs and GPUs. Moreover, when developing wafer-level chips, there is usually no physical wafer-level chip hardware yet.

[0138] To improve the development efficiency of wafer-level chip simulation programs, this invention provides a simulation program generation method, such as... Figure 7 As shown, it includes the following steps:

[0139] Step 701: Obtain the compilation information associated with the wafer-level chip in the configuration file;

[0140] The compilation information includes at least the kernel function name, the computing cluster contained in the target hardware, and the task load of the computing units contained in the computing cluster;

[0141] Step 702: Based on the kernel function names in the compilation information, determine the target kernel function in the preset MLIR code that matches the kernel function name;

[0142] Step 703: Based on the values ​​of the corresponding parameters in the heterogeneous and simulation dialect code blocks, convert the parallel task operation sequence into a cyclic execution operation sequence for different computing clusters;

[0143] Step 704: Convert the cyclic execution operation sequence for different computing clusters into specific implementations on different computing units of the corresponding computing clusters to obtain MLIR code for the target hardware;

[0144] Step 705: Extract the MLIR code for the GPU in the target hardware from the MLIR code for the target hardware, and generate a PTX file based on the MLIR code for the GPU in the target hardware.

[0145] Step 706: Convert the MLIR code for the target hardware (CPU and GPU) into MLIR code for the simulation platform;

[0146] Step 707: Perform a descent transformation on the MLIR code for the simulation platform to obtain the simulation program for the simulation platform.

[0147] Furthermore, the present invention provides a simulation testing method, such as... Figure 8 As shown, it includes the following steps:

[0148] Step 801: Pass the simulation program to the CPU component in the simulation platform, and pass the PTX file to the GPU component in the simulation platform;

[0149] Step 802: After the CPU component recognizes the function call to the GPU in the simulation program, it starts the GPU component by calling it. The network connection component connects to the various components on the simulation platform. The memory component manages the memory usage on the wafer-level chip simulation platform in a unified manner, so that the various components on the simulation platform can work together to complete the simulation test of the simulation program.

[0150] Step 803: Output simulation results.

[0151] The simulation program generation apparatus provided by the present invention is described below. The simulation program generation apparatus described below and the simulation program generation method described above can be referred to in correspondence.

[0152] like Figure 9 As shown, the simulation program generation device 900 provided by the present invention includes:

[0153] Module 901 is used to obtain compilation information associated with the target hardware from the configuration file;

[0154] The generation module 902 is used to encapsulate the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks by using the start operation and the termination operation, so as to obtain MLIR code containing heterogeneous dialect and simulation dialect.

[0155] Mapping module 903 is used to map MLIR code containing heterogeneous dialects and simulation dialects to MLIR code oriented towards the target hardware;

[0156] The conversion module 904 converts MLIR code for the target hardware into a simulation program for the simulation platform.

[0157] In one embodiment, the compilation information includes at least the kernel function name, the computing cluster contained in the target hardware, and the task load of the computing units contained in the computing cluster;

[0158] Accordingly, the generation module 902 is specifically used for:

[0159] Traverse the preset MLIR code and identify the kernel function code block in the preset MLIR code that matches the kernel function name;

[0160] By using start and stop operations, the kernel function code blocks in the preset MLIR code are encapsulated into heterogeneous and simulation dialect code blocks;

[0161] The corresponding parameters in the heterogeneous and simulated dialect code blocks are assigned values ​​based on the task load of the computing cluster and computing units, resulting in MLIR code containing both heterogeneous and simulated dialects. In one embodiment, the mapping module 903 is specifically used for:

[0162] Identify parallel task operation sequences in heterogeneous and simulated dialect code blocks;

[0163] Based on the values ​​of the corresponding parameters in the heterogeneous and simulation dialect code blocks, the parallel task operation sequence is converted into a cyclic execution operation sequence for different computing clusters;

[0164] The cyclic execution sequence of operations for different computing clusters is converted into specific implementations on different computing units of the corresponding computing clusters to obtain MLIR code for the target hardware.

[0165] In one embodiment, the computing unit includes a CPU and a GPU;

[0166] Accordingly, the conversion module 904 is specifically used for:

[0167] Extract the MLIR code for the GPU in the target hardware from the MLIR code for the target hardware, and generate a PTX file based on the MLIR code for the GPU in the target hardware.

[0168] and,

[0169] Extract the MLIR code for the CPU and GPU in the target hardware from the MLIR code for the target hardware, identify the operation sequences in the MLIR code for the CPU and GPU in the target hardware, convert the operation sequences in the MLIR code for the CPU and GPU in the target hardware into interface functions for the simulation platform, and obtain the simulation program for the simulation platform.

[0170] In one embodiment, the conversion module 904 is further specifically used for:

[0171] By utilizing simulation dialects and operations, MLIR code for the target hardware (CPU and GPU) is converted into MLIR code for the simulation platform.

[0172] The MLIR code for the simulation platform is subjected to a first descent transform to obtain the LLVM intermediate representation;

[0173] Perform a second descent transformation on the LLVM intermediate representation to obtain the simulation program for the simulation platform.

[0174] The simulation testing device provided by the present invention is described below. The simulation testing device described below and the simulation testing method described above can be referred to in correspondence.

[0175] like Figure 10 As shown, the simulation testing device 1000 provided by the present invention includes:

[0176] Module 1001 is used to build a simulation platform based on the structural information of the target hardware.

[0177] The simulation test module 1002 is used to run a simulation program on a simulation platform to obtain the simulation test results of the target hardware; the simulation program is generated using the simulation program generation method provided by this invention.

[0178] Figure 11 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 11 As shown, the electronic device may include: a processor 1101, a communications interface 1102, a memory 1103, and a communication bus 1104, wherein the processor 1101, the communications interface 1102, and the memory 1103 communicate with each other through the communication bus 1104. The processor 1101 can call logical instructions in the memory 1103 to execute the simulation program generation method or simulation testing method provided by the present invention.

[0179] Furthermore, the logical instructions in the aforementioned memory 1103 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0180] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the simulation program generation method or simulation testing method provided by the present invention.

[0181] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the simulation program generation method or simulation testing method provided by the present invention.

[0182] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0183] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating simulation programs, characterized in that, include: Retrieve compilation information associated with the target hardware from the configuration file; By using start and stop operations, the code blocks corresponding to the compilation information in the preset MLIR code are encapsulated into heterogeneous and simulation dialect code blocks, resulting in MLIR code containing heterogeneous and simulation dialects; The MLIR code containing heterogeneous dialects and simulated dialects is mapped to MLIR code oriented towards the target hardware; The MLIR code for the target hardware is converted into a simulation program for the simulation platform. The compilation information includes kernel function names, the computing cluster contained in the target hardware, and the task load of the computing units contained in the computing cluster; The step of encapsulating the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks to obtain MLIR code containing heterogeneous and simulation dialects includes: Traverse the preset MLIR code and identify the kernel function code block in the preset MLIR code that matches the kernel function name; By using start and stop operations, the kernel function code block in the preset MLIR code is encapsulated into heterogeneous and simulation dialect code blocks; The corresponding parameters in the heterogeneous and simulated dialect code blocks are assigned values ​​using the task load of the computing cluster and the computing unit to obtain MLIR code containing heterogeneous dialects and simulated dialects; The step of mapping the MLIR code containing heterogeneous dialects and emulated dialects to MLIR code oriented towards the target hardware includes: Identify parallel task operation sequences in heterogeneous and simulated dialect code blocks; Based on the values ​​of the corresponding parameters in the heterogeneous and simulation dialect code blocks, the parallel task operation sequence is converted into a cyclic execution operation sequence for different computing clusters; The cyclic execution operation sequence for different computing clusters is converted into a specific implementation on different computing units of the corresponding computing cluster, thus obtaining MLIR code for the target hardware.

2. The method according to claim 1, characterized in that, The computing unit includes a CPU and a GPU; converting MLIR code for the target hardware into a simulation program for the simulation platform includes: Extract the MLIR code for the GPU in the target hardware from the MLIR code for the target hardware, and generate a PTX file based on the MLIR code for the GPU in the target hardware; and, Extract the MLIR code for the CPU and GPU in the target hardware from the MLIR code for the target hardware, identify the operation sequences in the MLIR code for the CPU and GPU in the target hardware, convert the operation sequences in the MLIR code for the CPU and GPU in the target hardware into interface functions for the simulation platform, and obtain the simulation program for the simulation platform.

3. The method according to claim 2, characterized in that, The step of converting the operation sequences in the MLIR code for the CPU and GPU in the target hardware into interface functions for the simulation platform, thereby obtaining a simulation program for the simulation platform, includes: Using simulation dialects and operations, the MLIR code for the CPU and GPU of the target hardware is converted into MLIR code for the simulation platform; The MLIR code for the simulation platform is subjected to a first descent transform to obtain the LLVM intermediate representation; The LLVM intermediate representation is subjected to a second descent transformation to obtain a simulation program for the simulation platform.

4. A simulation testing method, characterized in that, include: A simulation platform is constructed based on the structural information of the target hardware. The simulation program is run on the simulation platform to obtain the simulation test results of the target hardware; The simulation program is generated using the simulation program generation method according to any one of claims 1 to 3.

5. A simulation program generation device, characterized in that, include: The acquisition module is used to obtain the compilation information associated with the target hardware from the configuration file; The generation module is used to encapsulate the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks by using the start and stop operations, so as to obtain MLIR code containing heterogeneous dialect and simulation dialect. A mapping module is used to map the MLIR code containing heterogeneous dialects and simulated dialects into MLIR code oriented towards the target hardware; A conversion module is used to convert MLIR code for the target hardware into a simulation program for the simulation platform; The compilation information includes kernel function names, the computing cluster contained in the target hardware, and the task load of the computing units contained in the computing cluster; The step of encapsulating the code blocks corresponding to the compilation information in the preset MLIR code into heterogeneous and simulation dialect code blocks to obtain MLIR code containing heterogeneous and simulation dialects includes: Traverse the preset MLIR code and identify the kernel function code block in the preset MLIR code that matches the kernel function name; By using start and stop operations, the kernel function code block in the preset MLIR code is encapsulated into heterogeneous and simulation dialect code blocks; The corresponding parameters in the heterogeneous and simulated dialect code blocks are assigned values ​​using the task load of the computing cluster and the computing unit to obtain MLIR code containing heterogeneous dialects and simulated dialects; The step of mapping the MLIR code containing heterogeneous dialects and emulated dialects to MLIR code oriented towards the target hardware includes: Identify parallel task operation sequences in heterogeneous and simulated dialect code blocks; Based on the values ​​of the corresponding parameters in the heterogeneous and simulation dialect code blocks, the parallel task operation sequence is converted into a cyclic execution operation sequence for different computing clusters; The cyclic execution operation sequence for different computing clusters is converted into a specific implementation on different computing units of the corresponding computing cluster, thus obtaining MLIR code for the target hardware.

6. A simulation testing device, characterized in that, include: The building block is used to construct a simulation platform based on the structural information of the target hardware. The simulation test module is used to run a simulation program on the simulation platform to obtain the simulation test results of the target hardware; the simulation program is generated using the simulation program generation method according to any one of claims 1 to 3.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the simulation program generation method as described in any one of claims 1 to 3, or the simulation testing method as described in claim 4.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the simulation program generation method as described in any one of claims 1 to 3, or the simulation testing method as described in claim 4.

Citation Information

Patent Citations

  • Task flow-based classical-quantum cooperative computing programming method and model

    CN116257222A

  • Probability model AI field compiling system and method based on multistage intermediate representation

    CN116679933A