Performance-transplantable programming framework system supporting various heterogeneous supercomputing architectures
By extending the Kokkos programming model and supporting a variety of heterogeneous supercomputing architectures, the existing programming model has been solved, and the existing programming model is highly complex and the development environment is incomplete, and the efficient and easy-to-use performance portable programming framework system is realized, which improves the development efficiency and application potential of high-performance computing.
Patent Information
- Application Number
- CN202510099459.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
The existing high-performance computing programming model has problems such as low abstraction level and close to the underlying hardware, which leads to complex programming logic, difficult programming and error-prone. Especially on Shenwei series processors and Haiguang DCUs, the software development environment is imperfect, forming a "programming wall", which hinders the efficient development and deployment of application software.
By extending the open source performance portable programming model Kokkos, it increases support for a variety of heterogeneous supercomputing architectures (including the Shenwei series computing platform and the Haiguang DCU series computing platform), and provides core library modules, programming interface modules, performance optimization modules, cross-platform support modules, document and example modules, and community and maintenance modules to form a performance portable programming framework system.
It improves the ease of use of computing platforms, reduces development costs, improves development efficiency, realizes the seamless operation of applications on multiple heterogeneous supercomputing architectures, and promotes the rapid development of high-performance computing technology.
Smart Images

Figure CN120029668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high performance computing technology, and more specifically to a performance portable programming framework system that supports multiple heterogeneous supercomputing architectures. Background Art
[0002] The architecture of high-performance computers is undergoing unprecedented changes to adapt to increasingly complex and diverse computing needs. In order to fully utilize the advantages of different hardware devices, the industry has introduced a variety of high-performance computing programming models, such as MPI (Message Passing Interface), CUDA (Unified Compute Device Architecture), HIP (Heterogeneous Computing Platform), OpenMP (Open Multiprocessing) and SYCL (Heterogeneous System-Level Computing). However, these programming models generally have the problem of low abstraction level and interface design close to the underlying hardware, which leads to complex programming logic, high programming difficulty and easy errors.
[0003] As a representative of E-class computing, the new generation of heterogeneous Sunway supercomputers has powerful computing capabilities that provide strong support for scientific research and engineering applications. The Sunway series processors use independent instruction sets and have completely independent intellectual property rights, and are a model of China's scientific and technological self-reliance. However, the software development environment of the Sunway processor is still imperfect, especially its slave core has powerful computing power but cannot fully support C++, and the learning threshold for calling the Athread programming model of the slave core is high, forming the so-called "programming wall", which seriously hinders the efficient development and deployment of application software.
[0004] In addition, Haiguang DCU, as a GPU-like accelerator card, also faces similar programming challenges. Although the HIP programming model it adopts has improved programming flexibility to a certain extent, it still has problems such as a steep learning curve and high programming complexity, which limits the widespread application of application software on Haiguang DCU.
[0005] In response to the above problems, the industry urgently needs a programming model that can achieve performance portability across platforms to reduce development costs, improve development efficiency, and meet the application requirements of the new generation of high-performance computing devices. As a library-based open source programming model, the Kokkos C++ library provides a performance-portable solution for programs in the scientific and engineering fields. However, the current Kokkos library has a limited support scope and does not yet support high-performance computing devices such as the Sunway heavy-core heterogeneous processor and the Hygon DCU, which limits its application potential in the field of high-performance computing.
[0006] Currently, the programming models of different supercomputer architectures vary greatly, and the programming logic is complex. The development cost of implementing multiple sets of customized application software for different high-performance computing devices is high, the cycle is long, and the readability is poor. In addition, application software is also constantly iterating and updating, and developers need to maintain multiple sets of programs at the same time, which has high maintenance costs. Therefore, the traditional solution of constantly recoding and porting application software using different programming models to different high-performance computers is not only costly but also unsustainable.
[0007] Therefore, how to provide a performance-portable programming framework system that supports multiple heterogeneous supercomputing architectures is an urgent problem that technical personnel in this field need to solve. Summary of the invention
[0008] In view of this, the present invention provides a performance portable programming framework system that supports multiple heterogeneous supercomputing architectures. By expanding the open source performance portable programming model Kokkos, support for multiple heterogeneous supercomputing architectures (including Shenwei series computing platforms and Haiguang DCU series computing platforms) is increased to improve the usability of the computing platform, reduce development costs, improve development efficiency, and promote the rapid development of high-performance computing technology.
[0009] In order to achieve the above object, the present invention adopts the following technical solution:
[0010] A performance-portable programming framework system that supports multiple heterogeneous supercomputing architectures, including:
[0011] Core library module, which is used to provide abstraction and data management functions for parallel execution of code, and supports backend programming models for multiple heterogeneous supercomputing architectures;
[0012] The programming interface module is used to define a unified programming interface, allowing developers to write high-performance computing applications without considering the differences in the underlying hardware;
[0013] Performance optimization module, which is used to improve the execution efficiency and performance of applications on the target heterogeneous supercomputing architecture through algorithm optimization and scheduling strategies by leveraging fine-grained data parallelism and memory access pattern abstraction;
[0014] Cross-platform support module, which is used to achieve code portability on different hardware platforms and ensure that applications can run seamlessly on a variety of heterogeneous supercomputing architectures;
[0015] Documentation and Examples module, which provides detailed programming guides, API reference documentation, and sample codes;
[0016] The community and maintenance module is used to build an open developer community and provide technical support, problem solving, and code contribution platforms.
[0017] Preferably, the cross-platform support module includes:
[0018] The template function registration and instantiation unit is used to register preset functions with different template parameters and instantiate the preset functions to adapt to the limitations of the target heterogeneous supercomputing architecture.
[0019] Preferably, the template function registration and instantiation unit includes a registration subunit, and uses a linked list data structure to store registered preset functions, and each node in the linked list data structure contains a pointer to the preset function and an execution mode configuration.
[0020] Preferably, the performance optimization module includes:
[0021] A parallel task allocation and mapping unit, used for realizing load balancing by adopting a task block strategy and a task block allocation strategy;
[0022] The callback mechanism and execution unit are used to find and call the instantiated preset function matching the request through the callback mechanism when a kernel call is initiated;
[0023] The acceleration unit is used to accelerate the matching process of the preset function during the kernel startup process.
[0024] Preferably, the acceleration unit utilizes the local data memory and single instruction multiple data vectorization technology of the target heterogeneous supercomputing architecture to improve matching efficiency.
[0025] Preferably, the callback mechanism utilizes a linked list data structure to perform efficient search operations in response to kernel startup requests.
[0026] Preferably, the template function registration and instantiation unit further includes:
[0027] A template parameter parsing and conversion subunit, used to parse the template parameters of the preset function and convert the template parameters into a format acceptable to the target heterogeneous supercomputing architecture;
[0028] The instantiation code generation subunit generates corresponding instantiation code according to the parsed template parameters.
[0029] Preferably, the parallel task allocation and mapping unit further includes:
[0030] The load balancing evaluation subunit is used to divide the task data into multiple tiles and calculate the total number of tiles according to the range of each dimension loop and the length of the tile:
[0031]
[0032] Among them, len_range n and len_tile n Respectively represent the loop body range and tile length in the n-th dimension loop;
[0033] Calculate the number of tiles that each processing unit should process from the number of cores and the total number of tiles:
[0034]
[0035] Where num_cpe is the number of slave cores in the coarse-grained process.
[0036] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a performance-portable programming framework system that supports multiple heterogeneous supercomputing architectures, and the framework system includes a core library module, a programming interface module, a performance optimization module, a cross-platform support module, a document and example module, and a community and maintenance module. The core library module provides abstraction and data management functions for parallel execution of code, and supports back-end programming models of multiple heterogeneous supercomputing architectures. The programming interface module defines a unified programming interface, allowing developers to write high-performance computing applications without considering the differences in underlying hardware. The performance optimization module improves the execution efficiency and performance of the application on the target heterogeneous supercomputing architecture through algorithm optimization and scheduling strategies. The cross-platform support module is one of the keys to the framework system, which includes a template function registration and instantiation unit for registering preset functions with different template parameters, and instantiating these functions to adapt to the limitations of the target heterogeneous supercomputing architecture. The unit also uses a linked list data structure to store registered preset functions to achieve efficient search and call. In addition, the performance optimization module also includes a parallel task allocation and mapping unit, a callback mechanism and execution unit, and an acceleration unit. These units work together to achieve load balancing, efficient matching of preset functions, and acceleration of the kernel startup process, thereby further improving the overall performance of the system. The documentation and example module provides detailed programming guidelines and sample codes to help developers get started quickly. The community and maintenance module establishes an open developer community, provides technical support and a code contribution platform, and promotes the exchange and sharing of technology. The performance portable programming framework system of the present invention has the characteristics of high efficiency, flexibility, and portability, provides developers with powerful tools and support, and helps promote the development and application of high-performance computing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0038] Figure 1 A schematic diagram of the structure provided by the present invention;
[0039] Figure 2 It is the overall framework diagram of the programming framework of the present invention;
[0040] Figure 3 The process of initiating a kernel function on a DCU of the present invention;
[0041] Figure 4 The present invention initiates the process of the kernel function on the Shenwei processor. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] The embodiment of the present invention discloses a performance-portable programming framework system that supports multiple heterogeneous supercomputing architectures. Kokkos uses complex template specialization to start functions. However, the AthreadAPI used to start the kernel on the Shenwei slave core only supports the old C language syntax and cannot pass template parameters to the kernel running from the slave core. Therefore, these template functions that cannot be implemented must be instantiated. Based on this, an embodiment of the present invention proposes a method for registering preset functions to reinterpret template parameters and record corresponding registered functions. When the programming framework initiates the kernel, the slave core calls these registered functions through a callback mechanism. It is only necessary to configure the execution mode (loop and specification) for the preset function in different dimensions. Time complexity, memory usage and code stability are easily affected by extreme scale numerical models. The linked list data structure is selected to implement the registration and search of preset functions. This decision takes into account the time, space complexity and robustness of the code. A linked list is a linear data structure with a time and space complexity of O(n), where n is related to the number of kernel functions. At the same time, the present invention also uses the local data memory (LDM) and single instruction multiple data (SIMD) vectorization of the Shenwei architecture to accelerate the matching process of the preset function during the kernel startup process. Figure 1-2 As shown, including:
[0044] The core library module is used to provide abstraction and data management functions for parallel execution code, and supports back-end programming models for a variety of heterogeneous supercomputing architectures. The core library module is the cornerstone of the programming framework system. It provides an abstract layer for parallel execution code, so that developers do not need to directly face the complex underlying parallel mechanism. This module is not only responsible for data management functions to ensure the consistency and efficiency of data in the parallel computing process, but also widely supports back-end programming models for a variety of heterogeneous supercomputing architectures, such as CUDA, OpenCL and MPI, etc., providing a solid foundation for applications to run on different hardware platforms. The core library module includes a computing power abstract parser, which is used to abstract and parse the development code.
[0045] The programming interface module is used to define a unified programming interface, allowing developers to write high-performance computing applications without considering the differences in the underlying hardware;
[0046] Performance optimization module, which is used to improve the execution efficiency and performance of applications on the target heterogeneous supercomputing architecture through algorithm optimization and scheduling strategies by leveraging fine-grained data parallelism and memory access pattern abstraction;
[0047] The cross-platform support module is used to achieve the portability of code on different hardware platforms, ensuring that applications can run seamlessly on a variety of heterogeneous supercomputing architectures; the cross-platform support module ensures the portability of code on different hardware platforms. It adapts to different heterogeneous supercomputing architectures, allowing applications to run seamlessly on these platforms. This module not only simplifies the process of code migration, but also improves the flexibility and scalability of the system.
[0048] The Documentation and Examples module provides detailed programming guides, API reference documentation, and sample codes, providing developers with comprehensive learning resources and references. These documents and examples not only help developers get started quickly, but also help them deeply understand the principles and implementation details of the system.
[0049] The community and maintenance module is used to build an open developer community, provide technical support, problem solving and code contribution platform, encourage developers to participate in the maintenance and expansion of the framework, and promote the continuous updating and optimization of technology.
[0050] Furthermore, the cross-platform support module includes:
[0051] The template function registration and instantiation unit is used to register preset functions with different template parameters and instantiate the preset functions to adapt to the limitations of the target heterogeneous supercomputing architecture.
[0052] Specifically, the template function registration and instantiation unit includes a registration subunit, which uses a linked list data structure to store registered preset functions, and each node in the linked list contains a pointer to the preset function and an execution mode configuration.
[0053] In another embodiment, the performance optimization module includes:
[0054] A parallel task allocation and mapping unit, used for realizing load balancing by adopting a task block strategy and a task block allocation strategy;
[0055] The callback mechanism and execution unit are used to find and call the instantiated preset function matching the request through the callback mechanism when a kernel call is initiated; it also supports the configuration of execution modes such as loop and specification, where the specification step uses the specification function provided by the Athread interface;
[0056] The acceleration unit is used to accelerate the matching process of the preset function during the kernel startup process.
[0057] Specifically, the acceleration unit utilizes the local data memory and single instruction multiple data vectorization technology of the target heterogeneous supercomputing architecture to improve the matching efficiency.
[0058] Specifically, the callback mechanism utilizes a linked list data structure to perform efficient search operations to respond to kernel startup requests.
[0059] Furthermore, the template function registration and instantiation unit also includes:
[0060] A template parameter parsing and conversion subunit, used to parse the template parameters of the preset function and convert the template parameters into a format acceptable to the target heterogeneous supercomputing architecture;
[0061] The instantiation code generation subunit generates corresponding instantiation code according to the parsed template parameters.
[0062] Furthermore, the Athread programming model requires developers to explicitly allocate parallel tasks / data to each CPE. Therefore, the parallel task allocation and mapping unit also includes:
[0063] The load balancing evaluation subunit is used to divide the task data into multiple tiles and calculate the total number of tiles based on the range of each dimension loop and the length of the tile:
[0064]
[0065] Among them, len_range n and len_tile n Respectively represent the loop body range and tile length in the n-th dimension loop;
[0066] Calculate the number of tiles that each processing unit should process from the number of cores and the total number of tiles:
[0067]
[0068] Where num_cpe is the number of slave cores in the coarse-grained process, which is usually 64. The parallel execution strategy of the "reduce" operator is similar to that of the "for" operator, except that the final reduction step utilizes the reduction function provided by the Athread interface. This function is used to reduce the data of the slave cores within each core group.
[0069] In another embodiment, the present invention further includes a memory model optimization module, which utilizes the characteristic of the master core and the slave core of the Shenwei architecture sharing memory space, reuses the Kokkos memory model in the host space, and provides two methods for optimizing memory latency using the local data memory (LDM) of the Shenwei architecture: method one is to define and use a local array in a function, and method two is to access a data pointer through the View.data interface of Kokkos, and then use the direct memory access (DMA) function provided by Athread.
[0070] Furthermore, the present invention also provides a performance-portable programming model application process:
[0071] Development phase: Developers develop application software based on the development rules of the performance portable programming model Kokkos.
[0072] Compilation phase: Developers port the application software to different supercomputing platforms and modify the compiler, compilation options, etc. based on the corresponding platforms and Kokkos rules.
[0073] Running phase: When the application software is running, the high-performance computing framework Kokkos is initialized first, including setting the hardware platform information and the functions required by the system by reading command line parameters / environment variables / program variables. At the same time, the variables that need to be run in the Kokkos framework are initialized. For heterogeneous computing platforms, the required variables will also be defined on the device side.
[0074] Running phase: When the numerical application software enters the iterative simulation, the pseudo function (functor) provided by C++ is used to recode the kernel function loop into a function object (function object). The operator in the function object defines the index of the work item, which is used for thread mapping and memory address index calculation when the kernel function is running. The function object is passed in as a parameter to an interface called "parallel_for" to implement the call of the kernel function and finally map it to different supercomputing architecture hardware. Figure 3-4They are respectively a process of initiating a kernel function on a domestic DCU and a process of initiating a kernel function on a domestic Shenwei processor of the present invention.
[0075] Running phase: The application software ends and Kokkos ends to release resources.
[0076] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0077] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A performance-portable programming framework system that supports multiple heterogeneous supercomputing architectures, characterized in that: include: Core library module, which is used to provide abstraction and data management functions for parallel execution of code, and supports backend programming models for multiple heterogeneous supercomputing architectures; The programming interface module is used to define a unified programming interface, allowing developers to write high-performance computing applications without considering the differences in the underlying hardware; Performance optimization module, which is used to improve the execution efficiency and performance of applications on the target heterogeneous supercomputing architecture through algorithm optimization and scheduling strategies by leveraging fine-grained data parallelism and memory access pattern abstraction; Cross-platform support module, which is used to achieve code portability on different hardware platforms and ensure that applications can run seamlessly on a variety of heterogeneous supercomputing architectures; Documentation and Examples module, which provides detailed programming guides, API reference documentation, and sample codes; The community and maintenance module is used to build an open developer community and provide technical support, problem solving, and code contribution platforms.
2. According to claim 1, a performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures is characterized in that: The cross-platform support module includes: The template function registration and instantiation unit is used to register preset functions with different template parameters and instantiate the preset functions to adapt to the limitations of the target heterogeneous supercomputing architecture.
3. A performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures according to claim 2, characterized in that: The template function registration and instantiation unit includes a registration subunit, and uses a linked list data structure to store registered preset functions. Each node in the linked list data structure contains a pointer to the preset function and an execution mode configuration.
4. According to claim 2, a performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures is characterized in that: Performance optimization modules include: A parallel task allocation and mapping unit, used for realizing load balancing by adopting a task block strategy and a task block allocation strategy; The callback mechanism and execution unit are used to find and call the instantiated preset function matching the request through the callback mechanism when a kernel call is initiated; The acceleration unit is used to accelerate the matching process of the preset function during the kernel startup process.
5. A performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures according to claim 4, characterized in that: The acceleration unit utilizes the local data storage and single instruction multiple data vectorization technology of the target heterogeneous supercomputing architecture to improve the matching efficiency.
6. A performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures according to claim 4, characterized in that: The callback mechanism utilizes a linked list data structure to perform efficient search operations and respond to kernel startup requests.
7. The performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures according to claim 3, characterized in that: The template function registration and instantiation unit also includes: A template parameter parsing and conversion subunit, used to parse the template parameters of the preset function and convert the template parameters into a format acceptable to the target heterogeneous supercomputing architecture; The instantiation code generation subunit generates corresponding instantiation code according to the parsed template parameters.
8. The performance-portable programming framework system supporting multiple heterogeneous supercomputing architectures according to claim 4, characterized in that: The parallel task allocation and mapping unit also includes: The load balancing evaluation subunit is used to divide the task data into multiple tiles and calculate the total number of tiles according to the range of each dimension loop and the length of the tile: Among them, len_range n and len_tile n Respectively represent the loop body range and tile length in the n-th dimension loop; Calculate the number of tiles that each processing unit should process from the number of cores and the total number of tiles: Where num_cpe is the number of slave cores in the coarse-grained process.