Method for Porting Pytorch Based on Domestic Shenwei Processor
Through the optimization of the architecture adaptation layer and high-performance computing library, the efficient operation of the pytorch framework on the Shenwei processor is achieved, which solves the compatibility problem, improves the computing performance and efficiency, and adapts to the hardware characteristics of the Shenwei processor.
Patent Information
- Application Number
- CN202510307476.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing pytorch framework has compatibility issues on the Shenwei processor platform, and cannot fully utilize its computing performance, resulting in inefficiency in deep learning tasks.
Through the architecture adaptation layer, high-performance computing libraries and compilation adaptation are built and optimized, the pytorch framework and Shenwei processors are realized seamlessly, and the parallel computing power and memory access characteristics of Shenwei processors are used to optimize computing task scheduling and resource management.
It improves the operating performance and computing efficiency of the pytorch framework on Shenwei processors, ensures efficient execution of deep learning tasks, lowers technical thresholds, and maintains compatibility with the original ecosystem.
Smart Images

Figure CN119806638B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method for porting PyTorch based on domestic Shenwei processors. Background Art
[0002] With the rapid development of artificial intelligence technology, deep learning frameworks have become one of the core tools for researching and developing artificial intelligence applications. In particular, as an open-source deep learning framework, PyTorch has become the first choice for researchers and developers worldwide due to its flexibility, ease of use, rich APIs, and model libraries. PyTorch can efficiently support tasks such as the construction, training, debugging, and inference of deep learning models, greatly improving the research and application efficiency of artificial intelligence technology.
[0003] However, with the continuous increase in the complexity of deep learning models, especially in the training and inference tasks of large-scale datasets, traditional computing resources (such as CPUs and GPUs) often face performance bottlenecks. Especially in high-computation-load deep learning tasks, the limitations of the performance of CPUs and GPUs in data processing and computing power have gradually become the key factors restricting the development of deep learning.
[0004] Existing PyTorch frameworks, especially in terms of hardware acceleration, are mainly optimized for x86 architecture CPUs and NVIDIA GPUs. These hardware platforms have specific instruction sets and programming models, and PyTorch has implemented efficient computing libraries and optimization algorithms on these hardware platforms. These computing libraries such as cuDNN, NCCL, etc., utilize the parallel computing power of NVIDIA GPUs to significantly improve the computing efficiency of deep learning tasks.
[0005] However, for the Shenwei processor platform, the PyTorch framework has certain limitations in the existing technology. The Shenwei processor has a unique architecture and instruction set, and its programming environment and toolchain are completely different from those of the x86 architecture and NVIDIA GPUs. Many high-performance computing libraries and modules in the existing PyTorch framework are developed for specific hardware platforms (such as NVIDIA GPUs), which may not be directly compatible on the Shenwei processor, thus affecting the running performance of the PyTorch framework on this platform. Summary of the Invention
[0006] The embodiments of this application provide a method for porting PyTorch based on domestic Shenwei processors to achieve the efficient operation of the PyTorch framework on Shenwei processors and improve the computing efficiency and performance of deep learning tasks.
[0007] To achieve the above object, an embodiment of the present application provides a pytorch porting method based on a domestic Shenwei processor, including:
[0008] Step of building an architecture adaptation layer. Configure an architecture adaptation layer swMath between the pytorch framework and the Shenwei processor. The architecture adaptation layer swMath is provided with multiple adaptation interfaces. Define a corresponding Shenwei implementation for each pytorch operator in the pytorch framework, implement a one-to-one mapping between the multiple adaptation interfaces and the pytorch operators in the pytorch framework, and when the adaptation interface is called, the Shenwei processor can be called for operator operations, so as to bridge the pytorch framework to the Shenwei processor through the architecture adaptation layer swMath to bridge the architectural differences between the two.
[0009] Based on this step, when performing corresponding operator operations, pytorch will call the adaptation interface of the corresponding operator and perform calculations on the Shenwei processor, directly mapping the standard operator operations of pytorch to the optimized implementation for the Shenwei processor.
[0010] Step of porting the high-performance computing library. Identify the high-performance computing libraries that depend on the CUDA platform in the pytorch framework, implement the high-performance computing libraries and optimize them based on the Shenwei processor to obtain optimized high-performance computing libraries. The optimized high-performance computing libraries include a basic mathematical operation library swBLAS, a deep learning operator library swDNN, and a tensor operation library swTensor. Integrate the basic mathematical operation library swBLAS, the deep learning operator library swDNN, and the tensor operation library swTensor into a unified dynamic link library libswmath.so and modify the connection configuration of pytorch to link the dynamic link library libswmath.so with the libtorch.so library of the pytorch framework to achieve correct linking of the system libraries and runtime libraries specific to the Shenwei platform;
[0011] Step of compilation adaptation. Use cross-compilers sw-gcc and sw-g++ for the Shenwei processor to compile the C / C++ library file code on the X86 platform, generate executable files for the Shenwei processor and run them on the Shenwei processor. Specifically, use the CMake tool to modify CMakeLists.txt to add cross-compilation configuration, realize cross-compiling C / C++ library files on the X86 platform, and complete the construction and registration of Python extension modules on the Shenwei platform, and control the file transfer and compilation process between the X86 platform and the Shenwei platform through an automated script.
[0012] In some of these embodiments, in the step of building the architecture adaptation layer, an environment variable is configured for each of the PyTorch operators. By configuring the environment variable, the PyTorch operator is enabled to call the default CPU for computing or call the ShenWei processor for computing, so as to utilize the parallel computing power of the ShenWei processor. Among them, when the value of the environment variable is 0, the default CPU implementation of the PyTorch architecture is called; when the value of the environment variable is 1, the operation is executed on the ShenWei processor by calling the adaptation interface corresponding to the PyTorch operator, so as to control the execution path of the operator.
[0013] In some of these embodiments, the step of transplanting the high-performance computing library further includes:
[0014] The step of optimizing the basic mathematical operation library: splitting the matrix operation into basic computing units, performing vectorization processing using the SIMD instructions of the ShenWei processor to achieve single-instruction multiple-data parallelism, and directly transmitting data between the main memory of the ShenWei processor and the slave-core processor by using the DMA method, reducing the data transmission overhead and improving the operation speed.
[0015] In some of these embodiments, the step of transplanting the high-performance computing library further includes:
[0016] The step of optimizing the deep learning operator library: dividing the input data and weight data into multiple data blocks and then performing operation operations. Among them, the size of each data block matches the cache size of the ShenWei processor, so that during the calculation process, the data in the cache can be used as much as possible, reducing the number of memory accesses, improving the calculation efficiency, and dividing the operation operations of the computationally intensive operators in the deep learning operator library swDNN into multiple operation sub-tasks. Each sub-task is independently executed in parallel on the slave cores of the ShenWei processor, and through the pipeline technology, the overlap of calculation and data transmission is realized, reducing the waiting time. Among them, the computationally intensive operators include, but are not limited to, the convolution operator Conv2d and the pooling operator Pooling.
[0017] In some of these embodiments, during the execution of the sub-tasks, the parallel programming model (such as OpenMP, MPI, etc.) of the ShenWei processor is used to achieve communication and synchronization between multiple cores.
[0018] In some of these embodiments, the step of transplanting the high-performance computing library further includes:
[0019] The step of optimizing the tensor operation library: developing the swTensor library to implement tensor operations, and performing memory access optimization and vectorization acceleration based on the ShenWei processor. The tensor operations include, but are not limited to, transpose operations, slice operations, and concatenation operations.
[0020] Based on the above steps, optimize the data layout and storage method according to the memory access characteristics of the Shenwei processor, and improve the performance of basic tensor operations through vectorization and parallelization techniques.
[0021] In some of these embodiments, the method further includes: configuring a performance counter interface, and reading core performance data through the hardware abstraction layer HAL provided by the Shenwei platform to collect the performance data of the Shenwei processor core.
[0022] Compared with the related art, the pytorch transplantation method based on the domestic Shenwei processor provided by the embodiments of the present application overcomes the differences in architecture and instruction set between the Shenwei processor and traditional CPUs and GPUs, ensuring that the pytorch framework can run smoothly on the Shenwei processor; aiming at the memory access mode and parallel computing model of the Shenwei processor, optimize the underlying implementation of pytorch to improve the computing performance of the pytorch framework on the Shenwei processor; re-implement or optimize the high-performance computing libraries and components in the pytorch framework that rely on specific hardware platforms to achieve the matching degree between the pytorch framework and the Shenwei processor, so as to improve the working performance of the pytorch framework; simplify the use of the programming environment and toolchain of the Shenwei processor, reduce the technical threshold for transplanting pytorch to the Shenwei processor; ensure the stability and reliability of the transplanted pytorch framework on the Shenwei processor, while maintaining compatibility with the original pytorch ecosystem, realizing the efficient transplantation and operation of the pytorch framework on the domestic Shenwei processor, and providing powerful computing support for deep learning models.
[0023] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. Description of the Drawings
[0024] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0025] Figure 1 is a flowchart of the transplantation method according to an embodiment of the present application. Detailed Embodiments
[0026] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be described and explained below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments provided by the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0027] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.
[0028] Referring to "embodiment" in the present application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appearing at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0029] Unless otherwise defined, the technical terms or scientific terms involved in the present application should be of the ordinary meaning understood by those with ordinary skills in the technical field to which the present application belongs. The words such as "a", "an", "one", "the", etc. involved in the present application do not indicate a limitation in quantity and can represent singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products, or devices. The terms "connected", "coupled", etc. involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0030] Based on the hardware specifications and programming manuals of the Shenwei processor, this embodiment of the application analyzes key features such as its instruction set, memory hierarchy, and parallel computing model, and designs an architecture adaptation layer on this basis.
[0031] Taking the Shenwei·TaihuLight as an example, the Shenwei processor uses the self-developed Shenwei instruction set, especially the SW64 architecture. In terms of the memory hierarchy, it adopts a hierarchical memory, with each computing unit configured with local memory, combined with global shared memory. The parallel computing model includes those based on MPI, OpenACC, or a unique parallel model, such as the slave core co-processing mode, as well as SIMD instructions, multi-level caches, high-bandwidth network interconnections, etc.
[0032] Porting the PyTorch framework to the Shenwei processor can make full use of its high-performance characteristics, accelerate the training and inference processes of deep learning models, and can achieve the efficient operation of the deep learning framework, make full use of the computing power of the Shenwei processor, and improve the execution efficiency of deep learning tasks. Therefore, in order to give full play to the computing power of the Shenwei processor, the PyTorch framework needs to be specifically improved and optimized. This not only includes the transplantation and re-implementation of existing libraries, but also the optimization of the scheduling and resource management of computing tasks to adapt to the unique hardware architecture and performance characteristics of the Shenwei processor.
[0033] This embodiment provides a method for porting PyTorch based on the domestic Shenwei processor. Figure 1 It is a flowchart of the method for porting PyTorch based on the domestic Shenwei processor according to the embodiment of the application. As Figure 1 shown, this process includes the following steps:
[0034] Step S1 of building the architecture adaptation layer: Configure an architecture adaptation layer swMath between the PyTorch framework and the Shenwei processor. The architecture adaptation layer swMath is provided with multiple adaptation interfaces, and a corresponding Shenwei implementation is defined for each PyTorch operator of the PyTorch framework, realizing a one-to-one mapping between the multiple adaptation interfaces and the PyTorch operators of the PyTorch framework. And when the adaptation interface is called, it can call the Shenwei processor for operator operations, so as to bridge the PyTorch framework and the Shenwei processor through this architecture adaptation layer swMath to bridge the architectural differences between the two. Among them, the PyTorch operators of the PyTorch framework, for example but not limited to, basic data operations such as add, mul, div, etc. that implement element-wise addition, subtraction, multiplication, and division, matrix operations such as matmul, etc. The adaptation interface for matrix multiplication operations will utilize the parallel computing power of the Shenwei processor to accelerate the operations. The same applies to conv2d and activation function operators (such as relu function, sigmoid function, etc.), which will not be elaborated here.
[0035] Based on this step, when performing the corresponding operator operations, PyTorch will call the adaptation interface of the corresponding operator to perform calculations on the Sunway processor, directly mapping the standard operator operations of PyTorch to the optimized implementation for the Sunway processor.
[0036] Figure 1 The process shown also includes: the high-performance computing library transplantation step S2, identifying the high-performance computing libraries in the PyTorch framework that rely on the CUDA platform, implementing the high-performance computing libraries and optimizing them based on the Sunway processor to obtain the optimized high-performance computing libraries. The optimized high-performance computing libraries include the basic mathematical operation library swBLAS, the deep learning operator library swDNN, and the tensor operation library swTensor. Integrate the basic mathematical operation library swBLAS, the deep learning operator library swDNN, and the tensor operation library swTensor into a unified dynamic link library libswmath.so and modify the connection configuration of PyTorch to link the dynamic link library libswmath.so with the libtorch.so library of the PyTorch framework to achieve correct linking of the system libraries and runtime libraries specific to the Sunway platform; among them, the high-performance computing library refers to a series of libraries that rely on the CUDA platform (Compute Unified Device Architecture) for parallel computing, such as cuDNN (CUDA Deep Neural Network Library) and cuBLAS (CUDA Basic Linear Algebra Subprograms). cuDNN is designed specifically for deep neural networks and provides highly optimized GPU acceleration functions for calculations such as forward and backward propagation, which can significantly improve the performance of deep learning models. cuBLAS, on the other hand, provides GPU-accelerated implementations of basic linear algebra operations, such as matrix multiplication and vector dot product, which are the basis of deep learning algorithms.
[0037] Compilation adaptation step S3: Use the cross-compilers sw-gcc and sw-g++ for the ShenWei processor to compile the C / C++ library file code on the X86 platform, generate the executable file for the ShenWei processor and run it on the ShenWei processor. Specifically, use the CMake tool to modify CMakeLists.txt to add cross-compilation configuration, achieve cross-compiling the C / C++ library file on the X86 platform, and complete the construction and registration of the Python extension module on the ShenWei platform. Control the file transfer and compilation process between the X86 platform and the ShenWei platform through an automated script. This script can use scp to transfer files and ssh to remotely execute commands. Run the automated script on the X86 platform, which is responsible for transferring the source code to the ShenWei platform and executing the CMake and compilation commands there. Once the compilation is completed, directly run on the ShenWei platform.
[0038] Use the CMake tool to modify CMakeLists.txt to add cross-compilation configuration: In the CMakeLists.txt file in the root directory of the CMake project, add or modify the following content to configure cross-compilation:
[0039] cmake
[0040] # Set the C / C++ compiler
[0041] set(CMAKE_C_COMPILER ${SW_TOOLCHAIN_PATH} / sw-gcc)
[0042] set(CMAKE_CXX_COMPILER ${SW_TOOLCHAIN_PATH} / sw-g++)
[0043] # Set the cross-compilation flag
[0044] set(CMAKE_CROSSCOMPILING TRUE)
[0045] Based on the above configuration, developers can perform python-related work on the ShenWei platform without being aware of the underlying compilation details.
[0046] Based on the above steps, a one-to-one mapping from PyTorch operators to the optimized implementation on the ShenWei processors is achieved, ensuring that each key operator can find its corresponding optimized version on the ShenWei processors and guaranteeing that the PyTorch framework can seamlessly utilize the high-performance computing capabilities of the ShenWei processors; the high-performance computing library is re-implemented in a modular refactoring manner, and the key computing functions in the high-performance computing library are vectorized for computationally intensive operations using the SIMD instruction set of the ShenWei processors and data is directly transferred between the main memory of the ShenWei processors and the slave core processors using the DMA method to improve the operation speed.
[0047] In some of these embodiments, in the step S1 of building the architecture adaptation layer, an environment variable is configured for each of the PyTorch operators, and by configuring the environment variable, the PyTorch operator is enabled to call the default CPU for computing or call the ShenWei processors for computing, so as to utilize the parallel computing capabilities of the ShenWei processors. Among them, when the value of the environment variable is 0, the default CPU implementation of the PyTorch architecture is called, and when the value of the environment variable is 1, the operation is executed on the ShenWei processors by calling the corresponding adaptation interface of the PyTorch operator to control the execution path of the operator. By way of example but not limitation, the environment variable name of the addition operator is configured as SW_ADD_ENABLE, the environment variable name of the conv2d matrix operation operator is configured as: SW_CONV2D_ENABLE, the environment variable name of the matmul matrix operation operator is configured as: SW_MATMUL_ENABLE, etc.
[0048] Based on the above environment variables, fine-grained control over the execution mode of each operator is achieved, thereby achieving an optimal balance between computing efficiency and computing accuracy.
[0049] In some of these embodiments, the step S2 of transplanting the high-performance computing library further includes:
[0050] Optimization step for the basic mathematical operation library: The matrix operation is split into basic computing units, vectorized using the SIMD instructions of the ShenWei processors to achieve single-instruction multiple-data parallelism, and data is directly transferred between the main memory of the ShenWei processors and the slave core processors using the DMA method to reduce the data transfer overhead and improve the operation speed. Among them, the matrix operation includes but is not limited to general matrix multiplication GEMM (General Matrix Multiply), general matrix-vector multiplication GEMV (General Matrix-Vector Multiply). The matrix operation is split into appropriate blocks as basic computing units, and the size of the block here is set according to the SIMD bit width. For example, for data of the FP32 type, the SIMD bit width is 512, then the size of this block is configured as a multiple of 512 / 32 = 16.
[0051] In addition, the embodiments of the present application also optimize the data access mode and calculation order for the cache hierarchy of the Shenwei processor. Specifically, it includes: analyzing the calculation dependency relationship of matrix operations, rearranging the calculation order to reduce the waiting time and data movement during the calculation process; for example but not limited to, in matrix multiplication, the order of multiplication operations can be adjusted so that each multiplication can use the data in the cache, thereby reducing memory access; through loop unrolling, multiple iterations are merged into one iteration to reduce the overhead of loop control and improve the pipeline efficiency of instructions, and through loop fusion, multiple related loops are merged into one loop to reduce the number of data transfers between memory and cache.
[0052] In some of the embodiments, the high-performance computing library transplantation step S2 further includes:
[0053] Deep learning operator library optimization step: After dividing the input data and weight data into multiple data blocks, perform arithmetic operations. Among them, the size of each data block matches the cache size of the Shenwei processor, so that during the calculation process, the data in the cache can be used as much as possible, reducing the number of memory accesses and improving the calculation efficiency. And the arithmetic operations of the computationally intensive operators in the deep learning operator library swDNN are divided into multiple operator sub-tasks, and each sub-task is independently executed in parallel on the slave cores of the Shenwei processor, and through pipeline technology, the overlap of calculation and data transmission is realized, reducing the waiting time. Among them, the computationally intensive operators include but are not limited to the convolution operator Conv2d and the pooling operator Pooling.
[0054] In some of the embodiments, during the execution of the sub-tasks, the parallel programming model (such as OpenMP, MPI, etc.) of the Shenwei processor is used to realize communication and synchronization between multiple cores. In the convolution operation, the convolution operator Conv2d can divide the processes such as loading of input data, loading of weight data, calculation, and storage of output data into different operator sub-tasks and optimize them through pipeline technology.
[0055] In some of the embodiments, the high-performance computing library transplantation step S2 further includes:
[0056] Tensor operation library optimization step: Develop the swTensor library to implement tensor operations, and perform memory access optimization and vectorization acceleration based on the Shenwei processor. The tensor operations include but are not limited to transpose operation, slicing operation, and splicing operation.
[0057] Among them, the memory access optimization and vectorization acceleration based on the Shenwei processor include:
[0058] Data layout optimization steps: divide the tensor into multiple tensor blocks, and each tensor block is aligned and allocated according to the LDM capacity (Local Device Memory) of the SW processor. Use the DMA engine to achieve efficient main memory - LDM data transfer; at the same time, utilize the asymmetric storage mode ASM of the SW processor to place frequently accessed data into the fast storage area. Among them, stride pre - calculation is adopted for continuous access patterns.
[0059] Among them, the memory access optimization and vectorization acceleration based on the SW processor also include:
[0060] Vectorization acceleration steps: utilize the SIMD instructions of the SW processor to vectorize tensor operations. Taking the transpose operation as an example, for the transpose of a 64×64 floating - point matrix, 4 - way vector parallelism is adopted.
[0061] Furthermore, it also includes: the application of pipeline parallel acceleration. Taking the tensor concatenation operation as an example, use DMA to load the input block from the main memory to the LDM, use vector instructions to reorganize the data, and then write the result block back to the main memory through DMA.
[0062] Based on the above steps, optimize the data layout and storage method according to the memory access characteristics of the SW processor, and improve the performance of basic tensor operations through vectorization and parallelization techniques.
[0063] In some embodiments, the method further includes: configuring a performance counter interface, reading core performance data through the hardware abstraction layer HAL provided by the SW platform to collect the performance data of the SW processor core. Specifically, define the types of performance events supported by the SW processor, including but not limited to the number of clock cycles, the number of retired instructions, memory bandwidth utilization rate, etc.; provide C / C++ interfaces for pytorch to call, and generate dynamic libraries for pytorch to call. In addition, it also supports the pytorch performance analysis tool, supporting hotspot function analysis and performance bottleneck identification. Specifically, combine the SW performance data with the Profiler of pytorch, add SW performance data collection logic in torch / autograd / profiler.py, associate pytorch operators (such as aten::matmul) with SW performance data, and combine the output of the pytorch Profiler with SW data to generate an interactive flame graph.
[0064] In addition, to facilitate developers to get started quickly, the embodiments of this application also write detailed documents and tutorials, including installation guides, API references, and usage cases.
[0065] In another embodiment, the embodiment of the present application also designs a test case to cover the core functions and common models of PyTorch, such as tensor operations, automatic differentiation, data loading, etc., to ensure the correctness of the transplanted framework on the Shenwei processor. After testing, after the successful development of PyTorch, the Megatron framework was transplanted at the same time, and a model with 175 billion parameters of Llama was successfully trained on the Shenwei supercomputer, verifying the compatibility and stability of PyTorch.
[0066] In another embodiment, the embodiment of the present application also provides a model migration guide, including code modification suggestions and practical logic, to help users migrate the model to the Shenwei processor.
[0067] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0068] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0069] The above-described embodiments only represent several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for porting PyTorch based on domestic Shenwei processors, characterized in that, Including: Steps for building the architecture adaptation layer: Configure an architecture adaptation layer swMath between the PyTorch framework and the Sunway processor. The architecture adaptation layer swMath is provided with multiple adaptation interfaces, and a corresponding Sunway implementation is defined for each PyTorch operator in the PyTorch framework, realizing a one-to-one mapping between the multiple adaptation interfaces and the PyTorch operators in the PyTorch framework. When the adaptation interface is called, the Sunway processor can be called for operator operations; Steps for transplanting the high-performance computing library: Identify the high-performance computing libraries in the PyTorch framework that depend on the CUDA platform, implement the high-performance computing libraries and optimize them based on the Sunway processor to obtain the optimized high-performance computing libraries. The optimized high-performance computing libraries include the basic mathematical operation library swBLAS, the deep learning operator library swDNN, and the tensor operation library swTensor. Integrate the basic mathematical operation library swBLAS, the deep learning operator library swDNN, and the tensor operation library swTensor into a unified dynamic link library libswmath.so and link the dynamic link library libswmath.so with the libtorch.so library of the PyTorch framework. The high-performance computing library refers to the library for parallel computing that depends on the CUPA platform, including cuDNN and cuBLAS. cuDNN provides highly optimized GPU acceleration functions for deep neural networks and is used for forward propagation and backward propagation calculations. cuBLAS is used to provide GPU acceleration implementations for basic linear algebra operations; Steps for compilation adaptation: Use the cross-compilers sw-gcc and sw-g++ for the Sunway processor to compile the C / C++ library file code on the X86 platform, generate the executable file for the Sunway processor and run it on the Sunway processor. Use the CMake tool to modify CMakeLists.txt to add cross-compilation configurations, realize cross-compiling C / C++ library files on the X86 platform, and complete the construction and registration of Python extension modules on the Sunway platform. Control the file transfer and compilation process between the X86 platform and the Sunway platform through an automated script. This automated script uses scp to transfer files and uses ssh to remotely execute commands. Run the automated script on the X86 platform, which is responsible for transferring the source code to the Sunway platform and executing the CMake and compilation commands. Once the compilation is successful, directly run on the Sunway platform; The steps for transplanting the high-performance computing library further include: Steps for optimizing the deep learning operator library: Divide the input data and weight data into multiple data blocks and then perform operation operations, and divide the operation operations of the computationally intensive operators in the deep learning operator library swDNN into multiple operation subtasks, and each subtask is independently executed in parallel on the slave cores of the Sunway processor.
2. The pytorch porting method based on domestic Shenwei processor according to claim 1, wherein In the steps for building the architecture adaptation layer, configure an environment variable for each of the PyTorch operators, and enable the PyTorch operators to call the default CPU for calculation or call the Sunway processor for calculation through the configuration of the environment variable.
3. The pytorch porting method based on domestic Shenwei processor according to claim 2, characterized in that, When the value of the environment variable is 0, the default CPU implementation of the PyTorch architecture is called. When the value of the environment variable is 1, operations are executed on the ShenWei processor by calling the corresponding adaptation interface of the PyTorch operator to control the execution path of the operator.
4. The pytorch transplantation method based on domestic Shenwei processor according to claim 2, characterized in that, The high-performance computing library transplantation steps further include: The basic mathematical operation library optimization step, which splits matrix operations into basic computing units, performs vectorization using the SIMD instructions of the ShenWei processor, and directly transfers data between the main memory of the ShenWei processor and the slave core processor by using the DMA method.
5. The pytorch porting method based on domestic Shenwei processor according to claim 1, characterized in that, During the execution of the sub-tasks, the parallel programming model of the ShenWei processor is used to achieve communication and synchronization between multiple cores.
6. The pytorch transplantation method based on domestic Shenwei processor according to claim 1, characterized in that The high-performance computing library transplantation steps further include: The tensor operation library optimization step, which develops the swTensor library to implement tensor operations, and performs memory access optimization and vectorization acceleration based on the ShenWei processor.
7. The pytorch porting method based on the domestic Shenwei processor according to claim 5, characterized in that It also includes: Configuring a performance counter interface to read core performance data through the hardware abstraction layer HAL provided by the ShenWei platform.
8. The pytorch transplantation method based on the domestic Shenwei processor according to claim 6, characterized in that The tensor operations include transpose operation, slicing operation, and concatenation operation.