An Implicit Data Dynamic Reuse Method for Many-Core Distributed Local Memory
By marking data variables in the compilation instructions and creating address mapping tables, the management problem of data reuse in distributed local memory systems is solved, and the dynamic reuse of data between multiple computing cores is realized, and the program performance is improved.
Patent Information
- Application Number
- CN202110453214.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-04-26
AI Technical Summary
On a multi-core system with distributed local memory, the existing implicit acceleration programming language cannot effectively manage data reuse between local memory of multiple computing cores, resulting in multiple transmissions of data between main memory and local memory, reducing program performance.
By marking data variable names or array offsets in the compilation instructions, creating main and local address mapping tables, and dynamic reuse of data in multiple computing cores under the guidance of the compiler, including processing of undistributed and distributed data, reducing data transmission overhead.
Multi-level data reuse in multiple acceleration segments and the same acceleration segment is realized, reducing data transmission overhead and improving program performance.
Smart Images

Figure CN114217811B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an implicit data dynamic reuse method for a many-core distributed local memory, and belongs to the technical field of high-performance computing. Background Art
[0002] Generally, there are multi-level storage systems on many-core systems, including off-chip main memory and a distributed local memory with independent addressing for each or every few computing cores. The local memory is completely managed by the user or software. The capacity of the main memory is large, while the capacity of the local memory is small. Compared with the main memory, the computing cores can access the local memory faster and with higher memory access bandwidth. Most of the data of the program is in the main memory. The accelerating threads transfer the data required by the current program segment to the local memory, and write it back to the main memory after the calculation is completed, so as to vacate the space in the local memory for the data of the next calculation program segment.
[0003] For implicit acceleration programming languages such as OpenACC and OpenMP, the management and transmission of data are also implicit. Programmers manage the local memory space through data copy instructions such as copy / copyin / copyout, and control the data transmission between the main memory and the local memory. Therefore, in the implicit acceleration programming mode, data can only be dynamically reused between different program blocks through compilation directives. The compiler automatically generates control codes such as space management and data transmission according to the information of the compilation directives.
[0004] Some implicit programming languages such as OpenACC provide data region compilation directives to support data reuse between multiple acceleration segments, but mainly for non-distributed device memory shared by all accelerating computing cores such as GPUs. On a many-core system with a distributed local memory, these compilation directives cannot describe the distributed reuse of data in the local memories of multiple computing cores, and can only support such coarse-grained data reuse between multiple acceleration segments, but cannot support data reuse between multiple computing modules within the same acceleration segment.
[0005] Using implicit programming languages such as OpenACC to accelerate traditional application programs on heterogeneous many-core systems is a relatively common method for developing many-core applications. On a many-core system with a distributed local memory (LDM), due to the limitation of the local memory capacity, each program block for accelerating calculation needs to copy data from the main memory to the local memory, and write the data back to the main memory after the block calculation is completed. Only a small amount of data can reside in the local memory for a long time. If a piece of data is used in multiple program blocks, then the multiple repeated transmissions of the data between the main memory and the local memory reduce the performance of the program. Such program blocks may be multiple acceleration segments of the application program, or different calculation parts within the same acceleration segment. Summary of the Invention
[0006] The object of the present invention is to provide an implicit data dynamic reuse method for a many-core distributed local memory, which can not only dynamically apply for and release, make full use of the limited local memory space, but also enable the reused data to stay in the local memory as long as possible, reduce the overhead of data transmission, and improve the performance of the program.
[0007] To achieve the above object, the technical solution adopted by the present invention is: to provide an implicit data dynamic reuse method for a many-core distributed local memory. When accelerating calculations with a many-core processor, the calculation tasks are decomposed and executed on several acceleration calculation cores, so that each acceleration calculation core completes a part of the calculation tasks respectively;
[0008] Before accelerating the calculation, the data required for the acceleration calculation is transmitted and stored in the distributed local memories of multiple acceleration calculation cores;
[0009] When accelerating the calculation, access the data copies in the local memory;
[0010] Achieve multi-level dynamic reuse of the data in the distributed local memory through compiler directives;
[0011] It includes the following steps:
[0012] S1. According to the data access pattern, data volume, and the capacity of the local memory of the acceleration calculation core in the acceleration calculation, process the storage methods of the data in the local memory of the acceleration calculation core respectively, including two types: non-distributed data and distributed data:
[0013] S11. For data with a small data volume, apply for space in the local memory of each acceleration calculation core, and transmit the same data to the local memory of each acceleration calculation core, which is called non-distributed data;
[0014] S12. For data with a large data volume and different access intervals for each acceleration calculation core, apply for space in the local memory of each acceleration calculation core respectively, and transmit the data required by each respectively, which is called distributed data;
[0015] S2. Mark the data variable names or array offsets that are also reused in other functions and other calculation tasks in the program through "register compiler directives", and the scope of action of the compiler directives is the life cycle of the data;
[0016] S3. The compiler creates a mapping table of the main memory address and local memory address of the reused data according to the data variable names or array offsets specified by the "register compiler directives". The mapping table respectively records the address range, length of the data in the main memory and the corresponding address range in the local memory, and stores the mapping table in the local memory for quick query;
[0017] S4. When accessing data that has already been stored in the local memory in other functions or in another computing task, indicate the data variable name or array offset that needs to be reused through "reuse directives", and at the same time mark the code segment that needs to access the data stored in the local memory. The reused data can be distributed or non-distributed data;
[0018] S5. The compiler generates a query statement based on the information in the "reuse directive", and queries the address of the data in the local memory according to the main memory address of the data to be reused;
[0019] S6. The compiler replaces the access to the main memory variable of the reused data in the code segment marked by the "reuse directive" with the access to the corresponding space in the local memory of the data;
[0020] S7. Mark the end of the life cycle of the reused data through the "release directive", delete the relevant entries in the address mapping table, and release the corresponding local memory space.
[0021] The further improved solutions in the above technical solutions are as follows:
[0022] 1. In the above solution, in S11, the data with a small amount is the data that does not exceed the local memory capacity of the acceleration computing core.
[0023] 2. In the above solution, in S12, the data with a large amount is the data that exceeds the local memory capacity of the acceleration computing core.
[0024] Due to the application of the above technical solutions, the present invention has the following advantages compared with the prior art:
[0025] The present invention provides an implicit data dynamic reuse method for a many-core distributed local memory, which realizes the reuse of distributed and non-distributed local memory data within multiple effective ranges. It can not only dynamically apply for and release, make full use of the limited local memory space, but also keep the reused data in the local memory as long as possible, reduce the overhead of data transmission, and improve the performance of the program. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Attached Figure 1 is a schematic diagram of an implicit data dynamic reuse method for a distributed local memory according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] Embodiment: The present invention provides an implicit data dynamic reuse method for a many-core distributed local memory. When performing acceleration computing with a many-core processor, the computing task is decomposed and executed on a certain number of acceleration computing cores, so that each acceleration computing core completes a part of the computing task respectively, thereby shortening the processing time of the computing task;
[0028] Before performing accelerated computing, transfer and store the data required for accelerated computing to the distributed local memory (local local memory) of multiple accelerated computing cores;
[0029] During accelerated computing, access the data copies in the local memory and make full use of the high-speed access performance of the local memory;
[0030] Implement multi-level dynamic reuse of the data in the distributed local memory through compiler directives;
[0031] It includes the following steps:
[0032] S1. According to the data access pattern, data volume, and the capacity of the local memory of the accelerated computing core during accelerated computing, process the storage method of the data in the local memory of the accelerated computing core respectively, including two types: non-distributed data and distributed data:
[0033] S11. For data with a small data volume, apply for space in the local memory of each accelerated computing core and transfer the same data to the local memory of each accelerated computing core, which is called non-distributed data;
[0034] S12. For data with a large data volume and different access intervals for each accelerated computing core, apply for space in the local memory of each accelerated computing core respectively and transfer the data they need respectively, which is called distributed data;
[0035] S2. Use "register compiler directive" to mark the data variable names or array offsets that may be reused in other functions or other computing tasks in the program. The scope of action of the compiler directive is the life cycle of the data;
[0036] S3. The compiler creates a mapping table of the main memory address and local memory address of the reusable data according to the data variable names or array offsets specified by the "register compiler directive". The mapping table records the address range, length in the main memory and the corresponding address range in the local memory of the data respectively, and stores the mapping table in the local memory for quick query;
[0037] S4. When accessing the data that has been stored in the local memory in other functions or in another computing task, use "reuse compiler directive" to specify the data variable names or array offsets that need to be reused, and at the same time mark the code segment that needs to access this part of the data. The reusable data can be distributed data or non-distributed data;
[0038] S5. The compiler generates a query statement according to the information of the "reuse compiler directive" and queries the address of the data in the local memory according to the main memory address of the data to be reused;
[0039] S6. The compiler replaces the accesses to the main memory variables of the reusable data in the code segments marked with "reuse pragma" with accesses to the corresponding spaces in the local memory for the data.
[0040] S7. Mark the end of the life cycle of the reusable data through "release pragma", delete the relevant entries in the address mapping table, and release the corresponding local memory space.
[0041] In S11, the data with a small amount is the data that does not exceed the local memory capacity of the acceleration computing core.
[0042] In S12, the data with a large amount is the data that exceeds the local memory capacity of the acceleration computing core.
[0043] The further explanations for the above embodiments are as follows:
[0044] In the many-core system with distributed local memory of the present invention, adopting the implicit method of pragma, through the program transformation guided by compilation combined with the underlying runtime library, the dynamic reuse of local memory data with non-distributed and distributed attributes is realized, and the multi-level dynamic reuse between multiple acceleration segments and between multiple computing modules within the same acceleration segment can be achieved, improving the performance of the program;
[0045] A mechanism for the dynamic residence of data in local memory is provided, enabling multiple program blocks to reuse the same data and reducing the multiple repeated transmissions of the same data;
[0046] The present invention proposes a method for dynamically reusing local memory data guided by compilation, dynamically maintaining the storage space of local memory and the address mapping table according to the reuse information of the data, so as to realize the dynamic reuse of distributed and non-distributed data at multiple levels.
[0047] The specific implementation solution is as follows:
[0048] 1. Mark the data that needs to be reused in the program through pragma and distinguish the distribution attributes of the data (distributed or non-distributed). The scope of action of the pragma is the life cycle of the data.
[0049] 2. The compiler performs code transformation for data transmission according to the pragma, and in combination with the runtime library, allocates storage space for the reusable data and performs data transmission from the main memory to the local memory, and this space is valid during the life cycle of the data.
[0050] 3. Create a mapping table for the main memory address and local memory address of the reusable data, and store this table in the local memory for quick query.
[0051] 4. The compiler generates conversion statements from the main memory address to the local memory address of the reusable variables within the scope of action of the pragma, and replaces all accesses to the main memory address with the local memory address.
[0052] 5. After the end of the scope of the pragma, recycle the space of the relevant reusable data and delete the relevant entries in the address mapping table.
[0053] When adopting the above-mentioned implicit data dynamic reuse method for many-core distributed local memory, it realizes the reuse within multiple effective ranges for both distributed and non-distributed local memory data. It can not only dynamically apply for and release, make full use of the limited local memory space, but also keep the reusable data in the local memory for as long as possible, reduce the overhead of data transmission, and improve the performance of the program.
[0054] To facilitate a better understanding of the present invention, the terms used in this article will be briefly explained below:
[0055] Accelerating computing core: The accelerating computing component of a many-core processor, which can load the code and data that need to be accelerated for computing onto the computing core for execution.
[0056] Local Data Memory (LDM) of the accelerating computing core: Each or every few accelerating computing cores have an independently addressed on-chip memory, and the capacity is generally not large.
[0057] Accelerator thread: A program entity running on the accelerating computing core.
[0058] Implicit programming language: Add pragmas in the program, and realize the required functions through the recognition of pragmas and program transformation during compilation. Typical implicit programming languages include OpenACC, OpenMP, etc.
[0059] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. All equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. An implicit data dynamic reuse method for many-core distributed local memory. When using a many-core processor for accelerated computing, the computing task is decomposed and executed on several accelerated computing cores, so that each accelerated computing core completes a part of the computing task respectively; It is characterized in that: Before accelerated computing, the data required for accelerated computing is transmitted and stored in the distributed local memory of multiple accelerated computing cores; During accelerated computing, access the data copy in the local memory; Realize multi-level dynamic reuse of the data in the distributed local memory through compilation directives; It includes the following steps: S1. According to the data access pattern, data volume, and the capacity of the local memory of the accelerated computing core in the accelerated computing, the storage methods of the data in the local memory of the accelerated computing core are processed respectively, including two types: non-distributed data and distributed data: S11. For data with a small data volume, apply space in the local memory of each accelerated computing core, and transmit the same data to the local memory of each accelerated computing core, which is called non-distributed data; S12. For data with a large data volume and different access intervals for each accelerated computing core, apply space in the local memory of each accelerated computing core respectively, and transmit the data required by each respectively, which is called distributed data; S2. Use "register compilation directive" to mark the data variable names or array offsets that are also reused in other functions and other computing tasks in the program. The scope of action of the compilation directive is the life cycle of the data; S3. The compiler creates a mapping table of the main memory address and local memory address of the reused data according to the data variable names or array offsets specified by the "register compilation directive". The mapping table respectively records the address range, length of the data in the main memory and the corresponding address range in the local memory, and stores the mapping table in the local memory for quick query; S4. When accessing the data that has been stored in the local memory in other functions or in another computing task, use "reuse compilation directive" to specify the data variable names or array offsets that need to be reused, and at the same time mark the code segment that needs to access the data that has been stored in the local memory. The reused data is distributed data or non-distributed data; S5. The compiler generates a query statement according to the information of the "reuse compilation directive", and queries the address of the data in the local memory according to the main memory address of the data to be reused; S6. The compiler replaces the access to the main memory variable of the reused data in the code segment marked by the "reuse compilation directive" with the access to the corresponding space of the data in the local memory; S7. Mark the end of the life cycle of the reused data through "release compilation directive", delete the relevant entries in the address mapping table, and release the corresponding local memory space.
2. The implicit data dynamic reuse method for many-core distributed local memory according to claim 1, wherein: In S11, the data with a small data volume is the data that does not exceed the capacity of the local memory of the accelerated computing core.
3. An implicit data dynamic reuse method for a many-core distributed local memory according to claim 1 or 2, characterized in that: In S12, the data with a large data volume is the data that exceeds the capacity of the local memory of the accelerated computing core.
Citation Information
Patent Citations
Method and device of pre-fetching data of compiler
CN102981883A
Data distribution and local optimization method for heterogeneous many-core architecture multi-level storage structure
CN103226487A