Slave core memory access optimization method based on new generation SW many-core processor
Through data preprocessing and selecting appropriate memory access solutions, the problem of limited core LDM space and data volume exceeding the maximum capacity on Shenweizhong core processors is solved, and efficient memory access and program execution efficiency is achieved.
Patent Information
- Application Number
- CN202510626634.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
When optimizing programs on Shenweizhong core processors, the prior art faces problems such as limited nuclear LDM space, data volume exceeding the maximum capacity, and blindly transmitting data, resulting in waste of storage resources and large communication overhead, resulting in inefficiency of the program.
Through data preprocessing, the array subscript is calculated in advance, deduplication is deduplicated and a new array is built, to determine whether the data volume exceeds the maximum spatial capacity of the LDM continuous sharing mode in the entire array, and to select a suitable memory access solution, including five LDM continuous sharing modes and general memory access methods, and to optimize the memory access process from the core.
It effectively solves the problem of limited space of slave core LDM and data volume exceeding the maximum capacity, reduces the communication overhead between master core and slave core, improves program execution efficiency, and makes full use of space locality.
Smart Images

Figure CN120144503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an optimization method for the memory access of slave cores based on a new generation of ShenWei many-core processors, belonging to the technical field of electronic information. Background Art
[0002] The computing power of the new generation of ShenWei supercomputers is provided by the domestic many-core SW26010pro processor. The architecture of the SW26010pro processor is as Figure 1 shown. The processor includes 6 core groups (Core-Group, CG), and each core group includes 1 management core (Management Processing Element, MPE, also known as the master core) and 1 computation processing elements clusters (CPE clusters, also known as the slave core array). Each slave core array consists of 64 slave cores (Computing Processing Element, CPE) in an 8×8 mesh format. In the slave core array, each slave core cluster consists of four slave cores, and four adjacent slave cores share a slave core cluster management component and integrate a router. Each core group has its own memory controller (Memory Controller, MC), which is connected to 16GB of DDR4 memory with a bandwidth of 51.2GB / s.
[0003] The storage within a core group of the ShenWei many-core chip is divided into the main memory located in the master core and the LDM located in each slave core privately. Each slave core has a local data memory (Local Data Memory, LDM) with a capacity of 256 KB. The memory access method of the slave cores and the LDM space division are as Figure 2 shown. In the face of memory access requirements, the slave cores can access the main memory in a direct discrete manner (i.e., gld / gst); they can also access the main memory in batches through DMA, first moving the data from the main memory to the locally accessible LDM of the slave cores. When the slave cores are computing, only a memory access to the LDM needs to be initiated.
[0004] The LDM space of the slave cores can be divided into three parts: private space, continuous shared space, and slave core data cache space, and the three parts of the space can be adjusted in different grades. When part of the space is configured as a level-1 data cache, the slave cores can perform cache access to this space by loading or storing the cacheable space of the main memory, and the cache line size is 256 bytes. The SW26010pro processor supports three data cache capacity configurations of 0KB, 32KB, and 128KB.
[0005] The LDM continuous shared space is a space used by slave cores for LDM shared access within the array, and a slave core can access the LDM spaces of other slave cores. In the LDM continuous shared mode, each slave core within the array provides an LDM space of the same capacity for continuous addressing, and the capacity of the LDM space provided by each slave core for sharing can be adjusted in 7 capacities, such as 0 KB, 4 KB, 8 KB, 16 KB, 32 KB, 64 KB, and 128 KB.
[0006] The continuous shared mode includes 5 modes, namely single cluster (a slave core cluster is a 2*2 slave core array), double cluster (a 4*2 array composed of two slave core clusters), quadruple cluster (an 8*2 array composed of four slave core clusters), full array cluster (full array sharing, continuous addressing by slave core cluster), and full array row (full array sharing, continuous addressing by row), which is convenient for the nearby access of different working sets.
[0007] The method for optimizing a single-core group using the ShenWei many-core processor mainly involves the main core starting the program and migrating the hot code segments that can be optimized in parallel in the application program to the slave cores to achieve the effect of parallel acceleration. During the execution of the hot code on the slave cores, the slave cores must first obtain the data required for the computing task from the main memory through discrete memory access (i.e., gld / gst) or DMA. Only after the slave cores obtain the required data can they continue to execute the subsequent code.
[0008] There are two ways for slave cores to obtain data. First, the slave cores can obtain the data required for the computing task from the main memory through direct discrete memory access (i.e., gld / gst) without saving it to the slave core LDM. When using this part of the data, the slave cores need to maintain communication with the main core at all times to obtain data from the main memory. Second, the slave cores can obtain the data required for the computing task from the main memory through DMA and store it in the slave core LDM. When using this part of the data, the slave cores only need to call it from the LDM.
[0009] Facing the above two ways of obtaining data, generally the second way is adopted. Even though the second way requires additional space to store data, compared with the first way, the second way reduces the communication overhead between the slave cores and the main core, and it is more convenient and fast to call data, making the overall program run more efficiently.
[0010] However, there are three problems with the above method. First, the data required for the computing tasks in the application program is basically array data. When normally optimizing the program for the purpose of convenience and speed, when obtaining array data from the slave core, the actual data required for the computing tasks is not considered. Instead, the entire array data is transferred to the slave core's LDM through DMA. This often ignores the rarity of the slave core's LDM resources, and blindly transferring extra data will increase the communication overhead, resulting in wasted storage resources and low program efficiency. Second, the amount of array data in the application program is generally very large. Since the LDM space resources of the slave core are very limited, each slave core only has 256KB of LDM storage space. There is a situation where the amount of array data exceeds 256KB and cannot be stored in the slave core's private LDM. In response to this situation, the LDM continuous sharing mode can be enabled, and the appropriate LDM continuous sharing mode and space capacity can be selected according to different data volumes. Although the LDM continuous sharing mode can be enabled and the space capacity can be set to store the array data, since the data is not preprocessed, the actual required data volume cannot be determined, and blindly transferring the entire array to the LDM will result in a significant reduction in the utilization rate of storage resources. Third, if there is a situation where the data volume exceeds the maximum capacity in the LDM continuous sharing mode. That is, when the full-array LDM continuous sharing mode is enabled and a 128KB continuous sharing space capacity is selected, and the total space capacity reaches 8192KB (64 * 128KB) but still cannot store the array data. At this time, the method of direct discrete memory access (i.e., gld / gst) is generally adopted, and the Cache space capacity is adjusted according to the requirements to speed up the direct discrete memory access. However, even if the slave core's Cache is used to improve the efficiency of direct discrete memory access, the final efficiency of the program is still not high. Summary of the Invention
[0011] Aiming at the deficiencies of the prior art, the present invention provides a method for optimizing the memory access of the slave core based on the new generation Shenwei many-core processor; When optimizing a program on the SW64 processor, the computing tasks are assigned to slave cores, and the slave cores obtain the data required for the computing tasks through direct discontinuous memory access (i.e., gld / gst) or DMA. When normally optimizing a program, for the sake of convenience and simplicity, usually the entire array of data required for the computing tasks is transferred to the slave core's LDM through DMA. However, when the data volume of the entire array exceeds the maximum capacity of the LDM continuous sharing mode, the data cannot be transferred to the LDM through DMA. If a general memory access method is blindly adopted, that is, direct discontinuous memory access and using the Cache space, it will not only cause serious problems of communication overhead, but also may have a situation where the data is too discontinuous, and the spatial locality cannot be fully utilized, resulting in low memory access efficiency of the program. For a program with the above problems, selecting the method of the present invention for optimization can effectively solve this problem and improve the program execution efficiency. At the same time, the present invention also designs an automated interface STMA (Sunway Thread Memory Access), which is convenient for programmers to directly call and use.
[0012] The present invention specifically solves the situation where the limited local storage space of the slave core leads to the slave core needing to frequently access the master core's storage resources discontinuously, and successfully reduces the communication overhead between the master core and the slave core, and efficiently utilizes the spatial locality, and applies it to the slave core acceleration scheme.
[0013] The technical solution of the present invention is as follows: A method for optimizing memory access of the slave core based on a new generation of SW64 processor, including: First, design five memory access schemes and a general memory access method according to the LDM continuous sharing mode; Secondly, perform data preprocessing, calculate the array subscripts in advance in the master core for the actual required array data in the computing tasks, remove the duplicate subscripts and construct a new array; Finally, determine whether the actual data volume exceeds the maximum space capacity of 8192KB of the full-array LDM continuous sharing mode; if it exceeds, adopt a general memory access method, that is, direct discontinuous memory access and set the Cache space capacity of 128KB; if it does not exceed, select different memory access schemes according to the data volume.
[0014] According to the preference of the present invention, design five memory access schemes and a general memory access method according to the LDM continuous sharing mode; including: The space of the slave core LDM is designed into two categories: type-a LDM space and type-b LDM space; in the type-a LDM space, 128KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128KB is used as the continuous shared space of the slave core LDM; in the type-b LDM space, 128KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128KB is used as the data cache space of the slave core; the private space of the slave core is used to store other data during the execution of the computing task; The five memory access schemes are as follows: Scheme 1: All 256KB of the space of the slave core LDM is used as the private space of the slave core, 128KB of which is used to store the data required for the computing task, and the other 128KB is used to save other data during the execution of the computing task; the slave core transfers the actually required data of the computing task to the slave core LDM through DMA; Scheme 2: The single-cluster mode in the continuous sharing mode of the slave core LDM is enabled, the type-a LDM space is used, and the slave core transfers the actually required data of the computing task to the LDM continuous sharing space through the slave core DMA; Scheme 3: The dual-cluster mode in the continuous sharing mode of the LDM is enabled, the type-a LDM space is used, and the slave core transfers the actually required data of the computing task to the LDM continuous sharing space through the slave core DMA; Scheme 4: The four-cluster mode in the continuous sharing mode of the LDM is enabled, the type-a LDM space is used, and the slave core transfers the actually required data of the computing task to the LDM continuous sharing space through the slave core DMA; Scheme 5: The full-array cluster mode in the continuous sharing mode of the LDM is enabled, the type-a LDM space is used, and the slave core transfers the actually required data of the computing task to the LDM continuous sharing space through the slave core DMA; The general memory access method means that the slave core uses the type-b LDM space through direct discrete memory access.
[0015] Preferably according to the present invention, data preprocessing is performed in the master core; it includes: Calculating the subscripts of the original array A; the original array A refers to the array originally existing in the main memory in the program, and the original array A stores the data required for the computing task. After the task is divided among the slave cores, the slave cores need to obtain data from the original array A to execute the computing task; Removing duplicates using a boolean array: By using a boolean array to determine whether the subscripts of the original array A have appeared. If they have appeared, skip them; if not, record the subscripts of the original array A; Constructing a new array B to store the actually required data of the computing task; Establishing a mapping relationship between the subscripts of the new array B and the computing task, and constructing another array C to store the mapping relationship; Summarize and calculate the actual data volume N required for the task, with the data volume unit being KB.
[0016] Further preferably, construct a new array B to store the data actually required for the calculation task; including: creating a new array B in the main memory, and after de-duplicating using a boolean array, assign the data in the original array A to the new array B.
[0017] Preferably according to the present invention, select a memory access scheme according to the summarized data volume N in data preprocessing; including: When 0KB < N ≤ 128KB, select Scheme 1; when 128KB < N ≤ 512KB, select Scheme 2; when 512KB < N ≤ 1024KB, select Scheme 3; when 1024KB < N ≤ 2048KB, select Scheme 4; when 2048KB < N ≤ 8192KB, select Scheme 5; when 8192KB < N, select a general memory access method; After completing the above steps, the main core proceeds according to the selected memory access scheme, and transfers the new array B and the mapping array C to the slave core LDM through DMA and executes subsequent calculation tasks.
[0018] Preferably according to the present invention, implement the memory access optimization method for the slave core based on the new generation of ShenWei multi-core processors through the automated interface STMA.
[0019] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the memory access optimization method for the slave core based on the new generation of ShenWei multi-core processors.
[0020] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the memory access optimization method for the slave core based on the new generation of ShenWei multi-core processors.
[0021] The beneficial effects of the present invention are: When optimizing a program on the SW64 processor, the computing tasks are assigned to slave cores, and the slave cores obtain the data required for the computing tasks through direct discrete memory access (i.e., gld / gst) or DMA. When normally optimizing a program, for the sake of convenience and simplicity, usually the entire array of data required for the computing tasks is transferred to the LDM of the slave core through DMA. However, when the data volume of the entire array exceeds the maximum capacity of the LDM continuous sharing mode, the data cannot be transferred to the LDM through DMA. If a general memory access method is blindly adopted, that is, direct discrete memory access and using the Cache space, it will not only cause serious problems of communication overhead, but also there may be a situation where the data discreteness is too large, and the spatial locality cannot be fully utilized, resulting in low memory access efficiency of the program. For a program with the above problems, selecting the method of the present invention for optimization can effectively solve this problem and improve the program execution efficiency. At the same time, the present invention also designs an automated interface STMA (Sunway Thread Memory Access), which is convenient for programmers to directly call and use. Description of the Drawings
[0022] Figure 1 It is a schematic diagram of the architecture of SW26010pro; Figure 2 It is a schematic diagram of the slave core accessing the main memory method and the LDM space division; Figure 3 It is a schematic diagram of the flow of the slave core memory access optimization method based on the new generation of SW64 processor of the present invention; Figure 4 It is a schematic diagram of the space division of class a and class b; Figure 5 It is a schematic diagram of the space of Scheme 1; Figure 6 It is a schematic diagram of the space of Scheme 2; Figure 7 It is a schematic diagram of the space of Scheme 3; Figure 8 It is a schematic diagram of the space of Scheme 4; Figure 9 It is a schematic diagram of the space of Scheme 5; Figure 10 It is a schematic diagram of the data preprocessing flow of the present invention. Detailed Embodiment
[0023] The present invention will be further defined below in conjunction with the drawings of the specification and embodiments, but not limited thereto.
[0024] Term Explanation: 1. gld / gst memory access: Discrete memory access. At the slave core side, data can be directly retrieved from the main memory or stored in the main memory through the address of the main memory, without the need to be saved to the local local data storage space of the slave core.
[0025] 2. Cache: A smaller and faster storage device that serves as a buffer for a larger and slower storage device, thereby improving data access speed. A core can use the Cache to accelerate the discrete access speed of reading or storing data in the main memory.
[0026] 3. LDM: Local Data Memory. Each slave core of the ShenWei many-core processor has a high-speed local data storage space LDM, and the total capacity of the LDM is 256KB.
[0027] 4. DMA: Direct Memory Access. It is an operation initiated by a single slave core for bulk data transfer between the main memory and the local memory LDM of the slave core.
[0028] Embodiment 1 A method for optimizing the memory access of the slave core based on the new generation of ShenWei many-core processor includes: First, five memory access schemes and a general memory access method are designed according to the LDM continuous sharing mode (single cluster, dual cluster, four cluster, full array). Second, data preprocessing is performed. For the array data actually required in the computing task, the array subscripts are calculated in advance in the main core, duplicate subscripts are removed, and a new array is constructed. Finally, it is judged whether the actual data volume exceeds the maximum space capacity of 8192KB (64 * 128KB) of the full array LDM continuous sharing mode. If it exceeds, a general memory access method is adopted, that is, direct discrete memory access (i.e., gld / gst) and a Cache space capacity of 128KB is set. If it does not exceed, different memory access schemes are selected according to the data volume.
[0029] This innovative method solves the problem of low efficiency of direct discrete memory access of the slave core, makes full use of spatial locality, and reduces the communication overhead between the main and slave cores, thereby greatly improving the execution efficiency of the program. The specific process is as Figure 3 shown.
[0030] Embodiment 2 The method for optimizing the memory access of the slave core based on the new generation of ShenWei many-core processor according to Embodiment 1 is different in that: Five memory access schemes and a general memory access method are designed according to the LDM continuous sharing mode, including: When designing the memory access scheme, five memory access schemes and a general memory access method are designed in combination with the data volume based on three factors: the space division of the slave core LDM, the LDM continuous sharing mode, and the space capacity.
[0031] According to the characteristics of the hardware architecture of the new generation of Shenwei supercomputers and Shenwei many-core processors, the LDM space of slave cores is divided into three parts: private space, continuous shared space, and slave core data Cache space, and the three parts of the space can be adjusted in different grades. To describe the memory access scheme more vividly and intuitively, as Figure 4 shown, the present invention designs the space of the slave core LDM into two types: type-a LDM space and type-b LDM space; in the type-a LDM space, 128KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128KB is used as the continuous shared space of the slave core LDM; in the type-b LDM space, 128KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128KB is used as the data Cache space of the slave core; the private space of the slave core is used to store other data during the execution of computing tasks; for example, intermediate variables, result data, scalar parameters, etc. generated during the computing.
[0032] A part of the continuous shared space of the slave core LDM can be configured to be continuously shared so as to realize distributed storage and shared access of data in various modes within the slave core array. It is used to store the data required for slave core computing or the data obtained from slave core computing. The slave core data Cache space is a smaller and faster storage device, serving as a buffer for a larger and slower storage device. When the slave core reads or stores data from / to the main memory in a discrete memory access manner, the slave core data Cache space can accelerate the discrete access speed of reading or storing main memory data.
[0033] The five memory access schemes include the following: Scheme 1: As Figure 5 shown, all 256KB of the space of the slave core LDM is used as the private space of the slave core, where 128KB of the space is used to store the data required for computing tasks, and the other 128KB of the space is used to save other data during the execution of computing tasks; the slave core transfers the data actually required for computing tasks to the slave core LDM through DMA; Scheme 2: As Figure 6 shown, the single-cluster mode in the continuous shared mode of the slave core LDM is enabled, and the type-a LDM space is used. This scheme can accommodate 512KB (4 * 128KB) of data. The slave core transfers the data actually required for computing tasks to the LDM continuous shared space through slave core DMA; Scheme 3: As Figure 7 shown, the double-cluster mode in the LDM continuous shared mode is enabled, and the type-a LDM space is used. This scheme can accommodate 1024KB (8 * 128KB) of data. The slave core transfers the data actually required for computing tasks to the LDM continuous shared space through slave core DMA; Scheme 4: As Figure 8As shown, enable the four-cluster mode in the LDM continuous sharing mode and use the class-a LDM space. This solution can accommodate 2048 KB (16 * 128 KB) of data. The slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave-core DMA; Solution 5: As Figure 9 shown, enable the full-array cluster mode in the LDM continuous sharing mode and use the class-a LDM space. This solution can accommodate 8192 KB (64 * 128 KB) of data. The slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave-core DMA; The general memory access method means that the slave core uses the class-b LDM space through direct discrete memory access (i.e., gld / gst). This method allows the slave core to access the main memory data without storing it in the slave-core LDM. When the data is relatively discrete, the access speed is slow.
[0034] After determining the above memory access solutions, three problems are faced. First, how to select one of the five memory access solutions and the general memory access method. Second, if any of the five memory access solutions is selected, how to extract the actually required data from the original array and store it in the new array. Third, after transferring the new array to the slave-core LDM, how does the slave core determine the subscript of the array data to be called in the new array. Therefore, the present invention proposes an innovative solution. First, perform data preprocessing, construct a new array, and establish a mapping relationship between the new array subscript and the computing task to achieve the purpose that the slave core can accurately call the new array data. Then, select the memory access solution (method) according to the data volume, thereby solving the slave-core memory access problem and improving the program execution efficiency.
[0035] The present invention will introduce this part of the content in the following two steps.
[0036] As Figure 10 shown, perform data preprocessing in the master core; including: Calculate the subscripts of the original array A; the original array A refers to the array originally existing in the main memory in the program. The original array A stores the data required for the computing task. After dividing the task among the slave cores, the slave cores need to obtain data from the original array A to execute the computing task; Use a boolean array to remove duplicates: Determine whether the subscript of the original array A has appeared by using a boolean array. If it has appeared, skip it; if it has not appeared, record the subscript of the original array A; Construct a new array B to store the data actually required for the computing task; Establish a mapping relationship between the subscripts of the new array B and the computing tasks, and construct another array C to store the mapping relationship. Previously, the subscripts of the original array A changed along with the computing tasks. Now that the new array B is constructed to store the data required for the computing tasks, a mapping relationship between the new array B and the computing tasks needs to be established. For example, if B[2] = A[5], when the computing task is 5, originally array A[5] was needed, but now instead of using array A, only array B[2] is needed. In this case, a relationship between the computing task 5 and the subscript 2 of array B needs to be established so that when the computing task is 5, B[2] is called. Therefore, a new array C is created to save this layer of mapping relationship, C[5] = 2. Subsequently, the array C is transmitted to the slave core so that when the slave core executes the computing task, it can accurately call array B.
[0037] Summarize the actual data volume N required for the computing tasks, and the data volume unit is KB. As Figure 10 shown, the data in the original array A is relatively discrete, while the data in the new array B is continuous. The slave core can make full use of spatial locality to accelerate the memory access speed, thereby improving the program efficiency.
[0038] Construct a new array B to store the data actually required for the computing tasks, including: creating a new array B in the main memory, and after removing duplicates using a boolean array, assign the data in the original array A to the new array B. For example, if the subscript a_index of the original array A has not appeared before, then assign the value of the original array A to array B, B[b_index] = A[a_index], where b_index is the subscript of array B.
[0039] Select a memory access scheme according to the summarized data volume N in the data preprocessing, including: The main problem faced is how to select the most suitable memory access scheme according to the data volume. For example, when the data volume is 0KB < N ≤ 128KB, memory access scheme one can be selected, or scheme two, scheme three, scheme four, etc. can also be selected.
[0040] There are three aspects involved: The first aspect is the essence of the spatial configuration in the LDM continuous sharing mode. The LDM continuous sharing space is a space used by slave cores for LDM shared access within the array, and a slave core can access the LDM spaces of other slave cores. In the LDM continuous sharing mode, each slave core within the array provides the same capacity of LDM for consecutive addressing. The second aspect is the characteristics of the hardware architecture of the ShenWei many-core processor. The distance between slave core clusters affects the communication latency between slave cores. The farther the distance between slave core clusters, the higher the communication latency between slave cores; the closer the distance between slave core clusters, the lower the communication latency between slave cores. The third aspect is that different LDM continuous sharing modes are selected for different memory access schemes. The number of slave cores participating in communication and the positions of slave cores are different in each mode. Single-cluster mode: A slave core array of 2×2 size; Dual-cluster mode: An array of 4×2 size composed of two slave core clusters; Quad-cluster mode: An array of 8×2 size composed of four slave core clusters; Full-array mode: Full-array sharing, with consecutive addressing by slave core clusters (rows).
[0041] To sum up, the LDM continuous sharing space requires communication between slave cores, and the distance between slave core clusters affects the communication latency between slave cores. In different memory access schemes, different LDM continuous sharing modes are selected, and the number of slave cores participating in communication and the hardware positions are different in each mode. For example, when the data volume is 0KB < N ≤ 128KB, memory access scheme one is selected. Compared with schemes two, three, four, etc., scheme one has fewer slave cores participating in communication and closer slave core positions, so the communication latency is lower and the efficiency is faster. As shown in Table 1, taking the Orbit program of nuclear physics nuclear fusion plasma particles' orbits as an example, under the condition of inputting the same physical parameters, the time comparison of the program selecting different schemes for five data volumes is presented, and the time unit is milliseconds (ms).
[0042] Table 1 Time comparison for five data volumes; Data volume Solution 1 Solution 2 Solution 3 Solution 4 Solution 5 0KB < N ≤ 128KB 31.092 32.664 33.543 39.385 66.517 128KB < N ≤ 512KB 39.686 40.165 45.991 72.699 512KB < N ≤ 1024KB 60.363 62.013 74.737 1024KB < N ≤ 2048KB 86.643 100.228 2048KB < N ≤ 8192KB 165.746 It can be seen from Table 1 that the running times of different schemes corresponding to the same data volume are different. Thus, it is determined to select a matching memory access scheme according to the data volume. When 0KB < N ≤ 128KB, select scheme one; when 128KB < N ≤ 512KB, select scheme two; when 512KB < N ≤ 1024KB, select scheme three; when 1024KB < N ≤ 2048KB, select scheme four; when 2048KB < N ≤ 8192KB, select scheme five; when 8192KB < N, select a general memory access method; After completing the above steps, the master core proceeds according to the selected memory access scheme, transfers the new array B and the mapping array C to the LDM of the slave core through DMA, and executes subsequent calculation tasks. The present invention solves the situation where the slave core needs to frequently and discretely access the storage resources of the master core due to the limited LDM space of the slave core, successfully reduces the communication overhead between the master core and the slave core, and efficiently utilizes spatial locality to improve the utilization rate of storage resources, providing support for subsequent optimizations such as SIMD vectorization of the program.
[0043] An optimization method for memory access of the slave core based on the new generation of Sunway many-core processors is realized through the automated interface STMA (Sunway Thread Memory Access).
[0044] The present invention also designs and implements the automated interface STMA, which aims to solve the memory access problem of the slave core. The function of this interface is to automate the design of the above memory access scheme, data preprocessing, and the selection of the memory access scheme based on the data volume. The implementation of this interface reduces the difficulty of programming based on the Sunway supercomputer. Programmers can directly call this interface in the master core to solve problems. The interface description is shown in Table 2: Table 2 STMA interface description; STMA interface Function void STMA_Init(int loop) Initialize STMA. void STMA_Host(double *old_p, int old_dim, double *new_p, double *map_p, int total_n) Programmers call this interface on the main core to automatically calibrate the memory access scheme parameters, perform data preprocessing, and detect the data volume to select the memory access scheme. void STMA_Finalize() End STMA and release resources. When programmers call the STMA interface, they can set the interface parameters according to the specific program. An example of using the STMA interface is as follows: / / Call the STMA interface in the master core master.c #include “STMA.h” int mian() { / / Define the required variables / / Initialize the interface STMA_Init(int loop); / / Data preprocessing STMA_Host(double *old_p, int old_ dim, double *new_p, double *map_p,int total_n); / / Calculation starts / / The hot part of the program / / Calculation ends / / Serial code segment STMA_Finalize(); return 0; } Through the implementation of the method in this embodiment, the present invention tests the inventive method within a single core group.
[0045] This test takes the Orbit program of fusion plasma particles in nuclear physics as an example. Under the condition of inputting the same physical parameters, the program is tested multiple times on five data volumes and the average value is taken. The five data volumes include: 1) 0KB < N ≤ 128KB 2) 128KB < N ≤ 512KB 3) 512KB < N ≤ 1024KB 4) 1024KB < N ≤ 2048KB 5) 2048KB < N ≤ 8192KB The time and speedup ratio of the serial program, the ordinary many-core optimized program, and the inventive optimized program are compared. The specific results and comparisons are shown in Tables 3 to 7.
[0046] Table 3 Running time and speedup ratio of the test program before and after optimization when the data volume is 0KB < N ≤ 128KB; Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 226.384 191.425 1.18 Optimization by the inventive method 226.384 31.092 7.28 Table 4 Running time and speedup ratio of the test program before and after optimization when the data volume is 128KB < N ≤ 512KB; Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 231.451 196.223 1.18 Optimization by the inventive method 231.451 39.686 Table 5 Running time and speedup ratio of the test program before and after optimization when the data volume is 512KB < N ≤ 1024KB; Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 235.502 196.203 1.20 Optimization by the inventive method 235.502 60.363 Table 6 Running time and speedup ratio of the test program before and after optimization when the data volume is 1024KB < N ≤ 2048KB; Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 236.385 201.584 1.17 Optimization by the inventive method 236.385 86.643 Table 7 Running time and speedup ratio of the test program before and after optimization when the data volume is 2048KB < N ≤ 8192KB; Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 240.362 202.598 1.19 Optimization by the inventive method 240.362 165.746 1.45 The experimental results prove that compared with the serial program and the ordinary many-core optimized program, the method of the present invention has an obvious acceleration effect.
[0047] Example 3 A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for optimizing the slave-core memory access based on the new generation Shenwei many-core processor described in Example 1 or 2 are implemented.
[0048] Example 4 A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for optimizing the slave-core memory access based on the new generation Shenwei many-core processor described in Example 1 or 2 are implemented.
Claims
1. Based on the new generation of Shenwei multi-core processor, the slave core memory access optimization method is characterized by: include: Firstly, five memory access schemes and a general memory access method are designed according to the LDM continuous sharing mode; Secondly, data preprocessing is to calculate the array subscripts of the array data actually required in the computing task in advance in the main core, remove duplicate subscripts and construct a new array; Finally, determine whether the actual data volume exceeds the maximum space capacity of 8192KB in the full array LDM continuous sharing mode; If it exceeds, a general memory access method is adopted, that is, direct discrete memory access and setting a cache space capacity of 128KB; if it does not exceed, a different memory access scheme is selected according to the data volume.
2. The method for optimizing memory access from a core based on a new generation of Shenwei multi-core processor according to claim 1, characterized in that: Five memory access schemes and a general memory access method are designed based on the LDM continuous sharing mode; including: The space of slave core LDM is designed into two categories: class a LDM space and class b LDM space; class a LDM space uses 128KB of slave core LDM space as slave core private space, and the remaining 128KB as slave core LDM continuous shared space; class b LDM space uses 128KB of slave core LDM space as slave core private space, and the remaining 128KB as slave core data cache space; slave core private space is used to store other data during the execution of computing tasks; The five memory access schemes are as follows: Solution 1: All 256KB of the slave core LDM space is used as the slave core private space, of which 128KB is used to store the data required for the computing task, and the other 128KB is used to store other data during the computing task execution; the slave core transfers the data actually required for the computing task to the slave core LDM through DMA; Solution 2: Enable the single cluster mode in the slave core LDM continuous sharing mode, use the class a LDM space, and the slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA; Solution 3: Enable the dual-cluster mode in the LDM continuous sharing mode, use the class a LDM space, and transfer the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA. Solution 4: Enable the four-cluster mode in the LDM continuous sharing mode, use the class a LDM space, and transfer the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA. Solution 5: Enable the full array cluster mode in the LDM continuous sharing mode, use the class a LDM space, and use the slave core DMA to transfer the data actually required for the computing task to the LDM continuous sharing space; The general memory access method refers to: directly accessing the memory from the core through discrete memory, using the class b LDM space.
3. The method for optimizing memory access from a core based on a new generation of Shenwei multi-core processor according to claim 1, characterized in that: Data preprocessing is performed in the main core; including: Calculate the subscript of the original array A; the original array A refers to the array that originally exists in the main memory in the program. The original array A stores the data required for the computing task. After the task is divided among the slave cores, the slave cores need to obtain data from the original array A to perform the computing task. Use Boolean array to remove duplicates: Use Boolean array to determine whether the subscript of the original array A has appeared. If it has appeared, skip it; if it has not appeared, record the subscript of the original array A; Construct a new array B to store the data actually required for the computing task; Establish a mapping relationship between the subscripts of the new array B and the computing tasks, and then build an array C to store the mapping relationship; Summarizes the actual amount of data N required for the computing task, in KB.
4. The method for optimizing memory access from a core based on a new generation of Shenwei multi-core processor according to claim 3 is characterized in that: Construct a new array B to store the data actually required for the computing task; including: creating a new array B in the main memory, using a Boolean array to deduplicate, and assigning the data in the original array A to the new array B.
5. The method for optimizing memory access from a core based on a new generation of Shenwei multi-core processor according to any one of claims 1 to 4, characterized in that: Select a memory access solution based on the amount of data N collected in data preprocessing; include: When 0KB<N≤128KB, select option 1; when 128KB<N≤512KB, select option 2; when 512KB<N≤1024KB, select option 3; when 1024KB<N≤2048KB, select option 4; when 2048KB<N≤8192KB, select option 5; when 8192KB<N, select the general memory access method; After completing the above steps, the master core follows the selected memory access scheme, transfers the new array B and the mapped array C to the slave core LDM via DMA, and executes subsequent computing tasks.
6. The method for optimizing memory access from a core based on a new generation of Shenwei multi-core processor according to any one of claims 1 to 4, characterized in that: The automatic interface STMA is used to implement the slave core memory access optimization method based on the new generation Shenwei multi-core processor.
Citation Information
Patent Citations
AztecOO transplantation optimization method and system based on Shenwei super computer
CN118656126A
Data processing method and related equipment
CN118734061A
Communication network operation and maintenance fault positioning and tracking method and system
CN119420639A
De-Duplication Optimized Platform for Object Grouping
US20170344598A1
Method for resource adjustment in cache, data access method and device
WO2019127104A1
Cited By
Lightweight MD simulation method and system for heterogeneous many-core processor
CN120373058A
Discrete data stream type memory access optimization method based on SW many-core processor SW39000
CN122220298A