Method for Optimizing Memory Access of Slave Cores Based on New Generation ShenWei Many-Core Processors

By designing a memory access solution for LDM continuous sharing mode and data preprocessing in Shenweizhong core processor, the memory access from core is optimized, and the problem of limited space of LDM is solved, and the program execution efficiency and storage resource utilization are improved.

CN120144503BActive Publication Date: 2025-07-29QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510626634.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-29
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the Shenweizhong core processor, the limited slave core LDM space leads to frequent discrete access to the main core storage resources, large communication overhead, and insufficient data transmission, resulting in low program memory access efficiency.

Method used

Five storage access solutions and a general storage access method are designed using the LDM continuous sharing mode. Combined with data preprocessing, the storage access from the core is optimized through DMA or direct discrete storage access, and the space capacity of 128KB Cache is set, and the appropriate storage access solution is selected to improve efficiency.

Benefits of technology

It effectively reduces the overhead of master-slave core communication, makes full use of space locality, improves program execution efficiency, reduces storage resource waste, and optimizes slave core acceleration solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144503B_ABST
    Figure CN120144503B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for optimizing the memory access of slave cores based on the new generation of Shenwei many-core processors, belonging to the technical field of electronic information. It includes: First, five memory access schemes and a general memory access method are designed according to the LDM continuous sharing mode. Second, data preprocessing is carried out. For the array data actually required in the computing task, the array subscripts are calculated in advance in the master core, duplicate subscripts are removed, and a new array is constructed. Finally, it is judged whether the actual data volume exceeds the maximum space capacity of 8192KB of the full-array LDM continuous sharing mode. If it exceeds, a general memory access method is adopted, that is, direct discrete memory access is performed and the Cache space capacity of 128KB is set. If it does not exceed, different memory access schemes are selected according to the data volume. Selecting the method of the present invention for optimization can improve the program execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an optimization method for the memory access of slave cores based on a new generation of ShenWei many-core processors, belonging to the technical field of electronic information. Background Art

[0002] The computing power of the new generation of ShenWei supercomputer is provided by the domestic many-core SW26010pro processor. The architecture of the SW26010pro processor is as Figure 1 shown. This processor includes 6 core groups (Core-Group, CG), and each core group includes 1 management core (Management Processing Element, MPE, also known as the master core) and 1 computation processing elements cluster (Computation Processing Elements Clusters, CPE cluster, also known as the slave core array). Each slave core array is composed of 64 slave cores (Computing Processing Element, CPE) in an 8×8 mesh format. In the slave core array, each slave core cluster is composed of four slave cores, and the four adjacent slave cores share a slave core cluster management component and integrate a router. Each core group has its own memory controller (Memory Controller, MC), which is connected to 16GB of DDR4 memory with a bandwidth of 51.2GB / s.

[0003] The storage within a core group of the ShenWei many-core chip is divided into the main memory located in the master core and the LDM located in each private slave core. Each slave core has a local data memory space (Local Data Memory, LDM) with a capacity of 256 KB. The ways for slave cores to access the main memory and the division of the LDM space are as Figure 2 shown. Facing the memory access requirements, the slave cores can access the main memory in a direct discrete manner (i.e., gld / gst); they can also access the main memory in batches through DMA, first moving the data from the main memory to the locally accessible LDM of the slave cores. When the slave cores are computing, they only need to initiate a memory access to the LDM.

[0004] The LDM space of the slave cores can be divided into three parts: private space, continuous shared space, and slave core data cache space, and the three parts of the space can be adjusted in different grades. When part of the space is configured as a level-1 data cache, the slave cores can perform cache access to this space by loading or storing the cacheable space of the main memory, and the cache line size is 256 bytes. The SW26010pro processor supports three data cache capacity configurations of 0KB, 32KB, and 128KB.

[0005] The LDM continuous shared space is a space used by slave cores for LDM shared access within an array. A slave core can access the LDM spaces of other slave cores. In the LDM continuous shared mode, each slave core within the array provides an LDM space of the same capacity for continuous addressing. The capacity of the LDM space provided by each slave core for sharing can be adjusted in 7 capacities, such as 0 KB, 4 KB, 8 KB, 16 KB, 32 KB, 64 KB, and 128 KB.

[0006] The continuous shared mode includes 5 modes, namely single cluster (a slave core cluster is a 2×2 slave core array), double cluster (a 4×2 array composed of two slave core clusters), quadruple cluster (an 8×2 array composed of four slave core clusters), full array cluster (full array sharing, continuous addressing by slave core cluster), and full array row (full array sharing, continuous addressing by row), which facilitates the nearby access to different working sets.

[0007] The method for optimizing a single-core group using the ShenWei many-core processor mainly involves the main core starting the program and migrating the hot code segments that can be optimized in parallel in the application program to the slave cores to achieve the effect of parallel acceleration. During the execution of the hot code on the slave cores, the slave cores must first obtain the data required for the computing task from the main memory through discrete memory access (i.e., gld / gst) or DMA. Only after the slave cores obtain the required data can they continue to execute the subsequent code.

[0008] There are two ways for slave cores to obtain data. First, the slave cores can obtain the data required for the computing task from the main memory through direct discrete memory access (i.e., gld / gst) without saving it to the slave core LDM. When using this part of the data, the slave cores need to maintain communication with the main core at all times to obtain data from the main memory. Second, the slave cores can obtain the data required for the computing task from the main memory through DMA and store it in the slave core LDM. When using this part of the data, the slave cores only need to call it from the LDM.

[0009] Facing the above two ways of obtaining data, generally the second way is adopted. Even though the second way requires additional space to store data, compared with the first way, the second way reduces the communication overhead between the slave cores and the main core, and it is more convenient and fast to call data, making the overall program run more efficiently.

[0010] However, there are three problems with the above method. First, the data required for the computing tasks in the application program is basically array data. When normally optimizing the program for the purpose of convenience and speed, when obtaining array data from the slave core, the data actually required for the computing tasks is not considered. Instead, the entire array data is transferred to the slave core LDM through DMA. This often ignores the rarity of the slave core LDM resources, and blindly transferring extra data will increase the communication overhead, resulting in wasted storage resources and low program efficiency. Second, the amount of array data in the application program is generally very large. Since the LDM space resources of the slave core are very limited, each slave core only has 256KB of LDM storage space. There is a situation where the amount of array data exceeds 256KB and cannot be stored in the private LDM of the slave core. In response to this situation, the LDM continuous sharing mode can be enabled, and the appropriate LDM continuous sharing mode and space capacity can be selected according to different data volumes. Although the LDM continuous sharing mode can be enabled and the space capacity can be set to store the array data, since the data is not preprocessed, the actual required data volume cannot be determined, and blindly transferring the entire array to the LDM will result in a significant reduction in the utilization rate of storage resources. Third, if there is a situation where the data volume exceeds the maximum capacity in the LDM continuous sharing mode. That is, when the full-array LDM continuous sharing mode is enabled and a 128KB continuous sharing space capacity is selected, and the total space capacity reaches 8192KB (64 * 128KB) but still cannot store the array data. At this time, the method of direct discrete memory access (i.e., gld / gst) is generally adopted, and the Cache space capacity is adjusted according to the demand to speed up the direct discrete memory access. However, even if the slave core Cache is used to improve the efficiency of direct discrete memory access, the final efficiency of the program is still not high. Summary of the Invention

[0011] Aiming at the deficiencies of the prior art, the present invention provides an optimization method for slave core memory access based on the new generation ShenWei many-core processor;

[0012] When optimizing a program on a SW64 processor, the computing tasks are assigned to slave cores. The slave cores obtain the data required for the computing tasks through direct discontinuous memory access (i.e., gld / gst) or DMA. When normally optimizing a program, for the sake of convenience and simplicity, usually the entire array of data required for the computing tasks is transferred to the LDM of the slave cores through DMA. However, when the data volume of the entire array exceeds the maximum capacity of the LDM continuous sharing mode, the data cannot be transferred to the LDM through DMA. If a general memory access method is blindly adopted, that is, direct discontinuous memory access and using the Cache space, it will not only cause serious communication overhead problems, but also there may be a situation where the data is too discontinuous to fully utilize the spatial locality, resulting in low memory access efficiency of the program. For a program with the above problems, selecting the method of the present invention for optimization can effectively solve this problem and improve the program execution efficiency. At the same time, the present invention also designs an automated interface STMA (Sunway Thread Memory Access) for the convenience of programmers to directly call and use.

[0013] The present invention specifically solves the situation where the local storage space of the slave cores is limited, resulting in the slave cores needing to frequently and discontinuously access the main core storage resources, and successfully reduces the communication overhead between the main core and the slave cores, and efficiently utilizes the spatial locality, and applies it to the slave core acceleration scheme.

[0014] The technical solution of the present invention is as follows:

[0015] A memory access optimization method for the slave cores of a new generation of SW64 processors, including:

[0016] First, five memory access schemes and a general memory access method are designed according to the LDM continuous sharing mode;

[0017] Secondly, data preprocessing is performed. For the array data actually required in the computing tasks, the array subscripts are calculated in advance in the main core, the duplicate subscripts are removed, and a new array is constructed;

[0018] Finally, it is judged whether the actual data volume exceeds the maximum space capacity of 8192 KB of the full-array LDM continuous sharing mode; if it exceeds, a general memory access method is adopted, that is, direct discontinuous memory access and a Cache space capacity of 128 KB is set; if it does not exceed, different memory access schemes are selected according to the data volume.

[0019] According to the preference of the present invention, five memory access schemes and a general memory access method are designed according to the LDM continuous sharing mode; including:

[0020] The space of the slave core LDM is designed into two categories: type-a LDM space and type-b LDM space; in the type-a LDM space, 128 KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128 KB is used as the continuous shared space of the slave core LDM; in the type-b LDM space, 128 KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128 KB is used as the data cache space of the slave core; the private space of the slave core is used to store other data during the execution of the computing task.

[0021] The five memory access schemes are as follows:

[0022] Scheme 1: All 256 KB of the space of the slave core LDM is used as the private space of the slave core. Among them, 128 KB of the space is used to store the data required for the computing task, and the other 128 KB of the space is used to save other data during the execution of the computing task; the slave core transfers the data actually required for the computing task to the slave core LDM through DMA.

[0023] Scheme 2: The single-cluster mode in the continuous sharing mode of the slave core LDM is enabled, and the type-a LDM space is used. The slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA.

[0024] Scheme 3: The double-cluster mode in the LDM continuous sharing mode is enabled, and the type-a LDM space is used. The slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA.

[0025] Scheme 4: The four-cluster mode in the LDM continuous sharing mode is enabled, and the type-a LDM space is used. The slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA.

[0026] Scheme 5: The full-array cluster mode in the LDM continuous sharing mode is enabled, and the type-a LDM space is used. The slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA.

[0027] The general memory access method means that the slave core uses the type-b LDM space through direct discrete memory access.

[0028] According to the preference of the present invention, data preprocessing is performed in the master core; it includes:

[0029] Calculating the subscripts of the original array A; the original array A refers to the array originally existing in the main memory in the program. The original array A stores the data required for the computing task. After the task is divided among the slave cores, the slave cores need to obtain data from the original array A to execute the computing task.

[0030] Deduplication using a boolean array: By using a boolean array to determine whether the subscript of the original array A has appeared. If it has appeared, skip it; if not, record the subscript of the original array A.

[0031] Construct a new array B to store the data actually required for the computing task.

[0032] Establish a mapping relationship between the subscript of the new array B and the computing task, and construct another array C to store the mapping relationship.

[0033] Summarize the actual data volume N required for the computing task, and the data volume unit is KB.

[0034] Further preferably, construct a new array B to store the data actually required for the computing task, including: creating a new array B in the main memory, and after deduplication using the boolean array, assign the data in the original array A to the new array B.

[0035] According to the present invention preferably, select a memory access scheme according to the summarized data volume N in data preprocessing, including:

[0036] When 0KB < N ≤ 128KB, select Scheme 1; when 128KB < N ≤ 512KB, select Scheme 2; when 512KB < N ≤ 1024KB, select Scheme 3; when 1024KB < N ≤ 2048KB, select Scheme 4; when 2048KB < N ≤ 8192KB, select Scheme 5; when 8192KB < N, select a general memory access method.

[0037] After completing the above steps, the main core performs according to the selected memory access scheme, and transfers the new array B and the mapping array C to the slave core LDM through DMA and executes the subsequent computing tasks.

[0038] According to the present invention preferably, implement the memory access optimization method for the slave core based on the new generation Shenwei many-core processor through the automated interface STMA.

[0039] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it realizes the steps of the memory access optimization method for the slave core based on the new generation Shenwei many-core processor.

[0040] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it realizes the steps of the memory access optimization method for the slave core based on the new generation Shenwei many-core processor.

[0041] The beneficial effects of the present invention are:

[0042] When optimizing a program on a ShenWei many-core processor, the computing tasks are assigned to slave cores, and the slave cores obtain the data required for the computing tasks through direct discrete memory access (i.e., gld / gst) or DMA. When normally optimizing a program, for the sake of convenience and simplicity, usually the entire array of data required for the computing tasks is transferred to the LDM of the slave core through DMA. However, when the data volume of the entire array exceeds the maximum capacity of the LDM continuous sharing mode, the data cannot be transferred to the LDM through DMA. If a general memory access method is blindly adopted, that is, direct discrete memory access and using the Cache space, it will not only cause serious problems of communication overhead, but also there may be a situation where the data is too discrete, and the spatial locality cannot be fully utilized, resulting in low memory access efficiency of the program. For a program with the above problems, selecting the method of the present invention for optimization can effectively solve this problem and improve the program execution efficiency. At the same time, the present invention also designs an automated interface STMA (Sunway Thread Memory Access), which is convenient for programmers to directly call and use. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic diagram of the architecture of SW26010pro;

[0044] Figure 2 It is a schematic diagram of the slave core accessing the main memory method and the LDM space division;

[0045] Figure 3 It is a schematic diagram of the flow of the slave core memory access optimization method based on the new generation of ShenWei many-core processor of the present invention;

[0046] Figure 4 It is a schematic diagram of the space division of class a and class b;

[0047] Figure 5 It is a schematic diagram of the space of Scheme 1;

[0048] Figure 6 It is a schematic diagram of the space of Scheme 2;

[0049] Figure 7 It is a schematic diagram of the space of Scheme 3;

[0050] Figure 8 It is a schematic diagram of the space of Scheme 4;

[0051] Figure 9 It is a schematic diagram of the space of Scheme 5;

[0052] Figure 10 It is a schematic diagram of the data preprocessing flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0054] Glossary Explanation:

[0055] 1. gld / gst Memory Access: Discrete memory access. On the slave core side, data can be directly retrieved from or stored into the main memory through the main memory address, without the need to save it to the local local data storage space of the slave core.

[0056] 2. Cache: A smaller and faster storage device that serves as a buffer for a larger and slower storage device, thereby improving data access speed. The slave core can use the Cache to accelerate the discrete access speed of reading or storing main memory data.

[0057] 3. LDM: Local Data Memory. Each slave core of the ShenWei many-core processor has a high-speed local local data storage space LDM, and the total capacity of the LDM is 256KB.

[0058] 4. DMA: Direct Memory Access. Initiated by a single slave core, it is an operation for batch data transfer between the main memory and the local local memory LDM of the slave core.

[0059] Embodiment 1

[0060] Based on the memory access optimization method for the slave core of the new generation ShenWei many-core processor, it includes:

[0061] Firstly, design five memory access schemes and a general memory access method according to the LDM continuous sharing mode (single cluster, double cluster, four clusters, full array);

[0062] Secondly, data preprocessing. For the array data actually required in the computing task, calculate the array subscripts in advance in the main core, remove duplicate subscripts and construct a new array;

[0063] Finally, judge whether the actual data volume exceeds the maximum space capacity of 8192KB (64 * 128KB) of the full array LDM continuous sharing mode; if it exceeds, adopt the general memory access method, that is, direct discrete memory access (i.e., gld / gst) and set the Cache space capacity of 128KB; if it does not exceed, select different memory access schemes according to the data volume.

[0064] This innovative method solves the problem of low efficiency of direct discrete memory access of the slave core, fully utilizes spatial locality, and reduces the communication overhead between the main and slave cores, thereby greatly improving the execution efficiency of the program. The specific process is as Figure 3 shown.

[0065] Embodiment 2

[0066] The method for optimizing the memory access of slave cores based on the new generation of ShenWei many-core processors according to Embodiment 1 is characterized in that:

[0067] Design five memory access schemes and a general memory access method according to the LDM continuous sharing mode; including:

[0068] When designing the memory access scheme, based on three aspects of factors, namely the slave core LDM space division, the LDM continuous sharing mode, and the space capacity, and combined with the data volume, design five memory access schemes and a general memory access method.

[0069] According to the characteristics of the hardware architecture of the new generation of ShenWei supercomputers and ShenWei many-core processors, the slave core LDM space is divided into three parts: private space, continuous sharing space, and slave core data Cache space, and the three parts of the space can be adjusted in different grades. To describe the memory access scheme more vividly and intuitively, as Figure 4 shown, the space of the slave core LDM in the present invention is designed into two types: type a LDM space and type b LDM space; in the type a LDM space, 128KB of the space of the slave core LDM is used as the slave core private space, and the remaining 128KB is used as the slave core LDM continuous sharing space; in the type b LDM space, 128KB of the space of the slave core LDM is used as the slave core private space, and the remaining 128KB is used as the slave core data Cache space; the slave core private space is used to store other data during the execution of the computing task; for example, intermediate variables, result data, and scalar parameters generated during the calculation, etc.

[0070] A part of the slave core LDM continuous sharing space can be configured as continuous sharing to achieve distributed storage and shared access of data in various modes within the slave core array. It is used to store the data required for slave core computing or the data obtained from slave core computing. The slave core data Cache space is a smaller and faster storage device, serving as a buffer for a larger and slower storage device. When the slave core reads or stores data from the main memory in a discrete memory access manner, the slave core data Cache space can accelerate the discrete access speed of reading or storing main memory data.

[0071] The five memory access schemes are as follows:

[0072] Scheme 1: As Figure 5 shown, regard the 256KB space of the slave core LDM as the slave core private space, where 128KB of the space is used to store the data required for the computing task, and the other 128KB of the space is used to save other data during the execution of the computing task; the slave core transfers the data actually required for the computing task to the slave core LDM through DMA;

[0073] Scheme 2: As Figure 6As shown, enable the single-cluster mode in the continuous shared mode of the slave core LDM, and use the class-a LDM space. This solution can accommodate 512 KB (4 * 128 KB) of data. The slave core transfers the data actually required for the computing task to the LDM continuous shared space through the slave core DMA;

[0074] Solution 3: As Figure 7 shown, enable the dual-cluster mode in the continuous shared mode of the LDM, and use the class-a LDM space. This solution can accommodate 1024 KB (8 * 128 KB) of data. The slave core transfers the data actually required for the computing task to the LDM continuous shared space through the slave core DMA;

[0075] Solution 4: As Figure 8 shown, enable the four-cluster mode in the continuous shared mode of the LDM, and use the class-a LDM space. This solution can accommodate 2048 KB (16 * 128 KB) of data. The slave core transfers the data actually required for the computing task to the LDM continuous shared space through the slave core DMA;

[0076] Solution 5: As Figure 9 shown, enable the full-array cluster mode in the continuous shared mode of the LDM, and use the class-a LDM space. This solution can accommodate 8192 KB (64 * 128 KB) of data. The slave core transfers the data actually required for the computing task to the LDM continuous shared space through the slave core DMA;

[0077] The general memory access method means that the slave core accesses the main memory data through direct discrete memory access (i.e., gld / gst) and uses the class-b LDM space. This method allows the slave core to access the main memory data without storing it in the slave core LDM. When the data is relatively discrete, the access speed is slow.

[0078] After determining the above memory access solutions, three problems are faced. First, how to select one of the five memory access solutions and the general memory access method. Second, if any one of the five memory access solutions is selected, how to extract the actually required data from the original array and store it in the new array. Third, after transferring the new array to the slave core LDM, how does the slave core determine the subscript of the array data to be called in the new array. Therefore, the present invention proposes an innovative solution. First, perform data preprocessing, construct a new array, and establish a mapping relationship between the new array subscript and the computing task to achieve the purpose that the slave core can accurately call the new array data. Then, select the memory access solution (method) according to the data volume, thereby solving the slave core memory access problem and improving the program execution efficiency.

[0079] The present invention will introduce this part of the content in the following two steps.

[0080] As Figure 10As shown, data preprocessing is performed in the main core, including:

[0081] Calculate the subscripts of the original array A. The original array A refers to the array that originally exists in the main memory in the program. The original array A stores the data required for the calculation task. After dividing the task among the slave cores, the slave cores need to obtain data from the original array A to execute the calculation task.

[0082] Use a boolean array to remove duplicates: Determine whether the subscript of the original array A has appeared by using a boolean array. If it has appeared, skip it; if it has not appeared, record the subscript of the original array A.

[0083] Construct a new array B to store the data actually required for the calculation task.

[0084] Establish a mapping relationship between the subscripts of the new array B and the calculation tasks, and construct another array C to store the mapping relationship. Before, the subscripts of the original array A changed with the calculation tasks. Now that the new array B is constructed to store the data required for the calculation tasks, a mapping relationship between the new array B and the calculation tasks needs to be established. For example: Suppose B[2]=A[5]. When the calculation task is 5, originally array A[5] was needed, but now instead of using array A, only array B[2] is needed. In this way, a relationship between the calculation task 5 and the subscript 2 of array B needs to be established so that when the calculation task is 5, B[2] is called. Therefore, a new array C is created to save this mapping relationship, C[5]=2. Subsequently, the array C is transmitted to the slave cores so that when the slave cores execute the calculation tasks, they can accurately call array B.

[0085] Summarize the actual data volume N required for the calculation tasks, and the data volume unit is KB. As Figure 10 shown, the data in the original array A is relatively discrete, while the data in the new array B is continuous. The slave cores can make full use of spatial locality to accelerate the memory access speed, thereby improving the program efficiency.

[0086] Construct a new array B to store the data actually required for the calculation task, including: Create a new array B in the main memory. After using the boolean array to remove duplicates, assign the data in the original array A to the new array B. For example: If the subscript a_index of the original array A has not appeared before, then assign the value of the original array A to array B, B[b_index] = A[a_index], where b_index is the subscript of array B.

[0087] Select a memory access scheme according to the total data volume N summarized in the data preprocessing, including:

[0088] The main problem faced is how to select the most suitable memory access scheme according to the data volume. For example, when the data volume is 0KB < N ≤ 128KB, memory access scheme one can be selected, or scheme two, scheme three, scheme four, etc.

[0089] There are three aspects involved: The first aspect is the essence of the spatial configuration in the LDM continuous sharing mode. The LDM continuous sharing space is a space used by slave cores for LDM shared access within the array, and a slave core can access the LDM spaces of other slave cores. In the LDM continuous sharing mode, each slave core within the array provides the same capacity of LDM for consecutive addressing. The second aspect is the characteristics of the hardware architecture of the ShenWei many-core processor. The distance between slave core clusters affects the communication latency between slave cores. The farther the distance between slave core clusters, the higher the communication latency between slave cores; the closer the distance between slave core clusters, the lower the communication latency between slave cores. The third aspect is that different memory access schemes select different LDM continuous sharing modes. The number of slave cores participating in communication and the positions of slave cores are different in each mode. Single-cluster mode: A slave core array with a size of 2*2; Dual-cluster mode: An array with a size of 4*2 composed of two slave core clusters; Four-cluster mode: An array with a size of 8*2 composed of four slave core clusters; Full-array mode: Full-array sharing, with consecutive addressing by slave core clusters (rows).

[0090] In summary, the LDM continuous sharing space requires communication between slave cores, and the distance between slave core clusters affects the communication latency between slave cores. In different memory access schemes, different LDM continuous sharing modes are selected, and the number of slave cores participating in communication and the hardware positions are different in each mode. For example, when the data volume is 0KB < N ≤ 128KB, memory access scheme one is selected. Compared with schemes two, three, four, etc., scheme one has fewer slave cores participating in communication and closer slave core positions, so the communication latency is lower and the efficiency is faster. As shown in Table 1, taking the Orbit program of nuclear physics nuclear fusion plasma particles as an example, under the condition of inputting the same physical parameters, the time comparison of the program selecting different schemes for five data volumes, and the time unit is milliseconds (ms).

[0091] Table 1 Time comparison for five data volumes;

[0092] Data volume Solution 1 Solution 2 Solution 3 Solution 4 Solution 5 0KB < N ≤ 128KB 31.092 32.664 33.543 39.385 66.517 128KB < N ≤ 512KB 39.686 40.165 45.991 72.699 512KB < N ≤ 1024KB 60.363 62.013 74.737 1024KB < N ≤ 2048KB 86.643 100.228 2048KB < N ≤ 8192KB 165.746

[0093] It can be seen from Table 1 that the running times of different schemes corresponding to the same data volume are different. Therefore, it is determined to select a matching memory access scheme according to the data volume. When 0KB < N ≤ 128KB, select scheme one; when 128KB < N ≤ 512KB, select scheme two; when 512KB < N ≤ 1024KB, select scheme three; when 1024KB < N ≤ 2048KB, select scheme four; when 2048KB < N ≤ 8192KB, select scheme five; when 8192KB < N, select a general memory access method;

[0094] After completing the above steps, the master core proceeds according to the selected memory access scheme, transfers the new array B and the mapping array C to the LDM of the slave core through DMA, and executes subsequent computing tasks. The present invention solves the situation where the slave core needs to frequently and discretely access the storage resources of the master core due to the limited LDM space of the slave core, successfully reduces the communication overhead between the master core and the slave core, and efficiently utilizes spatial locality to improve the utilization rate of storage resources, providing support for subsequent optimizations such as SIMD vectorization of the program.

[0095] An optimization method for memory access of the slave core based on the new generation of Sunway many-core processors is realized through the automated interface STMA (Sunway Thread Memory Access).

[0096] The present invention also designs and implements the automated interface STMA, which is designed to solve the memory access problem of the slave core. The function of this interface is to automate the design of the above memory access scheme, data preprocessing, and selection of the memory access scheme based on the data volume. The implementation of this interface reduces the difficulty of programming based on the Sunway supercomputer. Programmers can directly call this interface in the master core to solve problems. The interface description is shown in Table 2:

[0097] Table 2 STMA Interface Description;

[0098] STMA interface Function void STMA_Init(int loop) Initialize STMA. void STMA_Host(double *old_p, int old_dim, double *new_p, double *map_p, int total_n) The programmer calls this interface on the main core to automatically calibrate the memory access scheme parameters, perform data preprocessing, and detect the data volume to select the memory access scheme. void STMA_Finalize() End STMA and release resources.

[0099] When programmers call the STMA interface, they can set the interface parameters according to the specific program. An example of using the STMA interface is as follows:

[0100] / / Call the STMA interface in the master core master.c

[0101] #include “STMA.h”

[0102] int mian()

[0103] {

[0104] / / Define the required variables

[0105] / / Initialize the interface

[0106] STMA_Init(int loop);

[0107] / / Data preprocessing

[0108] STMA_Host(double *old_p, int old_ dim, double *new_p, double *map_p,int total_n);

[0109] / / Calculation starts

[0110] / / Hot part of the program

[0111] / / Calculation completed

[0112] / / Serial code segment

[0113] STMA_Finalize();

[0114] return 0;

[0115] }

[0116] Through the implementation of the method in this embodiment, the present invention tests the inventive method within a single core group.

[0117] This test takes the Orbit program of nuclear fusion plasma particles in nuclear physics as an example. Under the condition of the same input physical parameters, the program is tested multiple times on five data volumes and the average value is taken. The five data volumes include:

[0118] 1) 0KB < N ≤ 128KB

[0119] 2) 128KB < N ≤ 512KB

[0120] 3) 512KB < N ≤ 1024KB

[0121] 4) 1024KB < N ≤ 2048KB

[0122] 5) 2048KB < N ≤ 8192KB

[0123] Compare the time and speedup ratio of the serial program, the general many-core optimized program, and the inventive optimized program. The specific results and comparisons are shown in Tables 3 to 7.

[0124] Table 3 Running time and speedup ratio of the test program before and after optimization when the data volume is 0KB < N ≤ 128KB;

[0125] Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 226.384 191.425 1.18 Optimization by the inventive method 226.384 31.092 7.28

[0126] Table 4 Running time and speedup ratio of the test program before and after optimization when the data volume is 128KB < N ≤ 512KB;

[0127] Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 231.451 196.223 1.18 Optimization by the inventive method 231.451 39.686

[0128] Table 5 Running time and speedup ratio of the test program before and after optimization when the data volume is 512KB < N ≤ 1024KB;

[0129] Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 235.502 196.203 1.20 Optimization by the inventive method 235.502 60.363

[0130] Table 6 Running time and speedup ratio of the test program before and after optimization when the data volume is 1024KB < N ≤ 2048KB;

[0131] Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 236.385 201.584 1.17 Optimization by the inventive method 236.385 86.643

[0132] Table 7 Running time and speedup ratio of the test program before and after optimization when the data volume is 2048KB < N ≤ 8192KB;

[0133] Optimization method Before optimization (ms) After optimization (ms) Speedup ratio General many-core optimization 240.362 202.598 1.19 Optimization by the inventive method 240.362 165.746 1.45

[0134] The experimental results prove that compared with the serial program and the ordinary many-core optimized program, the method of the present invention has an obvious acceleration effect.

[0135] Embodiment 3

[0136] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the method for optimizing the slave-core memory access based on the new generation Shenwei many-core processor described in Embodiment 1 or 2 are implemented.

[0137] Embodiment 4

[0138] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method for optimizing the slave-core memory access based on the new generation Shenwei many-core processor described in Embodiment 1 or 2 are implemented.

Claims

1. A method for optimizing the memory access of slave cores based on the new generation of ShenWei many-core processors, characterized in that Including: Firstly, design five memory access schemes and a general memory access method according to the LDM continuous sharing mode; Secondly, perform data preprocessing. For the array data actually required in the computing task, calculate the array subscripts in the main core in advance, remove duplicate subscripts, and construct a new array; Finally, determine whether the actual data volume exceeds the maximum space capacity of 8192 KB of the full-array LDM continuous sharing mode; If it exceeds, adopt a general memory access method, that is, directly perform discrete memory access and set the Cache space capacity to 128 KB; If it does not exceed, select different memory access schemes according to the data volume; Perform data preprocessing in the main core; including: Calculate the subscripts of the original array A; the original array A refers to the array originally existing in the main memory in the program. The original array A stores the data required for the computing task. After dividing the task among the slave cores, the slave cores need to obtain data from the original array A to execute the computing task; Use a boolean array to remove duplicates: judge whether the subscripts of the original array A have appeared by using a boolean array. If they have appeared, skip them; if they have not appeared, record the subscripts of the original array A; Construct a new array B to store the data actually required for the computing task; Establish a mapping relationship between the subscripts of the new array B and the computing task, and construct another array C to store the mapping relationship; Summarize the actual data volume N required for the computing task, and the data volume unit is KB.

2. The method for optimizing the memory access of slave cores based on the new generation ShenWei many-core processor according to claim 1, characterized in that Design five memory access schemes and a general memory access method according to the LDM continuous sharing mode; including: Design the space of the slave core LDM into two types: type a LDM space and type b LDM space; for type a LDM space, 128 KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128 KB is used as the continuous sharing space of the slave core LDM; for type b LDM space, 128 KB of the space of the slave core LDM is used as the private space of the slave core, and the remaining 128 KB is used as the data Cache space of the slave core; the private space of the slave core is used to store other data during the execution of the computing task; The five memory access schemes are as follows: Scheme 1: Use 256 KB of the space of the slave core LDM as the private space of the slave core. Among them, 128 KB of the space is used to store the data required for the computing task, and the other 128 KB of the space is used to store other data during the execution of the computing task; the slave core transfers the data actually required for the computing task to the slave core LDM through DMA; Scheme 2: Enable the single-cluster mode in the LDM continuous sharing mode of the slave core, use type a LDM space, and the slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA; Scheme 3: Enable the double-cluster mode in the LDM continuous sharing mode, use type a LDM space, and the slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA; Scheme 4: Enable the four-cluster mode in the LDM continuous sharing mode, use type a LDM space, and the slave core transfers the data actually required for the computing task to the LDM continuous sharing space through the slave core DMA; Solution 5: Enable the full-array cluster mode in the LDM continuous sharing mode, use the class-a LDM space, and the slave core transfers the data actually required for the computing task to the LDM continuous sharing space through slave-core DMA. The general memory access method refers to: the slave core uses the class-b LDM space through direct discrete memory access.

3. The method for optimizing the memory access of slave cores based on the new generation of ShenWei many-core processors according to claim 1, wherein Construct a new array B to store the data actually required for the computing task; including: create a new array B in the main memory, and after de-duplication using a boolean array, assign the data in the original array A to the new array B.

4. The method for optimizing memory access of slave cores based on the new generation ShenWei many-core processor according to any one of claims 1-3, characterized in that, Select a memory access solution according to the total data volume N in data preprocessing; Including: When 0KB < N ≤ 128KB, select Solution 1; when 128KB < N ≤ 512KB, select Solution 2; when 512KB < N ≤ 1024KB, select Solution 3; when 1024KB < N ≤ 2048KB, select Solution 4; when 2048KB < N ≤ 8192KB, select Solution 5; when 8192KB < N, select the general memory access method; After completing the above steps, the master core proceeds according to the selected memory access solution, transfers the new array B and the mapping array C to the slave-core LDM through DMA, and executes the subsequent computing tasks.

5. The method for optimizing memory access of slave cores based on the new generation ShenWei many-core processor according to any one of claims 1-3, characterized in that, Implement the slave-core memory access optimization method based on the new-generation Shenwei many-core processor through the automated interface STMA.

Citation Information

Patent Citations

  • AztecOO transplantation optimization method and system based on Shenwei super computer

    CN118656126A

  • Communication network operation and maintenance fault positioning and tracking method and system

    CN119420639A