Optimization Method for Limited Computing Resources and Communication Redundancy Based on the New Generation of Shenwei Processors

By applying for multi-core group collaborative computing and optimization of internal and external data transmission on the new generation of Shenwei processors, the problem of insufficient acceleration effect of programs in process-level optimization is solved, and more efficient computing resource utilization and program execution efficiency are achieved.

CN119292794BActive Publication Date: 2025-07-01SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411826361.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-07-01
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

In program optimization based on the new generation of Shenwei processors, when some programs are process-level optimization by adding MPI, the acceleration effect is not obvious, and due to limited computing resources and redundant communication, programs with many cycles and large calculations still have insufficient execution efficiency.

Method used

It is proposed to apply for multi-core groups to allow multi-core groups to calculate in a coordinated manner, and optimize data collection and transmission within and between core groups, reduce the number of master-slave core transmissions, and improve program execution efficiency. An automated interface MCGO is designed to simplify the selection of core group numbers and data aggregation transmission processes.

Benefits of technology

Through multi-core group collaborative computing and data transmission optimization, the efficient use of computing resources is significantly improved, especially when processing hot spot calculation parts with a large number of cycles and a large amount of calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119292794B_ABST
    Figure CN119292794B_ABST
Patent Text Reader

Abstract

The present invention relates to an optimization method for limited computing resources and communication redundancy based on the new generation of Shenwei processors, belonging to the field of electronic information technology. It includes: First, analyze the hot computing part to determine the number of core groups required for the loop count and computing volume of the hot computing part; among them, if the running time of the code in the computing part accounts for more than half of the total code running time, then this computing part is determined as the hot computing part; Then, apply for and number multiple core groups, and let the multiple core groups perform collaborative computing on the hot computing part; Again, after the collaborative computing is completed, group the slave cores within the core group to optimize the collection and transmission of data, and then collect and transmit data between the core groups; Finally, send the final result back to the main memory once. Through the application of multiple core group collaborative optimization, the present invention enables the optimization of slave cores to utilize more computing resources, greatly improving the execution efficiency of the program.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an optimization method for limited computing resources and communication redundancy based on a new generation of Shenwei processors, and belongs to the field of electronic information technology. Background Art

[0002] At present, the new generation of Shenwei supercomputer "Shenwei·Blu-ray II" in China has been put into use, and its Shenwei many-core processor SW26010pro is a representative work of domestic self-developed processors. The development of the Shenwei series of supercomputers has provided strong support for data simulation and scientific research in various fields.

[0003] The SW26010pro chip includes 6 core groups (Core-Group, CG), and each core group includes a management processing element (Management Processing Element, MPE), also known as the main core. An 8×8 computation processing element (Computation Processing Element, CPE), also known as the slave core array. Each core group has its own memory controller (Memory Controller, MC), and the memory is 16 GB of DDR4, which is connected to the memory controller through a direct memory access (Direct Memory Access, DMA) with a bandwidth of 51.2 GB / s. Data transmission between two slave cores in the same slave core array is achieved through remote memory access (Remote Memory Access, RMA). Each slave core has 256 KB of local data memory (Local Data Memory, LDM). Each SW26010pro processor consists of 390 processing units. A single SW26010pro processor can provide a floating-point operation rate of approximately 14 TB / s and a memory access bandwidth performance of approximately 307.2 GB / s through DMA. The hardware architecture of the SW26010pro processor is as Figure 1 shown.

[0004] When optimizing a program using the new generation of Shenwei supercomputer, a combination of process-level optimization and thread-level optimization is usually adopted. More specifically, MPI and Athread technologies are used. Process-level optimization refers to parallelly completing the computing tasks of an application by running processes on the main cores of the core groups, and further improving the program execution efficiency by reducing communication and other means. Thread-level optimization refers to parallelly completing the acceleration tasks started by the process by running threads on each slave core. The process can use the Athread interface to start the slave core array of the core group, and then further divide the tasks to each slave core for execution.

[0005] When optimizing and accelerating a program, for a program with a loop count that can reach thousands, tens of thousands, or even millions, and with a large amount of computation within the loop, more computing resources are often required to improve the program execution efficiency. The general parallel optimization method is a combination of task parallelism and thread parallelism. When optimizing a program based on the new generation of ShenWei many-core processors, the commonly used parallel optimization method is MPI + Athread.

[0006] Task parallelism mainly refers to process-level parallelism between core groups, that is, running processes on the main cores of core groups to parallelly complete the computational tasks of the application. According to the interfaces provided by the parallel programming framework for application development, different processes can be distinguished using MPI process numbers. Thread parallelism mainly refers to accelerating the tasks started by the process by running threads on each slave core. The process can use the Athread interface to start the slave core array of a single core group, further dividing the tasks for each slave core to execute, thus greatly improving the program execution efficiency.

[0007] However, although some programs can be optimized at the process level by adding MPI, due to the relatively long communication time consumption, the acceleration effect may not be obvious. For such programs, it is generally not recommended to use process-level optimization, but it is more suitable to improve the execution speed through better thread-level optimization. Since only thread-level optimization is performed, the available computing resources will be reduced. Therefore, when facing a program with a loop count of thousands, tens of thousands, or even millions and a large amount of computation within the loop, the acceleration effect may still be limited. At the same time, after the computing acceleration of the slave cores is completed, it is usually necessary to reduce or transfer the computing results of each slave core back to the main memory through DMA, but these operations often increase the overall execution time of the program and reduce the efficiency. Summary of the Invention

[0008] Aiming at the deficiencies of the prior art, the present invention provides an optimization method for limited computing resources and communication redundancy based on the new generation of ShenWei processors;

[0009] When optimizing a program on the ShenWei many-core processor, if the addition of MPI to the program does not result in obvious optimization effects, then only better slave-core optimization can be relied on. At this time, the lack of computing resources becomes the main bottleneck. Especially when the number of loops in the hot-spot computing part reaches thousands, tens of thousands, or even millions, and the amount of computation within each loop is relatively large, computing resources become particularly important. When computing resources are sufficient, the time taken by the slave cores to collect the final results and transfer data to the main memory is relatively long, affecting the program execution efficiency. For programs that do not show obvious acceleration effects when adding MPI for process-level optimization, the method of the present invention proposes to apply for a multi-core group, let the multi-core group perform collaborative computing, and optimize data collection and transmission within and between the core groups, reducing the number of master-slave core transmissions, thereby improving the program execution efficiency and performance. At the same time, the present invention also designs an automated interface MCGO (Multi-Core Group Optimization), with the master-core interface used to determine the number of core groups, and the slave-core interface used to perform multi-level summarization and transmission of data. This is convenient for programmers to directly call and use.

[0010] The technical solution of the present invention is as follows:

[0011] Based on the optimization method for limited computing resources and communication redundancy of the new-generation ShenWei processor, it includes:

[0012] First, analyze the hot-spot computing part to determine the number of core groups required based on the number of loops and the amount of computation in the hot-spot computing part; among them, if the running time of the code in the computing part accounts for more than half of the total code running time, then this computing part is determined as the hot-spot computing part;

[0013] Then, apply for and number the multi-core groups, and let the multi-core groups perform collaborative computing on the hot-spot computing part;

[0014] Again, after the collaborative computing is completed, group the slave cores within the core group, optimize data collection and transmission, and then perform data collection and transmission between the core groups;

[0015] Finally, send the final result back to the main memory once.

[0016] According to the preference of the present invention, applying for and numbering the multi-core groups; includes:

[0017] For a program running in a single process, when the number of loops in the hot-spot computing part reaches more than one million times, apply for a multi-core group and start running, number the slave cores of the multi-core group in sequence, and then allocate the hot-spot computing part to the numbered slave cores for computing;

[0018] For a program running in multiple processes, first, divide the tasks, allocate the tasks to multiple processes to run, and each process carries out its own part of the tasks to execute the program; then, each process further distributes its own hot-spot computing part to its own slave cores for execution; finally, perform reduction computing.

[0019] Further preferably, when applying for a multi-core group variable space, variables required for calculation are defined in each core group, and private LDM variables are selected and defined.

[0020] Preferably according to the present invention, the number of core groups required for the number of loop iterations and the amount of calculation in the hot-spot calculation part is determined; that is: predicting the number of core groups by training a linear regression model; including:

[0021] 1) Extracting data of three types of programs as training data; the three types of programs are programs with a computational complexity of , , , where n represents the number of loop executions; O() is a notation for asymptotic upper bound, used to describe the trend of the running time or space requirement of an algorithm as the input scale increases; indicates that the running time of the algorithm is proportional to the input scale; indicates that the running time of the algorithm is proportional to the square of the input scale; indicates that the running time of the algorithm is proportional to the cube of the input scale; the training data includes multiple samples, the data of the three types of programs includes features and target values, the features include the number of loop iterations and the computational complexity; the target value is the number of core groups required under each combination obtained through actual testing; by modifying the number of loop iterations of these three types of programs, the number of loop iterations and the computational complexity are arranged and combined, and the optimal number of core groups for each combination is tested to obtain multiple samples of the training data;

[0022] 2) Using the above training data to train a linear regression model, and the formula of the linear regression model is as follows:

[0023] ;

[0024] Wherein, is the predicted value, is the bias term, is the weight of each feature, is the value of each feature;

[0025] The training process includes: initializing the linear regression model according to the number of features, including weights and bias terms, and setting the initial values of the weights and bias terms to 0; using the training data to train the linear regression model, and through multiple iterations, adjusting the weights and bias terms to make the linear regression model fit the training data; using the mean square error as the loss function, and continuously optimizing the parameters of the linear regression model through the gradient descent method to obtain the trained linear regression model;

[0026] 3) Inputting new data into the trained linear regression model for prediction, and outputting the predicted number of core groups.

[0027] Preferably according to the present invention, a multi-core group performs collaborative computing on the hotspot computing part; including:

[0028] When selecting a 2-core group for collaborative computing, the thread number range is 0 - 127, the core group number range is 0 - 1, and each thread has its corresponding core group number; according to the characteristics of the hotspot computing part, the computing tasks are assigned to each slave core, specifically: dividing the total task volume by the total number of threads to obtain the task volume that each slave core specifically needs to operate, and each slave core executes the computing of its own part of the task; each slave core obtains the required data from the master core through the DMA method or directly accessing the main memory, each slave core independently completes its own part of the task, the slave core computations between each core group do not interfere with each other, and reduction operations and RMA transmission operations are performed on the computing part within the core group, and the results are sent back to the master core;

[0029] When selecting a 3-core group for collaborative computing, the thread number range is 0 - 191, the core group number range is 0 - 2, and each thread has its corresponding core group number; according to the characteristics of the hotspot computing part, the computing tasks are assigned to each slave core, specifically: dividing the total task volume by the total number of threads to obtain the task volume that each slave core specifically needs to operate, and each slave core executes the computing of its own part of the task; each slave core obtains the required data from the master core through the DMA method or directly accessing the main memory, each slave core independently completes its own part of the task, the slave core computations between each core group do not interfere with each other, and reduction operations and RMA transmission operations are performed on the computing part within the core group, and the results are sent back to the master core;

[0030] When selecting a 6-core group for collaborative computing, the thread number range is 0 - 383, the core group number range is 0 - 5, and each thread has its corresponding core group number; according to the characteristics of the hotspot computing part, the computing tasks are assigned to each slave core, specifically: dividing the total task volume by the total number of threads to obtain the task volume that each slave core specifically needs to operate, and each slave core executes the computing of its own part of the task; each slave core obtains the required data from the master core through the DMA method or directly accessing the main memory, each slave core independently completes its own part of the task, the slave core computations between each core group do not interfere with each other, and reduction operations and RMA transmission operations are performed on the computing part within the core group, and the results are sent back to the master core.

[0031] Preferably according to the present invention, when optimizing with a 3-core group, the slave cores within the core group are grouped to optimize the collection and transmission of data, and then data collection and transmission are performed between the core groups; including:

[0032] First, when multiple slave cores execute tasks, divide the slave cores into groups. Every four slave cores form a slave core group. The logical numbers of each slave core are 0, 1, 2, and 3, and allocate continuous shared space in the LDM for the slave cores in each slave core group. After each slave core completes its calculation, put the result into the continuous shared space in the LDM. Designate the slave core with logical number 0 in each slave core group as the third-level summary slave core, which is responsible for collecting the calculation results of other slave cores in this group.

[0033] Then, designate the slave core with number 0 in each core group as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core within this core group. The specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through RMA operation.

[0034] Finally, set the slave core with number 0 in core group number 0 as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave cores in core groups numbered 1 and 2, so as to obtain the final calculation result. Then, the first-level summary slave core sends this final calculation result back to the main memory through DMA.

[0035] According to the preference of the present invention, when 2 core groups cooperate for optimization, group the slave cores within the core group to optimize data collection and transmission, and then perform data collection and transmission between core groups. It includes:

[0036] First, when multiple slave cores execute tasks, divide the slave cores into groups. Every four slave cores form a slave core group. The logical numbers of each slave core are 0, 1, 2, and 3, and allocate continuous shared space in the LDM for the slave cores in each slave core group. After each slave core completes its calculation, put the result into the continuous shared space in the LDM. Designate the slave core with logical number 0 in each slave core group as the third-level summary slave core, which is responsible for collecting the calculation results of other slave cores in this group.

[0037] Then, designate the slave core with number 0 in each core group as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core within this core group. The specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through RMA operation.

[0038] Finally, set the slave core with number 0 in core group number 0 as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave core in core group numbered 1, so as to obtain the final calculation result. Then, the first-level summary slave core sends this final calculation result back to the main memory through DMA.

[0039] According to the preference of the present invention, when 6 core groups cooperate for optimization, group the slave cores within the core group to optimize data collection and transmission, and then perform data collection and transmission between core groups. It includes:

[0040] First, when multi-core tasks are executed, slave core groups are divided. Every four slave cores form a slave core group. The logical numbers of each slave core are 0, 1, 2, and 3, and continuous shared space in the LDM is allocated to the slave cores in each slave core group. After each slave core finishes its calculation, the result is put into the continuous shared space in the LDM. In each slave core group, the slave core with logical number 0 is designated as the third-level summary slave core, which is responsible for collecting the calculation results of other slave cores in this group.

[0041] Then, in each core group, the slave core numbered 0 is designated as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core within this core group. The specific operation is as follows: Each third-level summary slave core sends the result to the second-level summary slave core through RMA operations.

[0042] Finally, the slave core numbered 0 in core group numbered 0 is set as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave cores in core groups numbered 1, 2, 3, 4, and 5, so as to obtain the final calculation result. Then, the first-level summary slave core sends this final calculation result back to the main memory through DMA.

[0043] Preferably according to the present invention, data collection and transmission within and between core groups are realized through the automated interface MCGO.

[0044] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the optimization method for limited computing resources and communication redundancy based on the new generation of Shenwei processors are realized.

[0045] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the optimization method for limited computing resources and communication redundancy based on the new generation of Shenwei processors are realized.

[0046] The beneficial effects of the present invention are as follows:

[0047] 1. Considering the situation of insufficient computing resources in single-process optimization, the present invention enables slave core optimization to utilize more computing resources through applying multi-core group collaborative optimization, greatly improving the execution efficiency of the program.

[0048] 2. Considering that a lot of time is consumed for result reduction and transmission after the slave cores finish their tasks, the present invention improves the program execution efficiency by optimizing the summary and transmission operations within and between core groups.

[0049] 3. Considering the selection of the number of multi-core groups and the fact that the optimization of multi-core group data summary and transmission involves cumbersome steps, the present invention designs the MCGO automated interface to simplify user operations and improve the user experience. Description of the Drawings

[0050] Figure 1Schematic diagram of the hardware architecture of the SW26010pro processor;

[0051] Figure 2 Schematic diagram corresponding to a single-process multi-core group;

[0052] Figure 3 Schematic diagram of the process of predicting the number of core groups by training a linear regression model;

[0053] Figure 4 Schematic diagram of the collaborative computing process of multi-core groups;

[0054] Figure 5 Schematic diagram of the third-level summary slave cores;

[0055] Figure 6 Schematic diagram of the second-level summary slave cores;

[0056] Figure 7 Schematic diagram of the first-level summary slave cores. Detailed implementation manners

[0057] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0058] Embodiment 1

[0059] Based on the computing resource limitation and communication redundancy optimization method of the new generation ShenWei processor, including:

[0060] First, analyze the hot computing part to judge the number of core groups required for the loop times and computing amount of the hot computing part; among them, if the running time of the code in the computing part accounts for more than half of the total code running time, then it is determined that this computing part is the hot computing part; the hot computing part usually refers to the code or data processing area that requires high-frequency computing and performance optimization. These "hot spots" are often the bottlenecks of program performance and need special attention for optimization. Generally, it is specifically manifested as a large number of loop times and high computing complexity. There are generally two methods for analyzing the hot computing part of a program: one is to use the Swprof performance analysis tool carried by the ShenWei supercomputer itself, add the corresponding compilation options during compilation, run the program and output the specific function running time ratio of the program. The other is to use the method of manual instrumentation, insert a timer in the key code segment, output the running duration of this code segment, and determine the code segment with a long running time. If the running time of this part of the code accounts for more than half of the total code running time, then it is determined that this part is the hot computing part.

[0061] Then, apply for and number the multi-core groups, and let the multi-core groups perform collaborative computing on the hot computing part;

[0062] Again, after the collaborative computing is completed, group the slave cores within the core group to optimize the collection and transmission of data, and then collect and transmit data between the core groups;

[0063] Finally, the final result is transmitted back to the main memory once. This reduces the number of transmissions between the master and slave cores, greatly improving the program execution efficiency.

[0064] Embodiment 2

[0065] Based on the optimization method for computing resource limitation and communication redundancy of the new generation ShenWei processor described in Embodiment 1, the difference is as follows:

[0066] Multi-core group application and numbering; including:

[0067] For a program running in a single process, when the number of loops in the hot computing part reaches more than one million times, it will be very strenuous for a single core group to handle. Therefore, it is an inevitable requirement to adopt multi-core group collaborative computing and let more slave cores participate in the computing. Apply for a multi-core group and start running it. Number the slave cores of the multi-core group in sequence, and then allocate the hot computing part to the numbered slave cores for computing; effectively improving the computing efficiency. The corresponding schematic diagram of a single-process multi-core group is as Figure 2 shown.

[0068] For a program running in multiple processes, first, divide the tasks and allocate the tasks to multiple processes to run. Each process carries a part of its own tasks to execute the program; then, each process further distributes its own hot computing part to its own slave cores for execution; finally, perform reduction computing.

[0069] When applying for the variable space of multiple core groups, variables required for calculation are defined in each core group, and either private LDM variables or cross-segment variables can be selected for definition. A private LDM variable is a high-speed local local data storage space of each slave core of the ShenWei many-core processor, with a capacity of 256KB. The LDM private space is the local private space that can be quickly accessed by this slave core. A cross-segment is a shared space between core groups on the same chip. Both the master and slave cores can access this distributed shared area of the entire chip, and it supports cacheable and non-cacheable access to the main memory space. Due to the shared attribute of the cross-segment, if the required variables are defined as cross-segment variables, locking needs to be performed before and after the calculation to ensure the correctness of the calculation result. However, the locking operation will affect the running efficiency of the program. Therefore, private LDM variables are selected for definition here. Therefore, when applying for the variable space of multiple core groups, variables required for calculation are defined in each core group, and private LDM variables are selected for definition. Among them, for the Local Data Memory (LDM), each slave core of the ShenWei many-core processor has a high-speed local local data storage space LDM, and the total capacity of LDM is 256KB. The LDM space is mainly divided into a private space and a continuous shared space. The LDM private space is the local private space that can be quickly accessed by this slave core; the LDM continuous shared space is a space used by the slave core for shared access within the array of LDM. The slave core can access the LDM space of other slave cores. The LDM continuous shared space is a space used by the slave core for shared access within the array of LDM. The slave core can access the LDM space of other slave cores.

[0070] Judge the number of core groups required for the number of loop iterations and the amount of calculation in the hot-spot calculation part; that is: predicting the number of core groups by training a linear regression model; including:

[0071] When selecting the specific number of core groups, it is usually necessary to analyze the number of loop iterations and the computational complexity of the hot-spot calculation part, and estimate the required number of core groups. To more accurately determine the number of core groups, it is also necessary to run the program to test the execution effects of different numbers of core groups and compare them. To further simplify the process of determining the number of core groups, a machine learning model can be introduced. By training a model to predict the optimal number of core groups, the optimal configuration can be selected according to features such as the number of loop iterations and the computational complexity.

[0072] 1) Extract data of three types of programs as training data; the three types of programs are those with computational complexities of , , respectively, where n is the input scale of the problem, which refers to the size or quantity of the input data, representing the number of elements or the amount of data that the algorithm needs to process, and here it represents the number of times the loop is executed; O() is a notation for asymptotic upper bound, used to describe the trend of the running time or space requirement of the algorithm as the input scale grows. Indicates that the running time of the algorithm is proportional to the input size; Indicates that the running time of the algorithm is proportional to the square of the input size; Indicates that the running time of the algorithm is proportional to the cube of the input size; The selection of these programs is based on their characteristics of having a large number of loops and requiring a large amount of computing resources, which is in line with the application scenario of the present invention. The training data includes multiple samples. The data of the three types of programs includes features and target values. The features include the number of loops and the computational complexity; The computational complexity is obtained by judging the structure of the code; For example, the complexity of a single-layer loop depends on the number of loops. If the loop is executed n times, it is . The complexity of a nested loop is the product of the number of loops in the outer and inner layers. The outer layer is n, the inner layer is n, and the total is . The computational complexity is defined as a specific value for convenience as a parameter when the function is called. The target value is the number of core groups required under each combination obtained through actual testing; By modifying the number of loops of these three types of programs and arranging and combining the number of loops and the computational complexity, the optimal number of core groups for each combination is tested to obtain multiple samples of the training data; The training data is used to train a linear regression model, and the new feature data is used to predict the number of core groups.

[0073] 2) Using the above training data, train a linear regression model. The formula of the linear regression model is as follows:

[0074] ;

[0075] Among them, is the predicted value, is the bias term (intercept), is the weight of each feature, is the value of each feature;

[0076] The training process includes: initializing the linear regression model according to the number of features, including weights and bias terms, and setting the initial values of the weights and bias terms to 0; using the training data to train the linear regression model, and through multiple iterations, adjusting the weights and bias terms so that the linear regression model can better fit the training data; using the mean squared error as the loss function and continuously optimizing the parameters of the linear regression model through the gradient descent method to reduce the prediction error and obtain the trained linear regression model;

[0077] 3) Input the new data into the trained linear regression model for prediction, and output the predicted number of core groups. After obtaining the appropriate number of core groups, apply for the specific number of core groups in the program to optimize the use of computing resources. The specific process is as Figure 3 shown:

[0078] Multiple core groups perform collaborative computing on the hot computing part; including:

[0079] When selecting 2-core group collaborative computing, the thread number range is 0 - 127, the core group number range is 0 - 1, and each thread has its corresponding core group number; according to the characteristics of the hotspot calculation part, the computing tasks are reasonably allocated to each slave core. Specifically, it means dividing the total task volume by the total number of threads to obtain the task volume that each slave core specifically needs to operate, and each slave core executes the calculation of its own part of the task; each slave core obtains the required data from the master core through the DMA method or directly accessing the main memory. Each slave core independently completes its own part of the task, and the slave core calculations between core groups do not interfere with each other. The reduction operation and RMA transfer operation are performed on the calculation part within the core group, and the results are sent back to the master core; RMA is a remote data transfer operation between the slave core LDMs within the core group.

[0080] When selecting 3-core group collaborative computing, the thread number range is 0 - 191, the core group number range is 0 - 2, and each thread has its corresponding core group number; according to the characteristics of the hotspot calculation part, the computing tasks are reasonably allocated to each slave core. Specifically, it means dividing the total task volume by the total number of threads to obtain the task volume that each slave core specifically needs to operate, and each slave core executes the calculation of its own part of the task; each slave core obtains the required data from the master core through the DMA method or directly accessing the main memory. Each slave core independently completes its own part of the task, and the slave core calculations between core groups do not interfere with each other. The reduction operation and RMA transfer operation are performed on the calculation part within the core group, and the results are sent back to the master core;

[0081] When selecting 6-core group collaborative computing, the thread number range is 0 - 383, the core group number range is 0 - 5, and each thread has its corresponding core group number; according to the characteristics of the hotspot calculation part, the computing tasks are reasonably allocated to each slave core. Specifically, it means dividing the total task volume by the total number of threads to obtain the task volume that each slave core specifically needs to operate, and each slave core executes the calculation of its own part of the task; each slave core obtains the required data from the master core through the DMA method or directly accessing the main memory. Each slave core independently completes its own part of the task, and the slave core calculations between core groups do not interfere with each other. The reduction operation and RMA transfer operation are performed on the calculation part within the core group, and the results are sent back to the master core.

[0082] The specific process of multi-core group calculation is as Figure 4 shown:

[0083] When optimizing the single-core group and multi-core group, the tasks processed are often critical resources, resulting in possible competition problems when multiple slave cores access the same storage area. To avoid competition errors, it is necessary to allocate an independent private LDM space for each slave core so that it only processes and updates private variables. After the task is completed, each slave core needs to summarize or send back the calculation results to the master core, and then the master core performs the final calculation. However, this operation will increase the summarization and transmission overhead in the single-core group and multi-core group environments, so it is necessary to further optimize the summarization and transmission processes.

[0084] When optimizing with a 3-core group in cooperation, the slave cores within the core group are grouped, and the collection and transmission of data are optimized. Then, data collection and transmission are carried out between the core groups. It includes:

[0085] First, when multiple slave core tasks are executed, slave core groups are divided. Every four slave cores form a slave core group. The logical numbers of each slave core are 0, 1, 2, and 3, and continuous shared space in LDM is allocated to the slave cores of each slave core group. After each slave core finishes its calculation, the result is put into the continuous shared space in LDM. Each slave core group designates the slave core with logical number 0 as the third-level summary slave core, which is responsible for collecting the calculation results of other slave cores within this group. Because the continuous shared space in LDM is used within the group, the third-level summary slave core directly collects the data in this space. Specifically, as Figure 5 shown:

[0086] Then, in each core group, the slave core numbered 0 is designated as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core within this core group. The specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through RMA operation. Specifically, as Figure 6 shown:

[0087] Finally, the slave core numbered 0 in the core group numbered 0 is set as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave cores in core groups numbered 1 and 2, so as to obtain the final calculation result. Then, the first-level summary slave core sends this final calculation result back to the main memory through DMA. The specific operation is as Figure 7 shown.

[0088] When optimizing with a 2-core group in cooperation, the slave cores within the core group are grouped, and the collection and transmission of data are optimized. Then, data collection and transmission are carried out between the core groups. It includes:

[0089] First, when multiple slave core tasks are executed, slave core groups are divided. Every four slave cores form a slave core group. The logical numbers of each slave core are 0, 1, 2, and 3, and continuous shared space in LDM is allocated to the slave cores of each slave core group. After each slave core finishes its calculation, the result is put into the continuous shared space in LDM. Each slave core group designates the slave core with logical number 0 as the third-level summary slave core, which is responsible for collecting the calculation results of other slave cores within this group. Because the continuous shared space in LDM is used within the group, the third-level summary slave core directly collects the data in this space.

[0090] Then, in each core group, the slave core numbered 0 is designated as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core within this core group. The specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through RMA operation;

[0091] Finally, the slave core numbered 0 with core group number 0 is set as the primary summary slave core, responsible for summarizing the results of the secondary summary slave cores with core group number 1, so as to obtain the final calculation result, and then the primary summary slave core transmits this final calculation result back to the main memory through DMA.

[0092] When six core groups cooperate for optimization, the slave cores within the core group are grouped to optimize data collection and transmission, and then data collection and transmission are carried out between core groups, including:

[0093] First, when multi-slave core tasks are executed, slave core groups are divided. Every four slave cores form a slave core group, and the logical numbers of each slave core are 0, 1, 2, 3. And continuous shared space in LDM is allocated to the slave cores of each slave core group; after each slave core finishes calculation, the result is put into the continuous shared space in LDM; the slave core with logical number 0 in each slave core group is designated as the tertiary summary slave core, responsible for collecting the calculation results of other slave cores within this group; because the continuous shared space in LDM is used within the group, the tertiary summary slave core directly collects the data in this space.

[0094] Then, the slave core numbered 0 in each core group is designated as the secondary summary slave core, responsible for summarizing the results of each tertiary summary slave core within this core group; the specific operation is: each tertiary summary slave core sends the result to the secondary summary slave core through RMA operation; specifically as Figure 6 shown:

[0095] Finally, the slave core numbered 0 with core group number 0 is set as the primary summary slave core, responsible for summarizing the results of the secondary summary slave cores with core group numbers 1, 2, 3, 4, 5, so as to obtain the final calculation result, and then the primary summary slave core transmits this final calculation result back to the main memory through DMA.

[0096] Data collection and transmission within and between core groups are realized through the automated interface MCGO (Multi-Core Group Optimization).

[0097] To facilitate programmers to call the method of the present invention, the automated interface MCGO is designed. This automated interface MCGO has four functions, namely: initialization interface function, predicted core group number function, multi-level summary and transmission function for data, and release interface function. The design of the automated interface MCGO aims to better apply multi-core group data summary and transmission optimization. The implementation of this interface reduces the programming difficulty based on the Shenwei supercomputer. Programmers can directly call this interface to assist programming. The description of the automated interface MCGO is shown in Table 1:

[0098] Table 1 MCGO interface description;

[0099]

[0100] An example of calling the MCGO interface is as follows:

[0101] / / Call the MCGO interface in the main core host.c

[0102] #include “MCGO.h”

[0103] int main()

[0104] {

[0105] / / Define the required variables

[0106] / / Initialize the interface

[0107] MCGO_Init()

[0108] / / Call the interface

[0109] MCGO_Master(loop_count, complexity, morecore_count);

[0110] / / Terminate the interface

[0111] MCGO_Finalize();

[0112] return 0;

[0113] }

[0114] Another example of calling the MCGO interface is as follows:

[0115] / / Call the MCGO interface in the main core slave.c

[0116] #include “MCGO.h”

[0117] int main()

[0118] {

[0119] / / Define the required variables

[0120] / / Initialize the interface

[0121] MCGO_Init()

[0122] / / Call the interface

[0123] MCGO_Slave (tid, morecore_count, value, result);

[0124] / / Terminate the interface

[0125] MCGO_Finalize();

[0126] return 0;

[0127] }

[0128] Through the implementation of the above method, the method of the present invention was tested.

[0129] First, the program was optimized for general co-processor utilization, that is, the hot calculation part was distributed to each co-processor. After each co-processor completed the calculation, the results were reduced and transmitted back to the main memory to obtain the running time of the general optimization. Then, multi-core group collaborative optimization and transmission optimization were carried out to obtain the running time of the invention optimization. Due to the hardware of the ShenWei processor, a single process can run at most 6 core groups. Therefore, the following gives the test of up to 6 core groups.

[0130] Table 2 shows the comparison of the results between the general optimization and the optimization method of the present invention;

[0131] Table 2 Multi-core group prediction results;

[0132]

[0133] Example 3

[0134] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for optimizing calculation resource limitation and communication redundancy based on the new generation of ShenWei processor described in Example 1 or 2.

[0135] Example 4

[0136] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method for optimizing calculation resource limitation and communication redundancy based on the new generation of ShenWei processor described in Example 1 or 2.

Claims

1. Based on the new generation of Shenwei processors, the computing resource limitation and communication redundancy optimization method is characterized by: include: First, the hotspot calculation part is analyzed to determine the number of hotspot calculation part loops and the number of core groups required for the calculation amount; if the code running time of the calculation part accounts for more than half of the total code running time, then the calculation part is determined to be a hotspot calculation part; Then, the multi-core group applies for and numbers the hotspot computing part, allowing the multi-core group to perform collaborative computing; Thirdly, after the collaborative computing is completed, the slave cores in the core group are grouped to optimize the data collection and transmission, and the data collection and transmission are then performed between the core groups; Finally, the final result is transferred back to main memory in a single pass; Determine the number of core groups required for the number of cycles and the amount of calculation for the hotspot calculation part; refers to: predicting the number of core groups by training a linear regression model; including: 1) Extract data from three types of programs as training data; the three types of programs have computational complexity of O(n), O(n 2 ), O(n 3 ), n is the number of times the loop is executed; O() is an asymptotic upper bound, which is used to describe the changing trend of the running time or space requirement of the algorithm as the input size grows; O(n) means that the running time of the algorithm is proportional to the input size; O(n 2 ) indicates that the running time of the algorithm is proportional to the square of the input size; O(n 3 ) indicates that the running time of the algorithm is proportional to the cube of the input size; the training data includes multiple samples, and the data of the three types of programs include features and target values, and the features include the number of loops and the computational complexity; the target value is the number of core groups required for each combination obtained from the actual test; by modifying the number of loops of the three types of programs, the number of loops and the computational complexity are arranged and combined, and the optimal number of core groups for each combination is tested, and multiple samples of training data are obtained; 2) Use the above training data to train a linear regression model. The formula of the linear regression model is as follows: y=β0+β1x1+β2x2+…+β i x i …b n x n ; Among them, y is the predicted value, β0 is the bias term, and β i is the weight of each feature, x i is the value of each feature; The training process includes: initializing the linear regression model according to the number of features, including weights and bias terms, with the initial values ​​of weights and bias terms set to 0; using the training data to train the linear regression model, adjusting the weights and bias terms through multiple iterations so that the linear regression model fits the training data; using the mean square error as the loss function, continuously optimizing the parameters of the linear regression model through the gradient descent method to obtain a trained linear regression model; 3) Input the new data into the trained linear regression model for prediction and output the predicted number of core groups; The multi-core group performs collaborative computing on the hotspot computing part, including: When selecting 3-core group collaborative computing, the thread number range is 0-191, and the core group number range is 0-2. Each thread has its own corresponding core group number. According to the characteristics of the hot spot computing part, the computing tasks are allocated to each slave core. Specifically, the total task amount is divided by the total number of threads to obtain the specific task amount that each slave core needs to operate, and each slave core performs the calculation of its own part of the task. Each slave core obtains the required data from the master core through DMA or direct access to the main memory. Each slave core independently completes its own part of the task. The slave core calculations between the core groups do not interfere with each other. The calculation part is reduced and RMA transmission operations are performed within the core group, and the results are transmitted back to the master core. When three core groups are used for collaborative optimization, the slave cores in the core group are grouped to optimize data collection and transmission, and data collection and transmission are then performed between core groups; including: First, when executing a multi-slave core task, the cores are divided into groups. Every four cores form a core group. The logical number of each core is 0, 1, 2, 3, and LDM continuous shared space is allocated to each core in the core group. After each core is calculated, the result is placed in the LDM continuous shared space. Each core group designates the core with the logical number 0 as the third-level summary core, which is responsible for collecting the calculation results of other cores in the group. Then, in each core group, slave core No. 0 is designated as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core in the core group; the specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through the RMA operation; Finally, slave core No. 0 of core group number 0 is set as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave cores of core groups numbered 1 and 2 to obtain the final calculation result, and then the first-level summary slave core transfers the final calculation result back to the main memory through DMA.

2. The method for optimizing computing resource constraints and communication redundancy based on a new generation of Shenwei processors according to claim 1 is characterized in that: Multi-core group application and numbering; including: For a single-process program, when the number of cycles of the hotspot calculation part reaches more than one million times, apply for a multi-core group and start running, number the slave cores of the multi-core group in sequence, and then assign the hotspot calculation part to the numbered slave cores for calculation; For a program running in multiple processes, first, divide the tasks and assign them to multiple processes to run. Each process carries its own part of the task execution program. Then, each process assigns its own hot spot computing part to its own slave core for execution. Finally, perform the reduction calculation.

3. The method for optimizing computing resource constraints and communication redundancy based on a new generation of Shenwei processors according to claim 2 is characterized in that: When applying for multi-core group variable space, define the variables required for calculation in each core group and choose to define private LDM variables.

4. The method for optimizing computing resource constraints and communication redundancy based on a new generation of Shenwei processors according to claim 1 is characterized in that: The multi-core group performs collaborative computing on the hotspot computing part; it also includes: When selecting 2-core group collaborative computing, the thread number range is 0-127, and the core group number range is 0-1. Each thread has its own corresponding core group number. According to the characteristics of the hot spot computing part, the computing tasks are allocated to each slave core. Specifically, the total task amount is divided by the total number of threads to obtain the specific task amount that each slave core needs to operate, and each slave core performs the calculation of its own part of the task. Each slave core obtains the required data from the master core through DMA or direct access to the main memory. Each slave core independently completes its own part of the task. The slave core calculations between the core groups do not interfere with each other. The calculation part is reduced and RMA transmission operations are performed within the core group, and the results are transmitted back to the master core. When selecting 6-core group collaborative computing, the thread number range is 0-383, the core group number range is 0-5, and each thread has its own corresponding core group number; according to the characteristics of the hot spot computing part, the computing tasks are assigned to each slave core, specifically: the total task amount is divided by the total number of threads to obtain the specific task amount that each slave core needs to operate, and each slave core performs the calculation of its own part of the task; each slave core obtains the required data from the main core through DMA or direct access to the main memory. Each slave core independently completes its own part of the task, and the slave core calculations between the core groups do not interfere with each other. The calculation part is reduced and RMA transmission operations are performed within the core group, and the results are transmitted back to the main core.

5. The method for optimizing computing resource constraints and communication redundancy based on a new generation of Shenwei processors according to claim 4 is characterized in that: When the two-core group is used for collaborative optimization, the slave cores in the core group are grouped to optimize data collection and transmission, and data collection and transmission are performed between the core groups; including: First, when executing a multi-slave core task, the cores are divided into groups. Every four cores form a core group. The logical number of each core is 0, 1, 2, 3, and LDM continuous shared space is allocated to each core in the core group. After each core is calculated, the result is placed in the LDM continuous shared space. Each core group designates the core with the logical number 0 as the third-level summary core, which is responsible for collecting the calculation results of other cores in the group. Then, in each core group, slave core No. 0 is designated as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core in the core group; the specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through the RMA operation; Finally, slave core No. 0 of core group number 0 is set as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave cores of core group number 1 to obtain the final calculation result, and then the first-level summary slave core transfers the final calculation result back to the main memory through DMA.

6. The method for optimizing computing resource constraints and communication redundancy based on a new generation of Shenwei processors according to claim 4 is characterized in that: When the 6-core group is used for collaborative optimization, the slave cores in the core group are grouped to optimize data collection and transmission, and data collection and transmission are performed between core groups; including: First, when executing a multi-slave core task, the cores are divided into groups. Every four cores form a core group. The logical number of each core is 0, 1, 2, 3, and LDM continuous shared space is allocated to each core in the core group. After each core is calculated, the result is placed in the LDM continuous shared space. Each core group designates the core with the logical number 0 as the third-level summary core, which is responsible for collecting the calculation results of other cores in the group. Then, in each core group, slave core No. 0 is designated as the second-level summary slave core, which is responsible for summarizing the results of each third-level summary slave core in the core group; the specific operation is: each third-level summary slave core sends the result to the second-level summary slave core through the RMA operation; Finally, slave core No. 0 of core group number 0 is set as the first-level summary slave core, which is responsible for summarizing the results of the second-level summary slave cores of core groups numbered 1, 2, 3, 4, and 5 to obtain the final calculation result, and then the first-level summary slave core transfers the final calculation result back to the main memory through DMA.

7. The method for optimizing computing resource constraints and communication redundancy based on a new generation of Shenwei processors according to any one of claims 1 to 6, characterized in that: Data collection and transmission within and between nuclear groups are realized through the automated interface MCGO.

Citation Information

Patent Citations

  • Neural network heterogeneous many-core multi-level resource mapping method based on compilation

    CN114253545A

  • Three-dimensional strain simulation PCG parallel optimization method and system based on Shenwei architecture

    CN114970294A

  • Method and system for increasing speed of parallel writing of shared main memory critical resources by slave cores based on new generation SW many-core processor

    CN116909741A