Electromagnetic scattering-oriented sparse approximate inverse Shenwei parallel preprocessing method and system

By adopting performance hotspot analysis, master-slave core parallel computing and a variety of fine-grain optimization strategies in the sparse approximate inverse preprocessor, the problem of inefficiency in electromagnetic scattering simulation of traditional sparse approximate inverse preprocessors is solved, and more efficient solution performance and higher accuracy are achieved.

CN120104370APending Publication Date: 2025-06-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510079008.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional sparse approximate inverse preprocessors have problems such as long construction time and large memory consumption in electromagnetic scattering simulation, and fail to fully utilize the multi-core parallel computing power of Shenwei architecture, resulting in inefficient solution.

Method used

A sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering is proposed. Through performance hotspot analysis, parallel calculation of master-slave cores, adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on hybrid precision, the construction and solution efficiency of sparse approximate inverse preprocessors are improved.

Benefits of technology

It significantly improves the construction and solution efficiency of sparse approximate inverse preprocessors, enhances the solution speed and accuracy of complex linear systems in electromagnetic scattering simulation, and meets the needs of real-time and large-scale simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104370A_ABST
    Figure CN120104370A_ABST
Patent Text Reader

Abstract

The invention provides a sparse approximate inverse Shenwei parallel preprocessing method and system for electromagnetic scattering, and relates to the technical field of processor parallel computing, and the method comprises the steps: carrying out the performance hotspot analysis of a sparse approximate inverse preprocessor in an electromagnetic scattering simulation process, and determining a hotspot function of the sparse approximate inverse preprocessor; extracting a sparse matrix vector multiplication operation of a hotspot function, performing master-slave core parallel calculation on the sparse matrix vector multiplication operation, migrating the sparse matrix vector multiplication operation to slave core calculation, and accelerating hotspot calculation by using a heterogeneous parallel sparse approximate inverse algorithm; parallel computing of master-slave core intensive numerical computing tasks is achieved; in the parallel computing process, the slave cores communicate with the master core in a direct memory access mode, and during the period, the master core is in a waiting state to ensure data consistency and communication synchronization until all the slave cores compute distributed tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of processor parallel computing, and in particular to a sparse approximate inverse parallel preprocessing method and system for electromagnetic scattering. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] The new generation of Sunway supercomputer is equipped with the domestically developed SW26010-pro processor, which integrates 6 core groups, each of which contains an operation control core (MPE) and an 8x8 slave core array computing core (CPE). The core groups interact with each other through an on-chip interconnect network. This architecture supports multiple programming models and parallel computing modes, providing a strong hardware foundation for solving complex computing problems. The solution of sparse linear equations is an indispensable key link in scientific and engineering computing, and is widely used in electromagnetic scattering analysis, antenna design, radar cross-section calculation and other fields. Electromagnetic scattering refers to the phenomenon that when electromagnetic waves (such as light waves, microwaves, etc.) meet objects, electromagnetic waves are reflected, refracted or transmitted on the surface or body of the object, forming different scattering patterns. In electromagnetic scattering problems, it is usually necessary to solve large sparse linear equations caused by complex geometric structures and material properties. Especially in tasks such as high-frequency electromagnetic simulation, multi-physics field coupling analysis and inverse scattering imaging, these equations are not only large in scale but also highly sparse and irregular.

[0004] In order to effectively solve these large-scale sparse linear equations, iterative methods are one of the commonly used means, such as the conjugate gradient method (CG), the generalized minimum residual method (GMRES), etc. However, the convergence speed of the iterative method directly may be slow, especially when dealing with electromagnetic scattering problems with complex boundaries. To speed up the iterative process, preconditioners are widely used to improve the initial guess value, thereby reducing the number of iterations required to reach a solution.

[0005] Sparse Approximate Inverse (SPAI) is a special type of preprocessor that approximates the inverse matrix of the original coefficient matrix by constructing a sparse matrix. The SPAI preprocessor has good parallelism and good adaptability to irregular matrices, so it has been widely used in high-performance computing environments. However, in electromagnetic scattering simulations, traditional SPAI preprocessors may encounter problems such as long construction time and high memory consumption, especially when running on specific hardware architectures such as the Shenwei architecture, it fails to fully utilize its multi-core parallel computing capabilities, resulting in low solution efficiency and failure to meet the requirements of real-time and large-scale simulations. Summary of the invention

[0006] In order to solve the above problems, the present invention proposes a sparse approximate inverse Shenwei parallel preprocessing method and system for electromagnetic scattering. In combination with the characteristics of the Shenwei architecture, an efficient parallelization strategy suitable for the many-core environment is designed to improve the construction and solution efficiency of the sparse approximate inverse preprocessor, enhance the solution speed and accuracy of complex linear systems in electromagnetic scattering simulation, and is suitable for solving large-scale sparse linear equations in electromagnetic scattering simulation.

[0007] According to some embodiments, the present disclosure adopts the following technical solutions:

[0008] A sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering, including:

[0009] Perform performance hotspot analysis on the sparse approximate inverse preprocessor in the electromagnetic scattering simulation process and determine the hotspot function of the sparse approximate inverse preprocessor;

[0010] Extract the sparse matrix-vector multiplication operations of the hotspot functions, parallelize the calculation of the sparse matrix-vector multiplication operations on the master and slave cores, migrate the sparse matrix-vector multiplication operations to the slave cores, and use the heterogeneous parallel sparse approximate inverse algorithm to accelerate the calculation of hotspots;

[0011] Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

[0012] According to some embodiments, the present disclosure adopts the following technical solutions:

[0013] The sparse approximate inverse Shenwei parallel preprocessing system for electromagnetic scattering includes:

[0014] A hotspot function determination module is used to perform a performance hotspot analysis on a sparse approximate inverse preprocessor in an electromagnetic scattering simulation process and determine a hotspot function of the sparse approximate inverse preprocessor;

[0015] The accelerated parallel computing module is used to extract the sparse matrix-vector multiplication operations of hotspot functions, parallelize the sparse matrix-vector multiplication operations on the master and slave cores, migrate the sparse matrix-vector multiplication operations to the slave cores, and accelerate the calculation of hotspots using the heterogeneous parallel sparse approximate inverse algorithm;

[0016] Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

[0017] According to some embodiments, the present disclosure adopts the following technical solutions:

[0018] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering is implemented.

[0019] According to some embodiments, the present disclosure adopts the following technical solutions:

[0020] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering is implemented.

[0021] According to some embodiments, the present disclosure adopts the following technical solutions:

[0022] An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] The disclosed sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering adopts job-level performance detection tools and fine-grained plugging technology to perform performance hotspot analysis, determine the main hotspot functions of the sparse approximate inverse preprocessor ParaSails algorithm, use the Shenwei customized Athread acceleration thread library function interface to write the master-slave core code, use the master-slave acceleration model to accelerate the calculation of hotspots, and implement the heterogeneous parallel sparse approximate inverse algorithm swParaSails for the Shenwei architecture, thereby improving the overall iterative method performance. By designing an efficient parallelization strategy suitable for a multi-core environment, it not only improves the construction and solution efficiency of the sparse approximate inverse preprocessor, but also enhances the solution speed and accuracy of complex linear systems in electromagnetic scattering simulation. The disclosure provides a new technical approach for electromagnetic scattering simulation on large-scale parallel computing platforms such as the Shenwei supercomputer, and also provides reference and reference for solving similar large-scale scientific computing tasks.

[0025] The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering disclosed in this paper is equipped with the domestically developed SW26010-pro processor, which integrates 6 core groups, each of which contains an operation control core (MPE) and an 8x8 slave core array computing core (CPE). The core groups interact with each other through an on-chip interconnect network. This architecture supports multiple programming models and parallel computing modes, providing a strong hardware foundation for solving complex computing problems.

[0026] The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering disclosed in the present invention proposes an adaptive data mapping optimization strategy, which dynamically transmits the number of rows of the block matrix according to the different scales of the input sparse matrix, fully utilizes the slave core LDM space, and improves the utilization rate of the slave core LDM; and proposes a kernel aggregation optimization strategy, which aggregates two hot functions MatrixMatvec() and MatrixMatvecTrans(), reduces the time overhead of calling the slave core startup function, and further improves the computing efficiency; also proposes a data communication optimization strategy based on mixed precision, uses single precision to transmit part of the result value to the main memory, and proposes a control algorithm to balance the residual and performance, thereby reducing the communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.

[0028] Figure 1 A flowchart of the overall design of master-slave core acceleration of the sparse approximate inverse preprocessor of the embodiment of the present disclosure;

[0029] Figure 2 An adaptive data mapping optimization strategy under the Sunway Accelerated Computing Framework (SACA) of an embodiment of the present disclosure;

[0030] Figure 3 It is a flowchart of the sparse approximate inverse (SPAI) preprocessor swParaSails in the iterative method of the embodiment of the present disclosure, taking the conjugate gradient (CG) algorithm as an example;

[0031] Figure 4 The figure is a flowchart of the residual and performance balance algorithm of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.

[0033] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.

[0034] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0035] Explanation of terms

[0036] 1. Sunway Accelerate Computing Architecture (SACA) is one of the core components of the new generation Sunway supercomputer. SACA is designed specifically for the Sunway multi-core processor architecture and provides an efficient parallel computing platform and programming model to fully utilize the hardware characteristics of the Sunway system. SACA provides three levels of parallel acceleration, namely process (MPI), thread (Athread) and data (SIMD).

[0037] The new generation of Sunway supercomputers is equipped with the domestically developed SW26010-pro processor, which integrates 6 core groups, each of which contains a computing control core (MPE) and an 8x8 slave core array computing core (CPE). The core groups interact with each other through an on-chip interconnect network. This architecture supports multiple programming models and parallel computing modes, providing a strong hardware foundation for solving complex computing problems.

[0038] 2. ParaSails algorithm is a parallel sparse approximate inverse preconditioning technology. ParaSails is mainly used to accelerate the iterative method of solving large sparse linear systems. It constructs a sparse approximate inverse matrix as a preconditioner and supports parallel processing to improve computing efficiency. It is particularly suitable for sparse matrix calculation optimization in high-performance computing environments.

[0039] Example 1

[0040] In one embodiment of the present disclosure, a sparse approximate inverse parallel preprocessing method for electromagnetic scattering is provided, and the method steps include:

[0041] Step 1: Perform performance hotspot analysis on the sparse approximate inverse preprocessor in the electromagnetic scattering simulation process and determine the hotspot function of the sparse approximate inverse preprocessor;

[0042] Step 2: Extract the sparse matrix-vector multiplication operations of the hotspot functions, parallelize the calculation of the sparse matrix-vector multiplication operations on the master and slave cores, migrate the sparse matrix-vector multiplication operations to the slave cores, and use the heterogeneous parallel sparse approximate inverse algorithm to accelerate the calculation of the hotspots;

[0043] Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

[0044] As an embodiment, the present disclosure provides a parallel optimization method for a sparse approximate inverse preprocessor for electromagnetic scattering on a Sunway supercomputer, aiming to improve the construction efficiency and solution performance of a SPAI preprocessor, and is particularly suitable for solving large-scale sparse linear equations in electromagnetic scattering simulation.

[0045] In general, the propagation of electromagnetic fields follows Maxwell's equations, which can usually be written in the following form:

[0046]

[0047]

[0048] Among them, H and E are magnetic field and electric field respectively, B and D are magnetic induction intensity and electric displacement, J is current density, and ρ is charge density.

[0049] To speed up the iteration process, preconditioners are widely used to improve the initial guess and thus reduce the number of iterations required to reach a solution. Usually a left or right preconditioner M is used to put the linear system into a form that is easy to solve:

[0050] AMy=b,x=My,or MAx=Mb

[0051] Sparse Approximate Inverse (SPAI) is a special type of preprocessor that approximates the inverse matrix A of the original coefficient matrix A by constructing a sparse matrix M. -1 ,Right now:

[0052] M≈A -1

[0053] Sparse Approximate Inverse (SPAI) is a special type of preprocessor that approximates the inverse matrix A of the original coefficient matrix A by constructing a sparse matrix M. -1, the SPAI preprocessor has good parallelism and good adaptability to irregular matrices, so it has been widely used in high-performance computing environments. However, in electromagnetic scattering simulation, traditional SPAI preprocessors may encounter problems such as long construction time and high memory consumption, especially when running on specific hardware architectures such as the Shenwei architecture, it fails to fully utilize its multi-core parallel computing capabilities, resulting in low solution efficiency and unable to meet the requirements of real-time and large-scale simulation. The present disclosure therefore provides a sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering. This method combines the characteristics of the Shenwei architecture and designs an efficient parallelization strategy suitable for the multi-core environment. It not only improves the construction and solution efficiency of the sparse approximate inverse preprocessor, but also enhances the solution speed and accuracy of complex linear systems in electromagnetic scattering simulation. The specific implementation process is as follows:

[0054] Step 1: Preprocessing initialization

[0055] The present invention performs preprocessing initialization and provides a method for optimizing the preprocessing matrix M, especially for different types of coefficient matrices (asymmetric or symmetric positive definite), laying a theoretical and technical foundation for the subsequent implementation of efficient sparse approximate inverse preprocessors on the Shenwei architecture. This method realizes optimization at the algorithm level, and then subsequently designs and implements parallelization strategies under the Shenwei architecture for specific mathematical calculations contained in the sparse approximate inverse algorithm process (such as sparse matrix-vector multiplication, transposed matrix-vector multiplication, and vector dot product operations).

[0056] As an embodiment, it is particularly suitable for sparse approximate inverse preconditioner technology for large-scale linear systems. The method determines M by minimizing the Frobenius norm difference between the identity matrix I and the product of the preconditioning matrix M and the coefficient matrix A, that is:

[0057] min||I-MA|| F

[0058] When the coefficient matrix A is asymmetric, the above objective function is directly used for optimization. If A is a symmetric positive definite matrix, then through Cholesky decomposition A = LL T To simplify the calculation of the preprocessing matrix. At this time, the optimization problem is transformed into finding a lower triangular matrix G such that:

[0059] min||I-GL|| F

[0060] And satisfy M=GG T This method not only improves the computational efficiency of the preprocessing matrix, but also enhances the numerical stability.

[0061] Step 2: Perform a performance hotspot analysis on the sparse approximate inverse preprocessor in the electromagnetic scattering simulation process to determine the hotspot function of the sparse approximate inverse preprocessor;

[0062] Specifically, in order to further improve the optimization performance of the sparse approximate inverse preprocessor ParaSails, the present disclosure uses job-level performance detection tools and fine-grained instrumentation technology to perform performance hotspot analysis. Specifically, the job-level performance monitoring and analysis tool SWPROF configured with the Shenwei architecture is used to determine the hotspot functions of the preprocessor. SWPROF has been built into the Shenwei system environment, and specific parameters are added to the script for submitting jobs to obtain detailed performance hotspot statistics.

[0063] As an example, detailed performance hotspot statistics are obtained as follows: bsub-bq q_linpack-shared-n 4-cgsp 64-share_size 15000-swrunarg-po SPAIout.log. / ParaSails

[0064] in:

[0065] -swrunarg'-p': analyze all master-slave codes and obtain master-slave hotspot functions;

[0066] -swrunarg'-P master': only analyze the main core code and obtain the main core hotspot functions;

[0067] -swrunarg'-P slave': specifically analyzes the slave core code and obtains the slave core hotspot functions.

[0068] Furthermore, after determining the hotspot function, the time-consuming loops therein are further analyzed. To accurately measure the execution time of these loops, the present disclosure adopts two methods: MPI's own statistical time function MPI_Wtime() and a custom timing function timer().

[0069] Specifically, the execution time of the measurement loop is as follows:

[0070] #include<sys / time.h>

[0071] double timer()

[0072] {

[0073] double t;

[0074] struct timeval tv;

[0075] gettimeofday(&tv,NULL);

[0076] t=(double)tv.tv_sec*1.0+(double)tv.tv_usec*1e-6;

[0077] return t;

[0078] }.

[0079] Step 3: Extract the sparse matrix-vector multiplication operation of the hotspot function, parallelize the calculation of the sparse matrix-vector multiplication operation on the master and slave cores, migrate the sparse matrix-vector multiplication operation to the slave core calculation, and use the heterogeneous parallel sparse approximate inverse algorithm to accelerate the calculation of the hotspot;

[0080] Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

[0081] Specifically, by analyzing the performance hotspots of the ParaSails preprocessor, the present disclosure finds that the MatrixMatvec() and MatrixMatvecTrans() functions are the main performance bottlenecks. These two functions perform standard matrix-vector multiplication and transposed matrix-vector multiplication operations, respectively. Since the two functions are similar in function, the specific heterogeneous parallel algorithm design is implemented in the present disclosure using the MatrixMatvec() function as an example to implement the heterogeneous parallel sparse approximate inverse algorithm swParaSails as follows:

[0082] First, the MatrixMatvec() function is used to perform a sparse matrix-vector multiplication (SPMV) operation, where the sparse matrix is ​​in CSR compression format and implemented with three arrays: the val array stores all non-zero elements of the matrix, the pntr array stores row pointers, that is, the index position of each row of non-zero elements in val, and the indx array is used to store the column index of the non-zero elements. Although the CSR format improves memory utilization, its indirect and irregular memory access pattern leads to poor data reuse and low computational intensity.

[0083] Furthermore, the unique Athread acceleration thread library of the Shenwei system is used to design slave core parallelization for computing hotspots. The master-slave core code is written using the athread function interface, and the master-slave acceleration model is used to accelerate the above computing hotspots. Among them, the master core program uses the athread_spawn() and athread_join() functions to accelerate the startup and recovery of thread tasks respectively. In particular, for the unique master-slave core architecture of the multi-core processor in the SW26010-Pro heterogeneous system, the computationally intensive part of the hot function MatrixMatvec(), that is, the sparse matrix-vector multiplication operation in the CSR (compressed sparse row) format, is migrated to the slave core for acceleration, thereby improving the sparse matrix-vector multiplication calculation speed. The sparse matrix-vector multiplication operation is parallelized between the master and slave cores, and the heterogeneous parallel sparse approximate inverse algorithm is used to accelerate the calculation of hot spots. The master-slave core parallelization task allocation strategy is as follows: the master core is used to read the non-zero element data of sparse matrices and vectors and allocate the corresponding memory; in the construction process of the sparse approximate inverse preprocessor ParaSails, the master core focuses on the selection and optimization of the non-zero element positions (sparse patterns) of the sparse matrix; calls the slave core resources, manages and coordinates the CPE (computational processing unit) to perform computing tasks; and collects and reduces the partial results of asynchronous calculations of each slave core. The slave core is responsible for executing intensive numerical computing tasks in the construction of the ParaSails preprocessor, such as matrix-vector multiplication, sparse matrix construction, iterative optimization, etc. In the parallel computing process, the slave core communicates with the master core using DMA (direct memory access), during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

[0084] Furthermore, by kernelizing the MatrixMatvec() and MatrixMatvecTrans() functions, a heterogeneous parallel sparse approximate inverse algorithm swParaSails for the Shenwei architecture proposed in the present invention can be implemented. The heterogeneous parallel sparse approximate inverse algorithm swParaSails introduced in the present invention introduces a variety of fine-grained performance optimization methods, including: adaptive data mapping, kernel aggregation optimization, and data communication optimization based on mixed precision.

[0085] As an embodiment, the adaptive data mapping optimization strategy in the heterogeneous parallel sparse approximate inverse algorithm swParaSails is as follows:

[0086] In view of the fact that the slave core LDM capacity of the SW26010-Pro processor is limited to 256KB, when the matrix scale is large, even if the sparse matrix M in the preprocessor is divided into 64 parts, the entire matrix cannot be loaded into the LDM of a single slave core at one time. Therefore, the present disclosure proposes an efficient adaptive data mapping optimization strategy. Considering that the number of non-zero elements in each row and column of the sparse matrix is ​​different and unevenly distributed, it is impossible to simply assign a fixed number of rows to the LDM of each slave core. The present disclosure dynamically transmits the number of rows of the block matrix according to the different scales of the input sparse matrix, and sets a dynamic data buffer in the LDM of each slave core. The buffer uses the number of non-zero elements as the loading condition, and only loads part of the data into the buffer for processing each time.

[0087] First, define the storage space size of LDM shared variables as S ldm KB, the remaining space available for storing sparse matrix data is remaining_ldm_space KB, the sparse matrix is ​​N rows, the maximum number of non-zero elements is max_nnz, the sparsity of each row of the sparse matrix is ​​m, the number of slave cores is Thread_num, and the data segment is divided into L rows. The implementation of the adaptive data mapping optimization strategy is specifically divided into the following steps:

[0088] (1) In the SpMV calculation process in the hotspot function MatrixMatvec(), the memory access to the input vector x is indirect and the access behavior is very irregular, resulting in a lot of discontinuity and randomness in its memory access. The slave core LDM space under the Shenwei architecture can be divided into private segments and shared segments. The shared segment provides a shared, high-speed memory area that multiple slave cores can access and use simultaneously without copying data to their respective local storage. By dividing the vector x into the shared segment of the LDM, the reuse of the data in the vector x during the SpMV calculation process can be improved.

[0089] The present disclosure dynamically configures the shared variable storage space S according to the size (N) of a given sparse matrix. ldm The length of the LDM shared segment must be a power of 2, that is, the size of each continuous shared segment in the slave core can only be 0, 4, 8, 16, 32, 64, or 128 KB. For sparse matrix-vector multiplication, the length of vector x is equal to the number of rows N of the sparse matrix, so the minimum memory space S for storing vector x can be obtained based on N. ldm , as shown in the following formula:

[0090]

[0091] Wherein, data_type is the data type of vector x. In the example disclosed in the present disclosure, the data type of vector x is double, that is, sizeof(double)=8, Thread_num=64. Thus, the corresponding size (N rows) matrix corresponding to the S to be configured can be adaptively determined. ldm Thread_num in a core group are each split out of the LDM of the core ldm KB of storage space, forming a Thread_num*S ldm KB of shared space is used to store vector x. Then, according to the following algorithm 1, different S configurations can be obtained according to different sparse matrix sizes (N rows). ldm Size. The specific algorithm 1 is:

[0092]

[0093] (2) Remaining_ldm_space (i.e., the remaining space available for storing sparse matrix data) is 256-S ldm KB.

[0094] (3) Each non-zero element of a double-precision floating point number occupies 8 bytes, so each 1KB LDM can store 128 non-zero elements. Therefore, the maximum number of non-zero elements max_nnz = (256-S ldm )×128, if the sparsity of each row is m (i.e. the average number of non-zero elements in each row), then the number of data rows L processed in each batch is:

[0095]

[0096] The adaptive data mapping optimization strategy can determine the optimal number of matrix rows L for each transmission of matrices of different sizes, divide the sparse matrix M by rows, and evenly divide it to each slave core according to the number of slave cores Thread_num in the core group started by the master core. Each time, data of size L rows is transmitted from the main memory, and each slave core obtains a different data segment L for SpMV calculation. The vector x is placed in the shared LDM of the slave core so that all slave cores can access it. After completing the calculation of each batch of data, the slave core will send the value of the result vector y of each part back to the master core. The transmission, calculation, and writing are repeated again. After several cycles, the slave cores asynchronously calculate all SpMVs in parallel. When there are remaining rest row data segments that have not been divided, they can be assigned to the slave cores with the front slave core number for calculation.

[0097] As an embodiment, the kernel aggregation optimization strategy in the heterogeneous parallel sparse approximate inverse algorithm swParaSails is as follows:

[0098] The hot functions of the sparse approximate inverse preprocessor are MatrixMatvec() and MatrixMatvecTrans(). Both are sparse matrix-vector multiplication operations in terms of function. The function execution process is very similar, and the main difference is whether the sparse matrix needs to be transposed. When these two hot functions are de-coreized, it is necessary to call the de-core startup function twice, namely the athread_spawn() function. In addition, these two hot functions are executed in the iterative algorithm, and the number of function executions is of the same order of magnitude as the number of iterations. For example, when solving a sparse linear system problem requires 10,000 iterations, the de-core startup function athread_spawn() needs to be executed 20,000 times. Although a single function call takes less time, when accumulated together, the function call will occupy a certain proportion of time overhead. Therefore, the present disclosure proposes a kernel aggregation strategy. Specifically, the MatrixMatvec() and MatrixMatvecTrans() functions are encapsulated into one function. The function body includes two calculation-intensive parts in the MatrixMatvec() and MatrixMatvecTrans() functions. A matrix transposition flag is introduced as a formal parameter as the function prototype of the two calculation modules to determine whether the matrix is ​​transposed, and one of the calculation modules in the function body is executed according to the flag.

[0099] Through the proposed kernel aggregation optimization strategy, 50% of the calls to the athread_spawn() function can be saved, reducing the time overhead of starting threads and further improving time efficiency.

[0100] As an embodiment, the mixed-precision data communication optimization strategy in the heterogeneous parallel sparse approximate inverse algorithm swParaSails is as follows:

[0101] For the parallel sparse approximate inverse algorithm swParaSails proposed in the present disclosure for the Shenwei architecture, in addition to the most important hot function swMatrixMatvec(), the sparse approximate inverse algorithm process will frequently execute kernel operations such as vector dot products. It is relatively simple to perform slave kernelization on this type of kernel. It only needs to be segmented and assigned to different slave cores. After each slave core calculates its own vector segment, it needs to be reduced to a slave core and then transferred to the main memory. Obviously, this process will generate data communication. Double precision is often used as the default precision choice in most scientific computing projects due to its superior calculation accuracy and numerical stability. Therefore, the original SPAI algorithm is implemented with double precision. However, in general, the results of double precision calculations often exceed the accuracy requirements of many engineering applications. Therefore, the accuracy can be relaxed in exchange for certain performance. For example, the single-precision floating point format has a smaller data bit width than double precision, so single precision can move at a higher rate in memory, reducing the amount of data transmitted, thereby reducing the amount of communication.

[0102] The present disclosure attempts to use single precision to represent the data to be transmitted, that is, the result of the inner product of each slave core vector segment, thereby reducing the communication overhead in the swParaSails algorithm. Specifically, since this data is stored in the LDM of the slave core, it is only necessary to introduce a data format conversion operation between the calculation of the vector segment inner product in the slave core function and the transmission to the main memory to convert the double precision into single precision.

[0103] Although the data communication optimization method based on mixed precision can reduce the communication volume, the effective number of bits and the representation range of the transmitted data after the precision conversion will be affected, which may cause the residual error to fail to decrease during the iteration process, thereby causing the algorithm to converge. Therefore, the present disclosure provides a residual error and performance balance algorithm to solve the convergence problem. Specifically, according to the number of iterations iter set by the iterative method, the average number of iterations per iteration is After the error is checked, check whether the residual is lower than the previous one. If the residual is not lowered, the mixed precision strategy is prohibited. This can improve program performance while ensuring the convergence of the algorithm.

[0104] As an example, Figure 3 As shown in the figure, the swParaSails preprocessor proposed in the present disclosure solves the linear system in the conjugate gradient (CG) iteration method, and is then applied to solve the electromagnetic scattering problem. The specific calculation process includes: firstly, inputting the coefficient matrix A of the linear system, the right-hand side term b and the initial guess solution x 0, and configure the relevant parameters of the preprocessing and conjugate gradient algorithm. Then, a sparse approximate inverse preprocessor M is constructed by the swParaSails method of the present invention, and the corresponding sparse approximate inverse calculation is performed. Subsequently, the preprocessor M is applied to preprocess the original linear equations, and the conjugate gradient algorithm is initialized, including calculating the initial residual and the initial search direction. In the iterative solution process, the step size is calculated for each iteration and the solution vector and residual are updated accordingly, and then whether to terminate the iteration is determined according to the convergence condition; if the convergence criterion is not met, the conjugate direction parameters are calculated and the search direction is updated, and the next round of iteration is continued until the convergence condition is met, and finally the solution x is output to complete the solution of the linear system.

[0105] Example 2

[0106] In one embodiment of the present disclosure, a sparse approximate inverse Shenwei parallel preprocessing system for electromagnetic scattering is provided, comprising:

[0107] A hotspot function determination module is used to perform a performance hotspot analysis on a sparse approximate inverse preprocessor in an electromagnetic scattering simulation process and determine a hotspot function of the sparse approximate inverse preprocessor;

[0108] The accelerated parallel computing module is used to extract the sparse matrix-vector multiplication operations of hotspot functions, parallelize the sparse matrix-vector multiplication operations on the master and slave cores, migrate the sparse matrix-vector multiplication operations to the slave cores, and accelerate the calculation of hotspots using the heterogeneous parallel sparse approximate inverse algorithm;

[0109] Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

[0110] Example 3

[0111] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering.

[0112] Example 4

[0113] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering is implemented.

[0114] Example 5

[0115] In one embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering.

[0116] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0118] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering, characterized in that: include: Perform performance hotspot analysis on the sparse approximate inverse preprocessor in the electromagnetic scattering simulation process and determine the hotspot function of the sparse approximate inverse preprocessor; Extract the sparse matrix-vector multiplication operations of the hotspot functions, parallelize the calculation of the sparse matrix-vector multiplication operations on the master and slave cores, migrate the sparse matrix-vector multiplication operations to the slave cores, and use the heterogeneous parallel sparse approximate inverse algorithm to accelerate the calculation of hotspots; Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

2. The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering according to claim 1, characterized in that: Use job-level performance detection tools and fine-grained instrumentation technology to perform performance hotspot analysis. Use the job-level performance monitoring and analysis tool SWPROF configured with the Shenwei architecture to determine the hotspot functions of the preprocessor. In the Shenwei system environment, add specific parameters to the script for submitting the job to obtain performance hotspot function data.

3. The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering according to claim 1, characterized in that: Extract the sparse matrix-vector multiplication operation of the hot function and migrate the computationally intensive part of the hot function, that is, the sparse matrix-vector multiplication operation in the CSR format, to the slave core for acceleration. The CSR compression format is implemented with three arrays: the val array stores all non-zero elements of the matrix, the pntr array stores row pointers, that is, the index position of each row of non-zero elements in val, and the indx array is used to store the column index of the non-zero elements.

4. The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering according to claim 1, characterized in that: The sparse matrix-vector multiplication operation is parallelized between the master and slave cores, and the heterogeneous parallel sparse approximate inverse algorithm is used to accelerate the calculation of hot spots. The master-slave core parallelization task allocation strategy is as follows: the master core reads the non-zero element data of the sparse matrix and vector, and allocates the corresponding memory; During the construction of the sparse approximate inverse preprocessor, the master core focuses on the selection and optimization of the positions of non-zero elements in the sparse matrix, calls the slave core resources, manages and coordinates the computing processing units to execute computing tasks, and collects and reduces partial results of asynchronous calculations of each slave core; the slave core is responsible for executing intensive numerical computing tasks in the construction of the sparse approximate inverse preprocessor. During the parallel computing process, the slave core communicates with the master core using direct memory access. During this period, the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

5. The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering according to claim 1, characterized in that: The adaptive data mapping optimization process is as follows: dynamically transmit the number of rows of the block matrix according to the different scales of the input sparse matrix, set a dynamic data buffer in the LDM of each slave core, and use the number of non-zero elements as the loading condition of the dynamic data buffer. Only part of the data is loaded into the buffer for processing each time.

6. The sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering according to claim 1, characterized in that: In the data communication optimization process based on mixed precision, single precision is used to represent the results of the inner products of each slave kernel vector segment. It is only necessary to introduce a data format conversion operation between the calculation of the vector segment inner product in the slave kernel function and the transmission to the main memory to convert the double precision into single precision, and set a residual and performance balance algorithm. According to the number of iterations set by the iterative method, check once within a set time interval whether the residual is reduced compared with the previous time. When the residual does not decrease, the data communication optimization process based on mixed precision is prohibited, so as to ensure the convergence of the algorithm.

7. A sparse approximate inverse Shenwei parallel preprocessing system for electromagnetic scattering, characterized in that: include: A hotspot function determination module is used to perform a performance hotspot analysis on a sparse approximate inverse preprocessor in an electromagnetic scattering simulation process and determine a hotspot function of the sparse approximate inverse preprocessor; The accelerated parallel computing module is used to extract the sparse matrix-vector multiplication operations of hotspot functions, parallelize the sparse matrix-vector multiplication operations on the master and slave cores, migrate the sparse matrix-vector multiplication operations to the slave cores, and accelerate the calculation of hotspots using the heterogeneous parallel sparse approximate inverse algorithm; Among them, various fine-grained performance optimization processes such as adaptive data mapping optimization, kernel aggregation optimization and data communication optimization based on mixed precision are introduced in the heterogeneous parallel sparse approximate inverse algorithm to realize the parallel computing of master-slave core-intensive numerical computing tasks; during the parallel computing process, the slave core communicates with the master core using direct memory access, during which the master core is in a waiting state to ensure data consistency and communication synchronization until all slave cores complete their assigned tasks.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the sparse approximate inverse parallel preprocessing method for electromagnetic scattering described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the sparse approximate inverse Shenwei parallel preprocessing method for electromagnetic scattering as described in any one of claims 1 to 6 is implemented.

10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the sparse approximate inverse parallel preprocessing method for electromagnetic scattering as described in any one of claims 1-6.

Citation Information

Cited By

  • CNN-oriented batch matrix multiplication parallel optimization method and system on SW architecture

    CN120508740A

  • Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

    CN121278223A

  • Ocean multi-physical field coupling method, medium and system

    CN122088213A