Parallel Optimization Method and System for Seismic Wave Simulation Algorithm Based on Shenwei Architecture

The proposed data shuffling and vectorization strategy for seismic wave simulation on Sunway architecture addresses inefficiencies in data access and computation, enhancing performance and bandwidth by optimizing data distribution and computation.

CN115390922BActive Publication Date: 2025-07-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210842656.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2025-07-15
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

The existing seismic wave simulation algorithm based on Shenwei architecture has problems such as discontinuous access from the kernel from the main memory and data multiplexing of Halo area data, complex DMA operations, large overhead for computing communication overlap, and inapplicable naive vectorization strategies, resulting in low computing efficiency.

Method used

The RMA+DMA collaborative memory access optimization method is adopted to shuffle the main memory data from the kernel level, and convert the variables from the kernel into vectorization units through vectorization strategies to optimize the data reading and calculation process of the seismic wave simulation algorithm.

Benefits of technology

The amount of data accessed during seismic wave simulation is reduced, the DMA bandwidth is improved, the computing efficiency and performance is improved, and the bandwidth is increased from 21GB/s to 27GB/s.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115390922B_ABST
    Figure CN115390922B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of algorithm parallel optimization, and provides a method and system for parallel optimization of seismic wave simulation algorithms based on the Shenwei architecture, including: after reading the data in the continuous area in the main memory into the slave cores, performing data shuffling at the slave core level; after converting the variables for seismic wave simulation in each grid point in the slave cores into a number of vectorization units, performing seismic wave simulation. It can not only reduce the amount of memory access data in the seismic wave simulation process, but also improve the bandwidth of DMA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of algorithm parallel optimization, and particularly relates to a method and system for parallel optimization of seismic wave simulation algorithms based on the Shenwei architecture. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] The propagation form of natural earthquakes inside the earth is seismic waves, and scientists mainly use the acoustic wave equation or the elastic wave equation to describe the propagation of seismic waves in the study of natural earthquake simulation. Simulating natural earthquakes not only requires accurately simulating the processes of seismic wave generation, diffusion, and contact boundaries, but also requires discretizing the wave equation in space and time when implementing the seismic simulation algorithm on a computer, which is a very complex process.

[0004] Efficient algorithms and sufficient computing resources are the core technologies of the earthquake rapid response system. The traditional Finite Difference Method (FDM) has the characteristics of high precision, high efficiency, easy programming, and easy parallelization. However, the traditional FDM cannot flexibly divide grids, especially for complex undulating terrains. To address the above drawbacks, based on the Curved Grid-FDM (CG-FDM), the traction mirror method was proposed and used to solve two-dimensional and three-dimensional problems of the elastic wave equation. This algorithm inherits the advantages of the traditional FDM algorithm, such as high precision, high efficiency, easy programming, and easy parallelization. In addition, CG-FDM is more flexible than the traditional FDM, and it can provide flexible grid meshing according to the terrain.

[0005] High-resolution seismic simulation requires large-scale parallel computing, which places higher demands on computing resources. The Shenwei supercomputer has powerful parallel computing capabilities. By reasonably dividing computing tasks through parallel programming techniques, fast and efficient seismic simulation can be achieved. In particular, the Shenwei supercomputer can complete parallel computing of tens of millions of cores, which provides a strong guarantee for high-resolution and high-frequency seismic rapid simulation. The calculated seismic intensity is also more accurate, and it can timely provide more reliable disaster assessment results for relevant emergency departments.

[0006] The Shenwei supercomputer uses the Shenwei processor, each of which is equipped with 6 core groups. Each core group contains a master core and 64 slave cores. In terms of storage, the master core storage hierarchy includes registers, data cache, instruction cache, and the second-level cache shared by data and instructions, main memory, and many other levels. Slightly different from the master core, the slave core mainly includes registers, on-chip data storage space, instruction cache, main memory, and other storage levels. Among them, the on-chip data storage space (LDM) has the characteristics of fast access speed, low latency, and high bandwidth. How to give full play to the performance of LDM is the key issue to effectively exert the computing power of the slave core. In a single slave core array, each slave core provides a random multiple access (Random Multiple Access, RMA) communication mechanism to exchange data with each other. Compared with direct memory access (Direct Memory Access, DMA) operations, RMA operations have smaller latency and higher bandwidth, and are suitable for transmitting data between adjacent slave cores to reduce DMA operations. The new Shenwei processor supports a 512-bit vectorized processing unit, reaching an advanced level. Vectorized operations, to put it simply, mean that the computing core can use one hardware instruction but can run multiple scalar operations at the same time. Vectorized operations can read data in batches and perform one-time operations, which can not only improve the memory access efficiency of data loaded into registers by LDM, but also improve computing efficiency.

[0007] However, the existing seismic wave simulation algorithm based on the Shenwei architecture has the following problems:

[0008] (1) When performing finite differences, each slave core must access discontinuous data blocks from the main memory to its own LDM space;

[0009] (2) According to the calculation characteristics of the stencil algorithm, each slave core must also obtain the Halo area adjacent to the data block; the Halo area between adjacent blocks is reused by multiple slave cores; therefore, when DMA fetches data, some data will be read repeatedly by multiple slave cores;

[0010] (3) When memory optimization reaches a certain level, the computational communication overlap within the core group cannot completely cover up the computational overhead, so it is necessary to improve computational efficiency;

[0011] (4) Naive vectorization strategies are not applicable in some scenarios and it is difficult to improve performance. Summary of the invention

[0012] To solve the technical problems existing in the above-mentioned background art, the present invention provides a parallel optimization method and system for seismic wave simulation algorithms based on the Shenwei architecture, which can not only reduce the amount of memory access data during the seismic wave simulation, but also improve the bandwidth of DMA.

[0013] To achieve the above object, the present invention adopts the following technical solutions:

[0014] The first aspect of the present invention provides a parallel optimization method for seismic wave simulation algorithms based on the Shenwei architecture, which includes:

[0015] After reading the data in the continuous area of the main memory into the slave cores, perform data shuffling at the slave core level;

[0016] After converting the variables for seismic wave simulation in each grid point within the slave core into a number of vectorization units, perform seismic wave simulation.

[0017] Further, each slave core reads the long column data block of its own slave core group from the main memory through the DMA method;

[0018] The long column data block includes a left Halo area, a right Halo area, and the target calculation area of each slave core in the slave core group where the slave core is located.

[0019] Further, the specific method of the data shuffling at the slave core level is as follows:

[0020] Each slave core within a slave core group expands the long column data block into a number of short column data blocks, and the short column data block includes a left Halo area, a right Halo area, and a calculation area;

[0021] Distribute each short column data block to the slave cores within the slave core group through RMA communication.

[0022] Further, the short column data block received by each slave core is a short column data block containing the target calculation area of the slave core.

[0023] Further, the number of short column data blocks obtained after expanding the long column data block is the same as the number of slave cores included in the slave core group.

[0024] Further, adopt a simple variable fusion vectorization strategy to convert the variables for seismic wave simulation in each grid point within the slave core into a number of vectorization units.

[0025] Further, adopt a hybrid vectorization strategy to convert the variables for seismic wave simulation in each grid point within the slave core into a number of vectorization units.

[0026] The second aspect of the present invention provides a parallel optimization system for seismic wave simulation algorithms based on the Shenwei architecture, which includes:

[0027] A cooperative memory access module, which is configured to: after reading data in a continuous area in the main memory into the slave core, perform data shuffling at the slave core level;

[0028] A vectorization module, which is configured to: after converting variables for seismic wave simulation in each grid point in the slave core into a number of vectorization units, perform seismic wave simulation.

[0029] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture as described above are implemented.

[0030] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture as described above are implemented.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] The present invention provides a parallel optimization method for a seismic wave simulation algorithm based on the Shenwei architecture, which increases the size of each data block read by the slave core from a short column to a long column, and the data structure remains the same as the original. It can not only reduce the amount of memory access data in the seismic wave simulation process, make full use of the low-latency characteristics of RMA communication, but also improve the bandwidth of DMA. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0034] Figure 1 is a schematic diagram of an eighth-order cooperative access model according to Embodiment 1 of the present invention;

[0035] Figure 2 is a schematic diagram of a naive variable fusion vectorization strategy according to Embodiment 1 of the present invention;

[0036] Figure 3 is a schematic diagram of a first hybrid vectorization strategy according to Embodiment 1 of the present invention;

[0037] Figure 4 is a schematic diagram of a second hybrid vectorization strategy according to Embodiment 1 of the present invention;

[0038] Figure 5 is a bandwidth data comparison diagram according to Embodiment 1 of the present invention;

[0039] Figure 6 It is a comparison chart of speedup ratio and bandwidth efficiency in the first embodiment of the present invention. Specific implementation manners

[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0041] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0042] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or their combinations.

[0043] The first embodiment

[0044] This embodiment provides a parallel optimization method for seismic wave simulation algorithms based on the Shenwei architecture, specifically including the following steps:

[0045] Step 1: Obtain the main core version.

[0046] After obtaining the main core version, multi-threaded parallelization of slave cores, vectorized calculation, and calculation memory access masking need to be performed in sequence.

[0047] The main core version is a multi-process version that does not call slave cores and is a benchmark version; multi-threaded parallelization of slave cores is to perform thread-level parallelization on the basis of the main core version, call the slave core array, and implement a simple slave core parallel version; vectorized calculation and calculation memory access masking are to implement vectorized calculation within the slave cores on the basis of the previous version, improve the calculation speed of the slave cores, and perform DMA memory access and calculation simultaneously to mask a part of the time.

[0048] Step 2: RMA+DMA collaborative memory access optimization (collaborative memory access based on RMA thread-level communication). After reading the data in the continuous area of the main memory into the slave cores, data shuffling at the slave core level is performed.

[0049] The RMA+DMA collaborative memory access optimization method divides the operation of pure DMA into two steps: The first step is to first read the data in the continuous area of the memory into the local data memory (LDM) of the slave cores. Since the data mapping relationship at this time cannot correspond to the algorithm calculation process, to solve this problem, the second step uses the RMA method for inter-thread communication to achieve data exchange and shuffling at the slave core level, so that each slave core obtains the correct data arrangement. Since this implementation method avoids repeated memory access in the Halo area and increases the continuity of memory access, although this two-step scheme is slightly more complex to implement than the pure DMA method, the DMA memory access method is more friendly to the memory.

[0050] If the slave core array is logically arranged in an n×n matrix, then each row naturally forms a slave core group. As Figure 1 shown in the schematic diagram of the octal collaborative access model, the so-called octal collaborative access model means that eight slave cores form a slave core group; if the slave core array is logically arranged in an 8×8 matrix, then each row naturally forms a slave core group.

[0051] In seismic simulation, after multi-level grid division, the data block assigned to each slave core is a three-dimensional grid data block composed of multiple grid points. The target data of the slave cores in a slave core group have adjacent position relationships in the main memory. In the octal collaborative access model, in the DMA data reading stage, each slave core of the slave core array reads the long column data block of its own slave core group from the main memory through the DMA method. Each slave core target data is a short column composed of a calculation area and two Halo areas. The Halo areas between the slave core groups in the memory overlap with each other. Assuming that the length of the calculation area (the target calculation area of each slave core) to be fetched by each slave core is Nc, and the lengths of the two Halo areas (i.e., the length of the left Halo area + the length of the right Halo area) are Nh, then the length of the short column is Nc+Nh. In the collaborative memory access optimization method, each slave core will read a long column data block with a length of Nc*8+Nh in one DMA. This long column contains the short column data of eight slave cores in a slave core group. That is, the long column data block includes the left Halo area, the right Halo area, and the target calculation area of each slave core in the slave core group where the slave core is located.

[0052] In Figure 1 , the lengths of each left Halo area in x, y, and z are respectively represented as halo lx , halo ly and halo lz ; the lengths of each right Halo area in x, y, and z are respectively represented as halo rx , halo ry and halo rz ; the lengths of each calculation area in x, y, and z are respectively represented as wx, wy, and wz.

[0053] After the DMA data reading is completed, the data shuffling stage is entered. For each slave core within a slave core group, data needs to be distributed to all slave cores within the slave core group.

[0054] The specific method of data shuffling at the slave core level is as follows: Each slave core within a slave core group expands a long column data block into several short column data blocks. The short column data block includes a left Halo region, a right Halo region, and a calculation region; each short column data block is distributed to the slave cores within the slave core group through RMA communication, and the short column data block received by each slave core is the short column data block containing the target calculation region of the slave core. The number of short column data blocks obtained after expanding the long column data block is the same as the number of slave cores contained in the slave core group. The short column data blocks obtained after expanding the long column data block are the same as the corresponding slave core data blocks within the slave core group.

[0055] Taking the slave cores in the first row of the slave core array as an example, it is naturally a slave core group. This slave core group includes 8 slave cores from core 0 to core 7. For each long column read by each slave core during the DMA data reading stage, a part of it needs to be copied, and the long column is expanded into eight short column data blocks. Taking the slave core group composed of the first row as an example, the expanded long column from left to right is sequentially the short column data of core 0 to core 7. Each short column data block includes a left Halo region, a calculation region, and a right Halo region, where the left and right Halo regions overlap with the calculation regions of adjacent slave cores. When expanding, the long column starts from the left, and data is copied in the order of the left Halo region, the calculation region, and the right Halo region to form a short column. This process is repeated until 8 short columns are copied. Because there are overlapping parts in the short columns, adjacent data will be copied repeatedly. Each slave core needs to perform data distribution operations with all other slave cores within its slave core group through RMA communication. Taking slave core 0 as an example, for a row of long column data it reads, it is copied and expanded into eight data blocks. The first data block (including two Halo regions and a calculation region) is left for itself, and at the same time it needs to distribute its corresponding data to slave cores 1 to 7 through RMA communication respectively. The second data block is distributed to slave core 1, the third data block is distributed to slave core 2, and so on, and the eighth data block is distributed to slave core 7 (similarly, the data distribution process of slave cores 1 to 7 is the same as that of slave core 0, which is omitted here).

[0056] The last stage of the eight - order cooperative access model is also that each slave core in the second to eighth rows of the slave core array respectively obtains the data of its adjacent upper - row slave cores through RMA communication.

[0057] Thus, through the above - mentioned eight - order cooperative access model, each slave core has obtained the data block of the size it is theoretically allocated, and can place the data in the correct position, thereby obtaining data consistent with the target data. The data block size of each slave core in the LDM is the size of the target data.

[0058] In this embodiment, the size of each data block read by DMA each time is increased from a short column to a long column, while the data structure remains the same as the original. Compared with the DMA scheme, the cooperative access scheme using RMA not only makes full use of the low-latency characteristics of RMA communication, but also improves the bandwidth of DMA.

[0059] Step 3: Vectorize data in parallel. After converting the variables for seismic wave simulation in each grid point within the slave core into a number of vectorized units, seismic wave simulation is performed.

[0060] For the seismic wave simulation algorithm, two vectorization strategies are mainly summarized: the first is the naive variable fusion vectorization strategy; the other is the hybrid vectorization strategy.

[0061] As Figure 2 shown, take six stress components, namely σxx, σyy, σzz, σxy, σxz, σyz and three velocity components, namely vx, vy, vz as an example. Using the naive variable fusion vectorization strategy to integrate different variables into one variable, the variable arrangement within a single grid point can be expressed as vx, vy, vx, σxx, σyy, σzz, σxy, σxz, σyz. If the same floating-point operation is performed on different variables, then one vectorized unit consists of every four adjacent variables in a grid point. The computing core can load these four variables and only through one vector memory access load instruction, and use vectorized floating-point operation instructions to replace the same floating-point operations performed by these four variables. The advantage of this strategy is that it can not only be vectorized, but also improve the DMA memory bandwidth. However, every advantage has its disadvantage: since there are a total of 9 variables and 9 is not divisible by 4, the vectorized units obtained under the naive variable fusion vectorization strategy are not regular. The variables of the third vectorized unit are σxy, vx, vy, vz. This strategy will make the calculation logic cumbersome and the amount of calculation increase significantly, which also makes this vectorization strategy not universal.

[0062] The hybrid vectorization strategy includes two types: the first hybrid vectorization strategy and the second hybrid vectorization strategy.

[0063] As Figure 3The first hybrid vectorization strategy shown is for scenarios where the same operation is performed on all variables. Among them, the first hybrid vectorization strategy uses a variable fusion method based on memory alignment. A single variable W1 is composed of vx, vy, vz, σxx, σyy, σzz, σxy, σxz, σyz. At the same time, the remaining σyz is made into a separate variable. The first hybrid vectorization strategy divides the 8 variables of variable W1 into two vectorization units and performs vectorization operations on them using a variable-based vectorization strategy. Then the remaining variable σyz is operated according to the grid point-based vectorization strategy, and every four adjacent grid points can be regarded as a vectorization unit.

[0064] As Figure 4 Shown is the second hybrid vectorization strategy, which is for scenarios where different operations are performed between different variables. Usually, one operation is performed on six stress components, while another operation is performed on the other three velocity components. Different operations are taken for stress components and velocity components. For the vectorization unit composed of vx, vy, vz, σxx, velocity-related operations and stress-related operations are performed on it, and the vectorization result of the velocity-related operation is represented by A, and the vectorization result of the stress-related operation is represented by B. A and B respectively represent two vectors, namely [A0(0), A1(0), A2(0), A3(0)] and [B0(0), B1(0), B2(0), B3(0)]. To ensure the correctness of the result, a Vector Shuffle, that is, a vectorization shuffle operation, is performed on these two vectors, and finally the correct target vector, namely [A0(0), A1(0), A2(0), B3(0)], is obtained.

[0065] Based on the characteristics of the master-slave heterogeneous architecture of the new high-performance processor and the computational characteristics of the seismic wave simulation algorithm, a cooperative memory access method and a vectorization strategy based on RMA thread-level communication are proposed, and it is found that they can indeed bring great performance improvement to the seismic wave simulation algorithm.

[0066] Among them, the cooperative memory access based on RMA thread-level communication can bring relatively obvious performance improvement. In Figure 5 a specific case and the measured bandwidth data are given. By using RMA communication, in addition to being able to reduce the amount of memory access data, the overall bandwidth is also increased from 21 GB per second to 27 GB per second.

[0067] The performance improvement brought by this embodiment to the seismic wave simulation algorithm includes gradual acceleration and the improvement of the bandwidth ratio (bandwidth efficiency = actual bandwidth / pure memory access model bandwidth), which is shown in Figure 6

[0068] Embodiment 2

[0069] ​This embodiment provides a parallel optimization system for seismic wave simulation algorithms based on the Shenwei architecture, which specifically includes the following modules:

[0070] The cooperative memory access module is configured to: after reading the data in the continuous area of the main memory into the slave cores, perform data shuffling at the slave core level;

[0071] The vectorization module is configured to: after converting the variables for seismic wave simulation in each grid point within the slave cores into a number of vectorization units, perform seismic wave simulation.

[0072] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and their specific implementation processes are the same, so they will not be repeated here.

[0073] Embodiment 3

[0074] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the parallel optimization method for seismic wave simulation algorithms based on the Shenwei architecture as described in Embodiment 1 above.

[0075] Embodiment 4

[0076] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the parallel optimization method for seismic wave simulation algorithms based on the Shenwei architecture as described in Embodiment 1 above.

[0077] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0078] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.

[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 or more boxes.

[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 or more boxes.

[0081] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0082] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A parallel optimization method for seismic wave simulation algorithm based on the Shenwei architecture, characterized in that, Including: After reading the data in the continuous area of the main memory into the slave cores, perform data shuffling at the slave core level; wherein, each slave core reads the long column data block of its own slave core group from the main memory through DMA; the long column data block includes a left Halo area, a right Halo area, and the target calculation area of each slave core in the slave core group where the slave core is located; the specific method of data shuffling at the slave core level is: each slave core in a slave core group expands the long column data block into several short column data blocks, and the short column data block includes a left Halo area, a right Halo area, and a calculation area; each short column data block is distributed to the slave cores in the slave core group through RMA communication. After converting the variables for seismic wave simulation in each grid point in the slave core into a number of vectorization units, perform seismic wave simulation.

2. The parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture according to claim 1, characterized in that Each short column data block received by each slave core is a short column data block containing the target calculation area of the slave core.

3. The parallel optimization method for seismic wave simulation algorithm based on the Shenwei architecture according to claim 1, characterized in that The number of short column data blocks obtained after expanding the long column data block is the same as the number of slave cores included in the slave core group.

4. The parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture according to claim 1, characterized in that, Adopt a simple variable fusion vectorization strategy to convert the variables for seismic wave simulation in each grid point in the slave core into a number of vectorization units.

5. The parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture according to claim 1, wherein Adopt a hybrid vectorization strategy to convert the variables for seismic wave simulation in each grid point in the slave core into a number of vectorization units.

6. The parallel optimization system for seismic wave simulation algorithm based on the Shenwei architecture is characterized in that Including: A cooperative memory access module, which is configured to: after reading the data in the continuous area of the main memory into the slave cores, perform data shuffling at the slave core level; wherein, each slave core reads the long column data block of its own slave core group from the main memory through DMA; the long column data block includes a left Halo area, a right Halo area, and the target calculation area of each slave core in the slave core group where the slave core is located; the specific method of data shuffling at the slave core level is: each slave core in a slave core group expands the long column data block into several short column data blocks, and the short column data block includes a left Halo area, a right Halo area, and a calculation area; each short column data block is distributed to the slave cores in the slave core group through RMA communication. A vectorization module, which is configured to: after converting the variables for seismic wave simulation in each grid point in the slave core into a number of vectorization units, perform seismic wave simulation.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture as described in any one of claims 1-5.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the parallel optimization method of the seismic wave simulation algorithm based on the Shenwei architecture as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Gas kinetics algorithm optimization method based on Shenwei architecture

    CN111104765A

  • Small-scale symmetric matrix parallel tridiagonalization method for SW many-core processor

    CN113704691A