Graphic processor-based discontinuous finite element seismic numerical simulation method and device
By using the intermittent finite element method based on the graphics processor in the numerical simulation of earthquakes and configuring parallel computing parameters, the problem of large calculation volume and low simulation efficiency of the finite element method is solved, and more efficient seismic wave simulation calculation is achieved.
Patent Information
- Application Number
- CN202311676088.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
In numerical seismic simulation, the finite element method has a large amount of calculation, resulting in low simulation efficiency, especially when dealing with large-scale sparse matrices.
The intermittent finite element seismic numerical simulation method based on the graphics processor is adopted, and the data of each intermittent finite element unit is calculated in parallel loop by configuring the number of execution threads of the graphics processor, the shared memory of thread blocks, the parallel method and the shared memory address.
The computing efficiency of seismic wave simulation is improved, and the computing time is reduced by making full use of the parallel computing power of the graphics processor.
Smart Images

Figure CN120124338A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of seismic wave numerical simulation in seismic exploration, and more particularly, to a discontinuous finite element seismic numerical simulation method and apparatus based on a graphics processing unit. Background Art
[0002] Seismic numerical simulation is a reliable method for studying the propagation characteristics of seismic waves and analyzing geological structures. The methods of seismic numerical simulation are mainly divided into two categories: one is the ray-based method, and the other is the wave equation-based method. Compared with the ray-based method, the wave equation seismic numerical simulation method is more favored in the research of seismic response characteristic analysis because it can more detailedly examine the kinematic and dynamic information of seismic waves. The wave equation numerical simulation mainly includes the finite difference method, the finite element method, the pseudospectral method, etc. These methods all rely on the mesh discretization of the model. Among them, the finite element method has the characteristics of higher accuracy and stronger applicability of the method, and is also more widely used in seismology.
[0003] However, seismic numerical simulation has a large amount of calculation, especially for the finite element method. The main form of the finite element method is the calculation of matrices. Conventional finite element methods often need to solve equations composed of large-scale sparse matrices. At the same time, it is also necessary to invert such matrices. Some special finite element methods, such as the spectral element method and the discontinuous Galerkin finite element method, can construct a diagonal mass matrix by selecting appropriate basis functions and integration nodes, thus avoiding the inversion of large-scale sparse matrices. But it also faces a huge amount of calculation.
[0004] Therefore, how to efficiently perform seismic wave numerical simulation has always been a difficult problem that the academic community and the industrial level have been continuously tackling. The graphics processing unit is an effective method for accelerating calculations. Compared with parallel computing using a central processing unit, the graphics processing unit can have more concurrent computing units and faster computing efficiency. And the graphics processing unit acceleration method has been widely used in finite difference seismic numerical simulation. At present, since the grids of the finite element method are usually unstructured triangular or tetrahedral grids, it is more difficult to perform parallel loop calculations compared with differences, which in turn leads to a large amount of calculation in seismic numerical simulation and low simulation efficiency. Summary of the Invention
[0005] In view of this, the present invention discloses a discontinuous finite element seismic numerical simulation method based on a graphics processing unit, which can solve the problems of large calculation amount and low simulation efficiency in current seismic wave numerical simulation.
[0006] According to one aspect of the present invention, a discontinuous finite element seismic numerical simulation method based on a graphics processing unit is proposed. The method includes:
[0007] Step 1, obtaining the parameters and matrices of the discontinuous finite element according to the discontinuous finite element grid and the velocity model;
[0008] Step 2: Configure the number of execution threads for each graphics processing unit, the shared memory for each thread block, the parallelization method between threads, and the shared memory address for each thread according to the graphics processing unit device number corresponding to the central processing unit process;
[0009] Step 3: After simulating the seismic source at the preset position, calculate the data of each discontinuous finite element unit in parallel loop according to the number of execution threads for each graphics processing unit, the shared memory for each thread block, the parallelization method between threads, and the shared memory address for each thread;
[0010] Step 4: Store the intermediate results of the parallel loop calculation in the shared memory, and store the results saved in the shared memory at the end of each moment into the global memory as the final output results.
[0011] In some embodiments, the step of obtaining the parameters and matrices of the discontinuous finite element according to the discontinuous finite element mesh and the velocity model includes:
[0012] Obtain the parameters and matrices of the discontinuous finite element through finite element mesh reading, velocity model reading, seismic observation system reading, discontinuous finite element basis function calculation, and seismic source simulation parameter specification.
[0013] In some embodiments, the step of configuring the number of execution threads for each graphics processing unit includes:
[0014] Configure the number of thread blocks corresponding to each graphics processing unit and the number of threads in each thread block.
[0015] In some embodiments, the step of configuring the shared memory for each thread block of each graphics processing unit includes: Configure the shared memory for each thread block of each graphics processing unit according to the number of variables to be solved, the maximum number of degrees of freedom of the discontinuous finite element unit, the number of threads in the thread blocks corresponding to the graphics processing unit, and the number of bytes per data.
[0016] In some embodiments, the number of intermediate results is the same as the number of variables to be solved, and each intermediate result occupies a total size in the shared memory that is the product of the maximum number of degrees of freedom of the unit, the number of threads, and the number of bytes per data.
[0017] In some embodiments, the step of calculating the data of each discontinuous finite element unit in parallel loop according to the number of execution threads for each graphics processing unit, the shared memory for each thread block, the parallelization method between threads, and the shared memory address for each thread includes:
[0018] Calculate the data of each discontinuous finite element unit in parallel loop according to the number of execution threads of each graphics processor, the shared memory of each thread block, the thread identification information corresponding to different discontinuous finite element unit numbers, and the shared memory address of each thread.
[0019] In some embodiments, the step of storing the intermediate result of the parallel loop calculation into the shared memory includes:
[0020] Store the intermediate result of the parallel loop calculation into the shared memory in the order of discontinuous finite element unit first and degree of freedom second.
[0021] According to one aspect of the present invention, there is also provided a discontinuous finite element seismic numerical simulation device based on a graphics processor, the device includes:
[0022] An acquisition module, configured to acquire the parameters and matrices of the discontinuous finite element according to the discontinuous finite element grid and the velocity model;
[0023] A configuration module, configured to configure the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread according to the graphics processor device number corresponding to the central processing unit process;
[0024] A calculation module, configured to, after simulating a seismic source at a preset position, calculate the data of each discontinuous finite element unit in parallel loop according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread;
[0025] A storage module, configured to store the intermediate result of the parallel loop calculation into the shared memory, and store the result saved in the shared memory at the end of each final moment into the global memory as the final output result.
[0026] According to another aspect of the present invention, there is also provided an electronic device, the electronic device includes:
[0027] A memory, storing executable instructions;
[0028] A processor, the processor runs the executable instructions in the memory to implement the above-mentioned discontinuous finite element seismic numerical simulation method based on a graphics processor.
[0029] According to another aspect of the present invention, there is also provided a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned discontinuous finite element seismic numerical simulation method based on a graphics processor.
[0030] The technical solution has at least the following advantages: According to the graphics processor device number corresponding to the central processing unit process, the present invention configures the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread, so as to utilize the graphics processor to allocate the calculation tasks of each unit to each thread. After simulating the seismic source at the preset position, according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread, the data of each discontinuous finite element unit is calculated in parallel and circularly, thus making full use of the shared memory of the thread block, realizing the parallel circular calculation of discontinuous finite element seismic simulation, and improving the calculation efficiency of seismic wave simulation.
[0031] The method and device of the present invention have other characteristics and advantages, which will be obvious in the accompanying drawings and subsequent specific embodiments incorporated herein, or will be described in detail in the accompanying drawings and subsequent specific embodiments incorporated herein. These accompanying drawings and specific embodiments are jointly used to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0033] Figure 1 The flowchart of a discontinuous finite element seismic numerical simulation method based on a graphics processor according to an embodiment of the present invention is shown;
[0034] Figure 2 The schematic diagram of a discontinuous finite element seismic numerical simulation device based on a graphics processor according to another embodiment of the present invention is shown;
[0035] Figure 3 The schematic diagram of a discontinuous finite element seismic numerical simulation electronic device based on a graphics processor according to still another embodiment of the present invention is shown;
[0036] Figure 4 The schematic diagram of a discontinuous finite element seismic numerical simulation computer-readable storage medium based on a graphics processor according to yet another embodiment of the present invention is shown;
[0037] Figure 5 The schematic diagram of the shared memory address of a single intermediate variable provided by the embodiment of the present invention;
[0038] Figure 6 The schematic diagram of the simulation acceleration effect provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0040] The present invention provides a discontinuous finite element seismic numerical simulation method based on a graphics processing unit, and the method includes:
[0041] Step 1, obtaining the parameters and matrices of discontinuous finite elements according to the discontinuous finite element mesh and the velocity model;
[0042] Step 2, configuring the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread according to the graphics processing unit device number corresponding to the central processing unit process;
[0043] Step 3, after simulating the seismic source at a preset position, calculating the data of each discontinuous finite element unit in parallel loop according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread;
[0044] Step 4, storing the intermediate results of the parallel loop calculation in the shared memory, and storing the results saved in the shared memory at the end of each moment into the global memory as the final output result.
[0045] In some embodiments, the step of obtaining the parameters and matrices of discontinuous finite elements according to the discontinuous finite element mesh and the velocity model includes:
[0046] Obtaining the parameters and matrices of discontinuous finite elements through finite element mesh reading, velocity model reading, seismic observation system reading, discontinuous finite element basis function calculation, and seismic source simulation parameter specification.
[0047] In some embodiments, the step of configuring the number of execution threads of each graphics processing unit includes:
[0048] Configuring the number of thread blocks corresponding to each graphics processing unit and the number of threads in each thread block.
[0049] In some embodiments, the step of configuring the shared memory of each thread block of each graphics processing unit includes: configuring the shared memory of each thread block of each graphics processing unit according to the number of variables to be solved, the maximum number of degrees of freedom of discontinuous finite element units, the number of threads in the thread blocks corresponding to the graphics processing unit, and the number of bytes of a single data.
[0050] In some embodiments, the number of the intermediate results is the same as the number of variables to be solved, and each of the intermediate results occupies a total size in the shared memory that is the product of the maximum number of degrees of freedom per unit, the number of threads, and the number of bytes per single data item.
[0051] In some embodiments, the step of parallelly and circularly calculating data of each discontinuous finite element unit according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode among threads, and the shared memory address of each thread includes:
[0052] Parallelly and circularly calculating data of each discontinuous finite element unit according to the number of execution threads of each graphics processor, the shared memory of each thread block, the thread identification information corresponding to different discontinuous finite element unit numbers, and the shared memory address of each thread.
[0053] In some embodiments, the step of storing the intermediate results of the parallel circular calculation into the shared memory includes:
[0054] Storing the intermediate results of the parallel circular calculation into the shared memory in the order of the discontinuous finite element unit first and the degree of freedom second.
[0055] The present invention also provides a discontinuous finite element seismic numerical simulation device based on a graphics processor, and the device includes:
[0056] An acquisition module, configured to acquire parameters and matrices of discontinuous finite elements according to a discontinuous finite element grid and a velocity model;
[0057] A configuration module, configured to configure the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode among threads, and the shared memory address of each thread according to the graphics processor device number corresponding to a central processing unit process;
[0058] A calculation module, configured to, after simulating a seismic source at a preset position, parallelly and circularly calculate data of each discontinuous finite element unit according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode among threads, and the shared memory address of each thread;
[0059] A storage module, configured to store the intermediate results of the parallel circular calculation into the shared memory, and store the results saved in the shared memory at the end of each moment into the global memory as final output results.
[0060] The present invention also provides an electronic device, and the electronic device includes:
[0061] A memory, storing executable instructions;
[0062] A processor that runs the executable instructions in the memory to implement the GPU-based discontinuous finite element seismic numerical simulation method described above.
[0063] The present invention also provides a computer-readable storage medium storing a computer program which, when executed by a processor, implements the GPU-based discontinuous finite element seismic numerical simulation method described above.
[0064] Example 1
[0065] Figure 1 The flowchart of the GPU-based discontinuous finite element seismic numerical simulation method according to an embodiment of the present invention is shown. As shown, the method includes steps 1 to 4.
[0066] Step 1: Obtain the parameters and matrices of the discontinuous finite element according to the discontinuous finite element mesh and the velocity model.
[0067] Specifically, the parameters and matrices of the discontinuous finite element are obtained by finite element mesh reading, velocity model reading, seismic observation system reading, discontinuous finite element basis function calculation, and source simulation parameter specification.
[0068] For example, calculate according to the formula where U is the seismic wave field to be solved, and the other matrices are all prepared in the CPU process.
[0069] Step 2: Configure the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread according to the graphics processing unit device number corresponding to the central processing unit process.
[0070] Specifically, the step of configuring the number of execution threads of each graphics processing unit includes: configuring the number of thread blocks corresponding to each graphics processing unit and the number of threads in each thread block. The step of configuring the shared memory of each thread block of each graphics processing unit includes: configuring the shared memory of each thread block of each graphics processing unit according to the number of variables to be solved, the maximum number of degrees of freedom of the discontinuous finite element unit, the number of threads in the thread block corresponding to the graphics processing unit, and the number of bytes of a single data.
[0071] Wherein, the number of intermediate results is the same as the number of variables to be solved, and each intermediate result occupies a total size in the shared memory that is the product of the maximum number of degrees of freedom of the unit, the number of threads, and the number of bytes of a single data.
[0072] For example, first, the number of thread blocks used by the processor, gridDim, and the number of threads in each block, blockDim, are given; then, the size of the shared memory is given. The total size of the shared memory is given by: the number of variables to be solved * the maximum degrees of freedom per element * blockDim * the number of bytes per data item. The total size of this shared memory must not exceed the upper limit required by the hardware. It should be noted that during parallel loop calculations, each thread is individually responsible for one element, that is, the thread identification information is associated with the element number. After each element is calculated, the thread will start calculating the next element, and so on; finally, the address in the shared memory where the intermediate variables for the calculation are stored is specified. Among them, the number of intermediate variables is the same as the number of variables to be solved. Therefore, the overall size occupied by each intermediate variable in the shared memory is: the maximum degrees of freedom per element * blockDim * the number of bytes per data item.
[0073] Furthermore, to enable the threads in a warp that are concurrently executing to operate on a contiguous block of memory, the shared memory for each intermediate variable is sorted with the element first and the degrees of freedom second, that is, CacheID + idof * CacheGap. Where CacheID = threadID and CacheGap = blockDIM, as can be Figure 5 shown.
[0074] Step 3, after simulating the seismic source at the preset position, calculate the data of each discontinuous finite element cell in parallel loop according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread.
[0075] Specifically, calculate the data of each discontinuous finite element cell in parallel loop according to the number of execution threads of each graphics processor, the shared memory of each thread block, the thread identification information corresponding to different discontinuous finite element cell numbers, and the shared memory address of each thread.
[0076] For example, first, the shared memory is set to zero before each element is calculated. It should be noted that since concurrency is carried out by threads and each thread is responsible for one element, and each thread will calculate the next element after the calculation of the element is completed, the shared memory needs to be re-initialized before each calculation. Then, the seismic source is loaded. Seismic numerical simulation requires simulating the occurrence of the seismic source at an appropriate position.
[0077] Step 4, store the intermediate results of the parallel loop calculation in the shared memory, and store the results saved in the shared memory at the end of each moment as the final output result in the global memory.
[0078] Specifically, the step of storing the intermediate results of the parallel loop calculation into the shared memory includes: storing the intermediate results of the parallel loop calculation into the shared memory in the order of discontinuous finite element cells first and degrees of freedom second.
[0079] The method provided in Embodiment 1 of the present invention configures the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread according to the graphics processing unit device number corresponding to the central processing unit process. Thus, by using the graphics processing unit, the calculation task of each unit is assigned to each thread. After simulating the seismic source at the preset position, according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread, the data of each discontinuous finite element cell is calculated in parallel loop, thereby making full use of the shared memory of the thread block, realizing the parallel loop calculation of discontinuous finite element seismic simulation, and improving the calculation efficiency of seismic wave simulation.
[0080] Example 2
[0081] According to an embodiment of the present invention, a discontinuous finite element seismic numerical simulation device based on a graphics processing unit is provided, as Figure 2 shown. The device includes: an acquisition module 21, a configuration module 22, a calculation module 23, and a storage module 24;
[0082] The acquisition module 21 is used to acquire the parameters and matrices of discontinuous finite elements according to the discontinuous finite element grid and the velocity model;
[0083] The configuration module 22 is used to configure the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread according to the graphics processing unit device number corresponding to the central processing unit process;
[0084] The calculation module 23 is used to, after simulating the seismic source at the preset position, calculate the data of each discontinuous finite element cell in parallel loop according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread;
[0085] The storage module 24 is used to store the intermediate results of the parallel loop calculation into the shared memory, and store the results saved in the shared memory at the end of each moment into the global memory as the final output result.
[0086] The device provided in Embodiment 2 of the present invention configures the number of execution threads of each graphics processor, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread according to the graphics processor device number corresponding to the central processor process, so as to use the graphics processor to allocate the calculation tasks of each unit to each thread. After simulating the seismic source at the preset position, the data of each discontinuous finite element unit is calculated in parallel and circularly according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread, thus making full use of the shared memory of the thread block, realizing the parallel and circular calculation of discontinuous finite element seismic simulation, and improving the calculation efficiency of seismic wave simulation.
[0087] Example 3
[0088] According to another aspect of the present invention, an electronic device is further provided. As Figure 3 shown, the electronic device includes:
[0089] A memory 31 storing executable instructions:
[0090] A processor 32 that runs the executable instructions in the memory to implement the discontinuous finite element seismic numerical simulation method based on a graphics processor according to the present invention.
[0091] The method includes the following steps:
[0092] Step 1, obtaining the parameters and matrices of discontinuous finite elements according to the discontinuous finite element grid and the velocity model;
[0093] Step 2, configuring the number of execution threads of each graphics processor, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread according to the graphics processor device number corresponding to the central processor process;
[0094] Step 3, after simulating the seismic source at the preset position, calculating the data of each discontinuous finite element unit in parallel and circularly according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallelism mode between threads, and the shared memory address of each thread;
[0095] Step 4, storing the intermediate results of the parallel circular calculation in the shared memory, and storing the results saved in the shared memory at the end of each moment as the final output result in the global memory.
[0096] In some embodiments, the step of obtaining the parameters and matrices of discontinuous finite elements according to the discontinuous finite element grid and the velocity model includes:
[0097] Obtain the parameters and matrices of the discontinuous finite element by finite element mesh reading, velocity model reading, seismic observation system reading, discontinuous finite element basis function calculation, and source simulation parameter specification.
[0098] In some embodiments, the step of configuring the number of execution threads of each graphics processor includes:
[0099] Configure the number of thread blocks corresponding to each graphics processor and the number of threads in each thread block.
[0100] In some embodiments, the step of configuring the shared memory of each thread block of each graphics processor includes: configuring the shared memory of each thread block of each graphics processor according to the number of variables to be solved, the maximum number of degrees of freedom of the discontinuous finite element cell, the number of threads in the thread blocks corresponding to the graphics processor, and the number of bytes per data.
[0101] In some embodiments, the number of intermediate results is the same as the number of variables to be solved, and each intermediate result occupies a total size in the shared memory that is the product of the maximum number of degrees of freedom of the cell, the number of threads, and the number of bytes per data.
[0102] In some embodiments, the step of calculating the data of each discontinuous finite element cell in parallel loop according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread includes:
[0103] Calculate the data of each discontinuous finite element cell in parallel loop according to the number of execution threads of each graphics processor, the shared memory of each thread block, the thread identification information corresponding to different discontinuous finite element cell numbers, and the shared memory address of each thread.
[0104] In some embodiments, the step of storing the intermediate results of the parallel loop calculation into the shared memory includes:
[0105] Store the intermediate results of the parallel loop calculation into the shared memory in the order of discontinuous finite element cell first and degree of freedom second.
[0106] The electronic device provided in Embodiment 3 of the present invention configures the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread according to the graphics processor device number corresponding to the central processor process, so as to utilize the graphics processor to allocate the computing tasks of each unit to each thread. After simulating the seismic source at the preset position, according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread, the data of each discontinuous finite element unit is calculated in parallel and cyclically, thus making full use of the shared memory of the thread block, realizing the parallel and cyclic calculation of discontinuous finite element seismic simulation, and improving the calculation efficiency of seismic wave simulation.
[0107] Example 4
[0108] According to another aspect of the present invention, there is also provided a computer-readable storage medium, as Figure 4 shown. The computer-readable storage medium 41 stores a computer program, and when the computer program is executed by a processor, it implements the discontinuous finite element seismic numerical simulation method based on a graphics processor according to the present invention.
[0109] The method includes the following steps:
[0110] Step 1, obtaining the parameters and matrices of discontinuous finite elements according to the discontinuous finite element grid and the velocity model;
[0111] Step 2, configuring the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread according to the graphics processor device number corresponding to the central processor process;
[0112] Step 3, after simulating the seismic source at the preset position, calculating the data of each discontinuous finite element unit in parallel and cyclically according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread;
[0113] Step 4, storing the intermediate results of the parallel cyclic calculation into the shared memory, and storing the results saved in the shared memory at the end of each moment as the final output result into the global memory.
[0114] In some embodiments, the step of obtaining the parameters and matrices of discontinuous finite elements according to the discontinuous finite element grid and the velocity model includes:
[0115] Obtaining the parameters and matrices of discontinuous finite elements through finite element grid reading, velocity model reading, seismic observation system reading, discontinuous finite element basis function calculation, and seismic source simulation parameter specification.
[0116] In some embodiments, the step of configuring the number of execution threads of each graphics processor includes:
[0117] Configuring the number of thread blocks corresponding to each graphics processor and the number of threads in each thread block.
[0118] In some embodiments, the step of configuring the shared memory of each thread block of each graphics processor includes: configuring the shared memory of each thread block of each graphics processor according to the number of variables to be solved, the maximum degrees of freedom of discontinuous finite element cells, the number of threads in the thread blocks corresponding to the graphics processor, and the number of bytes per data.
[0119] In some embodiments, the number of intermediate results is the same as the number of variables to be solved, and each intermediate result occupies a total size in the shared memory that is the product of the maximum degrees of freedom of the unit, the number of threads, and the number of bytes per data.
[0120] In some embodiments, the step of parallel loop calculating the data of each discontinuous finite element cell according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread includes:
[0121] Parallel loop calculating the data of each discontinuous finite element cell according to the number of execution threads of each graphics processor, the shared memory of each thread block, the thread identification information corresponding to different discontinuous finite element cell numbers, and the shared memory address of each thread.
[0122] In some embodiments, the step of storing the intermediate results of the parallel loop calculation into the shared memory includes:
[0123] Storing the intermediate results of the parallel loop calculation into the shared memory in the order of discontinuous finite element cells first and degrees of freedom second.
[0124] The computer-readable storage medium provided in Embodiment 4 of the present invention configures the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread according to the graphics processor device number corresponding to the central processing unit process, so as to use the graphics processor to allocate the calculation task of each unit to each thread, and after simulating the seismic source at the preset position, parallel loop calculate the data of each discontinuous finite element cell according to the number of execution threads of each graphics processor, the shared memory of each thread block, the parallel mode between threads, and the shared memory address of each thread, thereby making full use of the shared memory of the thread blocks, realizing the parallel loop calculation of discontinuous finite element seismic simulation, and improving the calculation efficiency of seismic wave simulation.
[0125] Example 5
[0126] The present invention will be further described below in conjunction with the experimental test results. The experimental steps of the present invention include:
[0127] The first step is data preparation before finite element calculation, including but not limited to mesh reading, model reading, acquisition system reading, basis function calculation, simulation parameter specification, etc. The general form of the calculation equation for discontinuous finite elements is:
[0128]
[0129] Among them, U is the seismic wave field to be solved, and the rest of the matrices are all prepared in the CPU process.
[0130] The second step is to specify the number of execution threads for graphics processing. It is necessary to specify the number of thread blocks gridDim used by the processor and the number of threads blockDim in each block.
[0131] The third step is to specify the size of the shared memory. The total size of the shared memory is given as: the number of variables to be solved * the maximum number of degrees of freedom per element * blockDim * the number of bytes per single data. Here, the total size of the shared memory cannot exceed the upper limit required by the hardware. Therefore, attention needs to be paid to the setting of blockDim.
[0132] The fourth step is that each thread is responsible for a single element during parallel loop calculation. That is, the threadID is associated with the element number. After each element is calculated, the thread will start calculating the next element, and so on.
[0133] The fifth step is to specify the address in the shared memory where the intermediate variables of the calculation are stored. The number of intermediate variables is the same as the number of variables to be solved. Therefore, the overall size occupied by each intermediate variable in the shared memory is: the maximum number of degrees of freedom per element * blockDim * the number of bytes per single data. In order to enable the threads in a single warp that are concurrently executed to operate on a continuous block of memory. The shared memory of each intermediate variable is sorted with the element first and the degree of freedom second. That is, CacheID + idof * CacheGap. Where CacheID = threadID and CacheGap = blockDIM. For details, see Figure 5 .
[0134] The sixth step is to zero the shared memory before each element calculation. Since the concurrency is carried out according to threads, each thread is responsible for a single element, and after the calculation of each element, each thread will calculate the next element. Therefore, the shared memory needs to be re-initialized before each calculation.
[0135] The seventh step is source loading. Seismic numerical simulation needs to simulate the occurrence of the source at an appropriate location.
[0136] In the eighth step, the units perform concurrent calculations and the intermediate results are stored in the shared memory, whose address follows the provisions of the fifth step.
[0137] The ninth step is to reduce the intermediate results to the final results and complete the calculation at each moment.
[0138] Step 10. Time loops until the end.
[0139] It should be noted that this experiment provides Figure 5 The address allocation method for the shared memory used by each unit is shown in the figure. Each block represents a data, tmpvx represents the intermediate result of the wave field vx calculation, and idof represents the degree of freedom. Blocks of the same color represent all threads in a thread block. This allocation method allows each thread warp to operate on a whole block of continuous addresses, thus avoiding storage conflicts. Figure 6 The comparison of the calculation speedup ratio of the discontinuous finite element in models with different mesh numbers and the single core of the E5-2684 v4 processor is shown. The comparison results of three graphics cards are given here. This experiment fully demonstrates that the present invention makes full use of the shared memory of the graphics processor thread block and avoids storage conflicts as much as possible, thereby improving the efficiency of discontinuous finite element simulation.
[0140] For other detailed descriptions of this exemplary embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0141] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A discontinuous finite element seismic numerical simulation method based on a graphics processing unit, characterized in that, the method includes: Step 1, obtaining the parameters and matrices of discontinuous finite elements according to discontinuous finite element meshes and velocity models; Step 2, configuring the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread according to the graphics processing unit device number corresponding to the central processing unit process; Step 3, after simulating the seismic source at a preset position, calculating the data of each discontinuous finite element unit in parallel loop according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread; Step 4, storing the intermediate results of the parallel loop calculation into the shared memory, and storing the results saved in the shared memory at the end of each moment into the global memory as the final output result.
2. The discontinuous finite element seismic numerical simulation method based on a graphics processing unit according to claim 1, characterized in that, the step of obtaining the parameters and matrices of discontinuous finite elements according to discontinuous finite element meshes and velocity models includes: Obtaining the parameters and matrices of discontinuous finite elements through finite element mesh reading, velocity model reading, seismic observation system reading, discontinuous finite element basis function calculation, and seismic source simulation parameter specification.
3. The discontinuous finite element seismic numerical simulation method based on a graphics processing unit according to claim 1, characterized in that, the step of configuring the number of execution threads of each graphics processing unit includes: Configuring the number of thread blocks corresponding to each graphics processing unit and the number of threads in each thread block.
4. The discontinuous finite element seismic numerical simulation method based on a graphics processing unit according to claim 3, characterized in that, the step of configuring the shared memory of each thread block of each graphics processing unit includes: Configuring the shared memory of each thread block of each graphics processing unit according to the number of variables to be solved, the maximum number of degrees of freedom of discontinuous finite element units, the number of threads in the thread blocks corresponding to the graphics processing unit, and the number of bytes of a single data.
5. The discontinuous finite element seismic numerical simulation method based on a graphics processing unit according to claim 4, characterized in that, the number of intermediate results is the same as the number of variables to be solved, and each intermediate result occupies a total size in the shared memory that is the product of the maximum number of degrees of freedom of the unit, the number of threads, and the number of bytes of a single data.
6. The discontinuous finite element seismic numerical simulation method based on a graphics processing unit according to claim 1, characterized in that, the step of calculating the data of each discontinuous finite element unit in parallel loop according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread includes: Calculating the data of each discontinuous finite element unit in parallel loop according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the thread identification information corresponding to different discontinuous finite element unit numbers, and the shared memory address of each thread.
7. A discontinuous finite element seismic numerical simulation method based on a graphics processing unit according to claim 1, wherein, the step of storing the intermediate result of the parallel loop calculation into the shared memory includes: storing the intermediate result of the parallel loop calculation into the shared memory in the order of discontinuous finite element cells first and degrees of freedom second.
8. A discontinuous finite element seismic numerical simulation device based on a graphics processing unit, wherein, the device includes: an acquisition module, configured to acquire parameters and matrices of discontinuous finite elements according to a discontinuous finite element mesh and a velocity model; a configuration module, configured to configure the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread according to the graphics processing unit device number corresponding to the central processing unit process; a calculation module, configured to, after simulating a seismic source at a preset position, perform parallel loop calculation on data of each discontinuous finite element cell according to the number of execution threads of each graphics processing unit, the shared memory of each thread block, the parallel manner between threads, and the shared memory address of each thread; a storage module, configured to store the intermediate result of the parallel loop calculation into the shared memory, and store the result saved in the shared memory at the end of each final moment into the global memory as the final output result.
9. An electronic device, wherein, the electronic device includes: a memory storing executable instructions; a processor, the processor running the executable instructions in the memory to implement the method according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-7.