Parallel optimization method and device based on two-dimensional linear time-varying sea surface generation

By introducing parallel optimization technology into the sea surface generation method, combining single-machine multi-threading and cluster multi-node parallel strategies, the sea surface generation algorithm is optimized and designed, which solves the problems of large calculation volume and low efficiency of traditional methods, and realizes efficient sea surface generation calculation.

CN120086482APending Publication Date: 2025-06-03BEIJING INST OF ENVIRONMENTAL FEATURES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161475.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When the sea surface size increases, the calculation amount increases sharply, resulting in a rapid decline in calculation efficiency and facing the problem of failure.

Method used

The parallel optimization method based on two-dimensional linear time-varying sea surface generation is adopted. By obtaining the standard sea spectrum method, the performance analysis and communication overhead determination are performed. Combining a single-machine multi-threaded parallel optimization strategy and a cluster multi-node parallel optimization strategy, the design is optimized to achieve parallel optimization of sea surface generation.

Benefits of technology

It improves the calculation efficiency of two-dimensional linear time-varying sea surface generation, reduces communication overhead, solves the problems of large calculation volume and low efficiency of traditional methods, and provides a basis for the calculation of electromagnetic scattering characteristics of super-large time-varying sea surface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086482A_ABST
    Figure CN120086482A_ABST
Patent Text Reader

Abstract

The invention provides a parallel optimization method and device based on two-dimensional linear time-varying sea surface generation. The method comprises the following steps: obtaining each algorithm module of a sea surface generated by a standard sea spectrum method; wherein the algorithm module for generating the sea surface by the standard sea spectrum method comprises a random number generation module, a sea spectrum superposition module and a fast Fourier transform module; performing performance analysis on the serial program of each algorithm module, and determining the communication overhead of each algorithm module; and according to the characteristics of each algorithm module, performing optimization design on the algorithm module with high communication overhead by adopting a mode of combining a single-machine multi-thread parallel optimization strategy and / or a cluster multi-node parallel optimization strategy so as to complete parallel optimization of the two-dimensional linear time-varying sea surface generation. According to the scheme, the calculation efficiency of the sea surface generation method can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of electromagnetic scattering calculation, and particularly to a parallel optimization method and device based on the generation of a two-dimensional linearly time-varying sea surface. Background Art

[0002] The generation of a geometric sea surface is a basic step in the electromagnetic scattering calculation of a two-dimensional time-varying sea surface. Due to the complexity of the sea surface and the variety of sea spectra, a geometric sea surface is usually generated based on the standard sea spectrum method. The standard sea spectrum generally adopts three sea spectrum models, such as the PM sea spectrum, the JONSWAP sea spectrum, and the Fung sea spectrum.

[0003] However, when the size of the sea surface continues to increase, the computational amount of the standard sea spectrum method increases sharply, and the computational efficiency drops rapidly, making the commonly used sea surface generation methods face the problem of failure.

[0004] Therefore, there is an urgent need to provide a parallel optimization method and device based on the generation of a two-dimensional linearly time-varying sea surface. Summary of the Invention

[0005] In order to solve the problems of large computational amount and low computational efficiency of traditional sea surface generation methods, the embodiments of the present invention provide a parallel optimization method and device based on the generation of a two-dimensional linearly time-varying sea surface.

[0006] In a first aspect, the embodiments of the present invention provide a parallel optimization method based on the generation of a two-dimensional linearly time-varying sea surface. The method includes:

[0007] Obtaining each algorithm module for generating a sea surface by the standard sea spectrum method; wherein, the algorithm modules for generating a sea surface by the standard sea spectrum method include a random number generation module, a sea spectrum superposition module, and a fast Fourier transform module;

[0008] Performing performance analysis on the serial program of each algorithm module, and determining the communication overhead of each algorithm module;

[0009] According to the characteristics of each algorithm module, adopting a combination of a single-machine multi-thread parallel optimization strategy and / or a cluster multi-node parallel optimization strategy to optimize and design the algorithm module with large communication overhead, so as to complete the parallel optimization of the generation of the two-dimensional linearly time-varying sea surface.

[0010] In a second aspect, the embodiments of the present invention further provide a parallel optimization device based on the generation of a two-dimensional linearly time-varying sea surface. The device includes:

[0011] An obtaining unit, configured to obtain each algorithm module for generating a sea surface by the standard sea spectrum method; wherein, the algorithm modules for generating a sea surface by the standard sea spectrum method include a random number generation module, a sea spectrum superposition module, and a fast Fourier transform module;

[0012] An analysis unit for performing performance analysis on the serial programs of each algorithm module and determining the communication overhead of each algorithm module;

[0013] An optimization unit for optimizing and designing the algorithm modules with large communication overhead by combining a single-machine multi-thread parallel optimization strategy and / or a cluster multi-node parallel optimization strategy according to the characteristics of each algorithm module, so as to complete the parallel optimization of the two-dimensional linear time-varying sea surface generation.

[0014] In a third aspect, an embodiment of the present invention further provides a computing device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method described in any embodiment of this specification is implemented.

[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in any embodiment of this specification.

[0016] On the other hand, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method described in any of the above embodiments.

[0017] An embodiment of the present invention provides a parallel optimization method based on two-dimensional linear time-varying sea surface generation. Based on the standard sea spectrum method, first perform performance analysis on the serial programs of each algorithm module in the standard sea spectrum method to determine the algorithm modules with large communication overhead, and then for each algorithm module point, introduce multiple strategies such as MPI and OpenMP to achieve algorithm acceleration in the single-machine multi-node and cluster multi-node environments. By improving the computing efficiency of the existing serial computing model, the problems of large computational amount and low efficiency in two-dimensional linear time-varying sea surface generation are solved, laying a foundation for the calculation of the electromagnetic scattering characteristics of ultra-large time-varying sea surfaces. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a flowchart of a parallel optimization method based on two-dimensional linear time-varying sea surface generation provided by an embodiment of the present invention;

[0020] Figure 2 is a flowchart of an algorithm for generating the sea surface by the traditional standard sea spectrum method provided by an embodiment of the present invention;

[0021] Figure 3 is an overall partitioning strategy diagram of multi-node parallel sea surface generation provided by an embodiment of the present invention;

[0022] Figure 4 is an overall strategy flowchart of multi-node parallel sea surface generation provided by an embodiment of the present invention;

[0023] Figure 5 is a hardware architecture diagram of a computing device provided by an embodiment of the present invention;

[0024] Figure 6 is a structural diagram of a parallel optimization device for generating a two-dimensional linear time-varying sea surface provided by an embodiment of the present invention. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] The following describes the specific implementation manners of the above concepts.

[0027] Please refer to Figure 1 , an embodiment of the present invention provides a parallel optimization method for generating a two-dimensional linear time-varying sea surface, and the method includes:

[0028] Step 100, obtaining each algorithm module for generating the sea surface by the standard sea spectrum method; wherein, the algorithm modules for generating the sea surface by the standard sea spectrum method include a random number generation module, a sea spectrum superposition module, and a fast Fourier transform module;

[0029] Step 102, performing performance analysis on the serial programs of each algorithm module, and determining the communication overhead of each algorithm module;

[0030] Step 104, according to the characteristics of each algorithm module, adopting a combination of a single-machine multi-thread parallel optimization strategy and / or a cluster multi-node parallel optimization strategy to optimize the algorithm modules with large communication overhead, so as to complete the parallel optimization of the two-dimensional linear time-varying sea surface generation.

[0031] In the embodiments of the present invention, based on the standard sea spectrum method, the serial programs of each algorithm module in the standard sea spectrum method are first analyzed for performance to determine the algorithm modules with large communication overheads. Then, for each algorithm module point, multiple strategies such as MPI and OpenMP are introduced to achieve algorithm acceleration in the single-machine multi-node and cluster multi-node environments. By improving the computational efficiency of the existing serial computing model, the problems of large computational amount and low efficiency in the generation of two-dimensional linearly time-varying sea surfaces are solved, laying a foundation for the calculation of the electromagnetic scattering characteristics of ultra-large time-varying sea surfaces.

[0032] Regarding step 100:

[0033] In the embodiments of the present invention, by analyzing the algorithm principle of two-dimensional time-varying sea surface generation, the algorithm flow is sorted out according to the basic principle of generating the sea surface based on the standard sea spectrum method. The formula for generating the time-varying sea surface undulation height h(x, y, t) based on the linear filtering method of the standard sea spectrum is as follows:

[0034]

[0035] Wherein,

[0036] In the formula, L x and L y are the lengths of the two-dimensional linearly time-varying sea surface in the x-direction and y-direction respectively, i and j are the grid numbers, the complex random number G * is the reverse conjugate of G, A i,j (k xi , k yi , t) is the two-dimensional frequency spectrum, k is the wave number, k x = kcosφ, k y = ksinφ, W(k, φ) is the standard sea spectrum, having the characteristic of W(k, π - φ) = W(-k x , k y , φ), is the angle, and ω is the angular frequency;

[0037] Furthermore, A(k xi , k yi ) has the following characteristics:

[0038] A(k x , k y ) = A * (-k x , -k y ), A(k x , -k y ) = A * (-k x , k y )

[0039] According to the algorithm principle of generating the two-dimensional time-varying sea surface above, it can be known that the standard sea spectrum method for generating the sea surface mainly includes three algorithm modules: a random number generation module, a sea spectrum superposition module, and a fast Fourier transform module (FFT). Further, as Figure 2 shown, the sea spectrum superposition module includes the superposition of the sea spectrum of half of the sea surface and the mapping of half of the sea surface, and the fast Fourier transform module includes two FFTShift calculations and an FFT calculation.

[0040] In some embodiments, in the random number generation module, random numbers are generated by using a parallel library or a centralized random number generation method.

[0041] The random number generation module of the standard sea spectrum method for generating the sea surface generally generates two-dimensional random numbers by calling the Boost library. The Boost library is a serial library. It first needs to establish a large random number pool and then extract random numbers from the pool. The data storage amount is larger than the number of grid points of the actual sea surface dissection. Therefore, in the embodiments of the present invention, by replacing the Boost library with other parallel libraries or adopting a centralized random number generation method to send random numbers from the rank = 0 node to the remaining nodes (rank = 1, 2,...) through data transmission, it is possible to generate random numbers in parallel on multiple nodes, thereby avoiding too high correlation between random numbers of each node.

[0042] Regarding step 102:

[0043] In some embodiments, step 102 includes:

[0044] Obtain the time for generating the time-varying sea surface by the serial program of each algorithm module respectively;

[0045] Compare and sort the times of each algorithm module, and determine the algorithm module corresponding to the generation time exceeding the preset threshold as the algorithm module with large communication overhead; wherein, the algorithm module with large communication overhead includes the sea spectrum superposition module and the fast Fourier transform module.

[0046] In the embodiments of the present invention, first, the performance of the serial program for generating the sea surface based on the standard sea spectrum method is analyzed and tested, and the time taken by the serial programs of each algorithm module to implement the generation of the time-varying sea surface is compared to determine the algorithm module with a large communication overhead. The sea spectrum superposition module optimizes and enhances the data output by the random number generation module by using a two-dimensional matrix filling method. This step requires matrix data transfer, involves a large amount of data interaction, and has a large communication overhead. The fast Fourier transform module includes two FFT Shift calculations and FFT calculations. Based on the principle of FFT, it requires comprehensive data interaction and data aggregation operations, which is the link with the largest communication overhead. After the FFT calculation, FFT Shift is performed on the data to form the random fluctuation height, and this step also requires matrix transfer and involves a large amount of data interaction, resulting in a large communication overhead.

[0047] Table 1 Computational performance of the serial modules for sea surface generation (incident wave frequency: 4 GHz, grid size: 0.0125 m, test environment: 8 cores, 16 GB of memory, Windows operating system)

[0048]

[0049] Table 2 Computational performance of the serial modules for sea surface generation (incident wave frequency: 12 GHz, grid size: 0.0042 m, test environment: 8 cores, 16 GB of memory, Windows operating system)

[0050]

[0051] As can be seen from Table 1 and Table 2, the computational performance of the sea spectrum superposition module and the fast Fourier transform module is the lowest, and the communication overhead is relatively large. The inefficiency of the sea spectrum superposition module is mainly due to the need for sea spectrum calculation and random number superposition on the sea spectrum operation, and the number of loops is m / 2 × n / 2, with a large number of loops and high memory requirements, resulting in low performance. The performance of the FFT calculation module itself is relatively inefficient, and it needs to perform large-scale FFT operations on m × n data.

[0052] Regarding step 104:

[0053] In some embodiments, step 104 includes:

[0054] Perform single-machine multi-thread parallel optimization on the sea spectrum superposition module based on the OpenMP compilation directive statement;

[0055] Use MPI to perform cluster multi-node parallel optimization on the sea spectrum superposition module and the fast Fourier transform module respectively to complete the parallel optimization of the generation of the two-dimensional linear time-varying sea surface.

[0056] In the embodiment of the present invention, a single-machine multi-thread parallel strategy is designed, and the calculation link (Haipu superposition module) with large communication overhead is deeply analyzed to design a single-machine multi-thread parallel strategy. Since the Haipu superposition module is in the same for loop, a total of m / 2×n loops are required, which has natural parallelism, and the single-machine multi-thread acceleration is mainly performed on the Haipu superposition. For example, the OpenMP compilation guide command statement #pragma omp parallel for default(none)private(count,phi,omga,wkp,k)shared(rannum,gmn) can be used to parallelize the peripheral loop of the module.

[0057] Furthermore, there is a large amount of communication overhead in the sea surface generation link, which is the main link restricting the algorithm's computing power. The single-machine multi-threaded parallel strategy based purely on OpenMP has little improvement on the algorithm. It is necessary to combine the cluster multi-node parallel strategy to parallelize the sea spectrum overlay module, the two-dimensional fast Fourier transform module and all the modules of sea surface generation, so as to further improve the generation efficiency of the two-dimensional linear time-varying sea surface.

[0058] When clustering multi-node parallelism for all sea surface generation algorithm modules, by utilizing the data symmetry in the sea spectrum superposition module and the FFT module (they have the characteristics of axisymmetry, quadrant symmetry, center symmetry, etc.), the following is established: Figure 3 and Figure 4 The global parallel data partitioning strategy and strategy flow shown in the figure can avoid communication between nodes as much as possible.

[0059] In some specific implementations, the Haipu superposition module optimizes and enhances the data output by the random number generation module by filling in a two-dimensional matrix;

[0060] The use of MPI to perform cluster multi-node parallel optimization on the Haipu overlay module includes:

[0061] The data of the two-dimensional matrix is ​​evenly divided along the X-axis direction according to a preset number of nodes to obtain a 1 / 2 sea surface space spectrum; wherein each node in the 1 / 2 sea surface space spectrum is used to process one of the corresponding pieces of data in the two-dimensional matrix;

[0062] Each node in the 1 / 2 sea surface spatial spectrum is conjugate expanded along the Y-axis direction to obtain a sea surface mapping spectrum.

[0063] Due to the symmetry of the two-dimensional data matrix in the FFT algorithm, only 1 / 2 of the sea surface needs to be generated for matrix mapping during the calculation in the sea spectrum superposition module. According to this feature, in the embodiments of the present invention, the data of the two-dimensional matrix is first evenly divided along the X-axis direction according to the number of nodes, so that each node is only used to process a corresponding piece of data. As Figure 3 shown in, for a sea surface of size m×n, taking 4 nodes as an example, the data of the two-dimensional matrix is evenly divided along the X-axis direction according to the four nodes, and the 1 / 2 sea surface spatial spectrum data in the first and second quadrants is obtained, so that each node is only responsible for a piece of data in the two-dimensional matrix. After the 1 / 2 sea surface spatial spectrum is generated, conjugate expansion is directly performed on each node along the Y-axis direction without communication to obtain the sea surface mapping spectrum. Using the above method can not only avoid communication between nodes in the sea spectrum superposition module, but also enable each node to perform FFTshift without communication.

[0064] However, in order to adapt to two-dimensional linear FFT calculation (this process is actually an inverse Fourier transform process, and the data must conform to the distribution law after FFT), communication is required before two-dimensional linear FFT calculation, so that the data can meet the splitting requirements of FFT, and result merging is also required after two-dimensional linear FFT calculation. These communication links cannot be avoided.

[0065] In some embodiments, MPI is used to perform cluster multi-node parallel optimization on the fast Fourier transform module respectively, including:

[0066] The two-dimensional matrix output by the sea spectrum superposition module is split to obtain four solution modules; among them, the first solution module corresponds to the values of odd rows and odd columns of the two-dimensional matrix, the second solution module corresponds to the values of odd rows and even columns of the two-dimensional matrix, the third solution module corresponds to the values of even rows and odd columns of the two-dimensional matrix, and the fourth solution module corresponds to the values of even rows and even columns of the two-dimensional matrix;

[0067] The four solution modules are respectively assigned to the corresponding nodes for OpenMP task parallelism, so that each solution module independently completes the two-dimensional fast Fourier transform;

[0068] The MPI collective communication function is used to merge the seas calculated by the four solution modules to generate a two-dimensional linear time-varying sea surface.

[0069] The operation of the two-dimensional fast Fourier transform (FFT2) is the most time-consuming part of the entire sea surface generation. In order to balance the serial call of the fftw3 library and minimize the communication overhead to the greatest extent, in the embodiments of the present invention, it is considered to split the mathematical process of FFT2 into multiple solution modules, and then allocate the solution tasks of each module to the cluster nodes for parallel and independent processing, and finally perform aggregation to solve the problem. For example, the FFT2 mathematical solution process of a sea surface is split into 2×2 functional modules. Each module independently calls the FFT2 solution function in fftw3. Finally, the FFT2 calculation results of these 4 sea surfaces are aggregated using the MPI collective communication function MPI_ALLGATHER, and addition and subtraction operations are performed on the four sea surfaces in each node to finally achieve the merging of the sea surfaces.

[0070] It should be noted that the change of phase needs to be considered additionally during the addition and subtraction operations. At the same time, since the module splitting and merging operations will affect the calculation efficiency, the splitting module and the merging module are operated in parallel with OpenMP.

[0071] In summary, in the embodiments of the present invention, based on generating the sea surface by the standard spectrum method, a standard sea spectrum algorithm based on the OpenMP / MPI parallel optimization technology is established to accelerate the generation efficiency of the two-dimensional linear time-varying sea surface, thereby breaking through the bottleneck problems of calculation efficiency and memory requirements to meet the calculation requirements in practical engineering applications.

[0072] In order to verify the advantages of the parallel optimization method in the embodiments of the present invention, the incident wave is set to 1 GHz and the grid division density is 0.05 m. Tables 3 to 6 give the parallel efficiency test based on OpenMP for a single machine with multiple threads. The test environment is the Windows operating system, the CPU is Intel Xeon E312xx (16 cores), and the memory is 32 GB.

[0073] Table 3 OpenMP parallel results for generating original data (2 threads)

[0074]

[0075] Table 4 OpenMP parallel results for generating original data (4 threads)

[0076]

[0077] Table 5 OpenMP parallel results for generating original data (8 threads)

[0078]

[0079] Table 6 OpenMP parallel results for generating original data (16 threads)

[0080]

[0081] As can be seen from Tables 3 to 6, OpenMP has a good acceleration effect on the sea spectrum superposition module, and the parallel efficiency decreases as the number of threads increases. This is because the communication overhead increases. Thus, although OpenMP can significantly speed up the operation, it has high memory requirements, and only using OpenMP cannot improve the computing power of the algorithm.

[0082] Tables 7 and 8 present the performance test results of the cluster multi-node parallel FFT2 based on MPI. The test environment for all 4 nodes is the Windows operating system, the CPUs are all Intel Xeon E312xx (16 cores), and the memory is all 32GB.

[0083] Table 7 Comparison of Serial and Parallel Times for FFT2 Calculation (4 Nodes)

[0084]

[0085]

[0086] Table 8 Comparison of Serial and Parallel Times for All Modules in Sea Surface Generation (4 Nodes)

[0087]

[0088] MPI cluster multi-node parallel processing is respectively performed on the FFT2 module and all modules for sea surface generation. As can be seen from Tables 7 and 8, the parallel efficiency can reach more than 50% at most in the case of 4 nodes. Through analysis, the main reason for the low parallel efficiency in some links is that communication is required before and after the FFT2 calculation, and the communication overhead is large, resulting in low parallel efficiency. Subsequently, the sea surface generation part can be further optimized to improve the parallel efficiency.

[0089] Through the above examples, it can be seen that the parallel optimization technology based on two-dimensional linear time-varying sea surface generation proposed in the embodiments of the present invention introduces multiple strategies such as MPI and OpenMP to achieve algorithm acceleration in the single-machine multi-node and multi-machine multi-node environments, solves the problem of too long calculation time, improves the calculation efficiency of the existing serial calculation model, and lays a foundation for the calculation of the electromagnetic scattering characteristics of ultra-large time-varying sea surfaces.

[0090] As Figure 5 , Figure 6 shown, the embodiments of the present invention provide a parallel optimization device based on two-dimensional linear time-varying sea surface generation. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, as Figure 5 shown, it is a hardware architecture diagram of a computing device where the parallel optimization device based on two-dimensional linear time-varying sea surface generation provided by the embodiments of the present invention is located. Except for Figure 5In addition to the processor, memory, network interface, and non-volatile memory shown, the computing device where the device is located in the embodiments generally may also include other hardware, such as a forwarding chip responsible for processing packets, and so on. Taking software implementation as an example, as Figure 6 shown, as a logically meaningful device, it is formed by the CPU of the computing device where it is located reading the corresponding computer program in the non-volatile memory into the memory and running it.

[0091] A parallel optimization device based on two-dimensional linear time-varying sea surface generation provided in this embodiment, the device includes:

[0092] An acquisition unit 601, configured to acquire each algorithm module for generating a sea surface by the standard sea spectrum method; wherein, the algorithm modules for generating a sea surface by the standard sea spectrum method include a random number generation module, a sea spectrum superposition module, and a fast Fourier transform module;

[0093] An analysis unit 602, configured to perform performance analysis on the serial program of each algorithm module, and determine the communication overhead of each algorithm module;

[0094] An optimization unit 603, configured to perform an optimization design on the algorithm modules with large communication overhead by combining a single-machine multi-thread parallel optimization strategy and / or a cluster multi-node parallel optimization strategy according to the characteristics of each algorithm module, so as to complete the parallel optimization of the two-dimensional linear time-varying sea surface generation.

[0095] In the embodiments of the present invention, the acquisition unit 601 may be used to execute step 100 in the above method embodiments, the analysis unit 602 may be used to execute step 102 in the above method embodiments, and the optimization unit 603 may be used to execute step 104 in the above method embodiments.

[0096] In an embodiment of the present invention, in the acquisition unit 601, in the random number generation module, random numbers are generated by using a parallel library or a centralized random number generation method.

[0097] In an embodiment of the present invention, when the analysis unit 602 performs performance analysis on the serial program of each algorithm module and determines the communication overhead of each algorithm module, it is configured to perform the following operations:

[0098] Respectively obtain the time for each algorithm module's serial program to generate a time-varying sea surface;

[0099] Compare and sort the times of each algorithm module, and determine the algorithm modules corresponding to the generation time exceeding a preset threshold as the algorithm modules with large communication overhead; wherein, the algorithm modules with large communication overhead include a sea spectrum superposition module and a fast Fourier transform module.

[0100] In an embodiment of the present invention, when the optimization unit 603 optimizes the algorithm module with large communication overhead by combining the single - machine multi - thread parallel optimization strategy and the cluster multi - node parallel optimization strategy, it is used to perform the following operations:

[0101] Perform single - machine multi - thread parallel optimization on the sea spectrum superposition module based on the OpenMP compilation directive statement;

[0102] Use MPI to perform cluster multi - node parallel optimization on the sea spectrum superposition module and the fast Fourier transform module respectively to complete the parallel optimization of the generation of the two - dimensional linear time - varying sea surface.

[0103] In an embodiment of the present invention, in the optimization unit 603, the sea spectrum superposition module optimizes and enhances the data output by the random number generation module by means of two - dimensional matrix filling;

[0104] When the optimization unit 603 uses MPI to perform cluster multi - node parallel optimization on the sea spectrum superposition module, it is used to perform the following operations:

[0105] Divide the data of the two - dimensional matrix evenly along the X - axis direction according to the preset number of nodes to obtain a 1 / 2 sea surface spatial spectrum; wherein each node in the 1 / 2 sea surface spatial spectrum is respectively used to process one of the corresponding pieces of data in the two - dimensional matrix;

[0106] Conjugately expand each node in the 1 / 2 sea surface spatial spectrum along the Y - axis direction to obtain a sea surface mapping spectrum.

[0107] In an embodiment of the present invention, when the optimization unit 603 uses MPI to perform cluster multi - node parallel optimization on the fast Fourier transform module respectively, it is used to perform the following operations:

[0108] Split the two - dimensional matrix output by the sea spectrum superposition module to obtain four solution modules; wherein the first solution module corresponds to the values of the odd rows and odd columns of the two - dimensional matrix, the second solution module corresponds to the values of the odd rows and even columns of the two - dimensional matrix, the third solution module corresponds to the values of the even rows and odd columns of the two - dimensional matrix, and the fourth solution module corresponds to the values of the even rows and even columns of the two - dimensional matrix;

[0109] Allocate the four solution modules to the corresponding nodes for OpenMP task parallelism respectively, so that each solution module independently completes the two - dimensional fast Fourier transform;

[0110] Use the MPI collective communication function to merge the seas calculated by the four solution modules to generate a two - dimensional linear time - varying sea surface.

[0111] It is understandable that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on a parallel optimization device based on a two-dimensional linearly time-varying sea surface. In other embodiments of the present invention, a parallel optimization device based on a two-dimensional linearly time-varying sea surface may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0112] Regarding the information interaction, execution process, etc. among the various modules within the above-mentioned device, since they are based on the same concept as the method embodiments of the present invention, the specific content can be referred to the descriptions in the method embodiments of the present invention, and will not be elaborated here.

[0113] The embodiments of the present invention further provide a computing device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, a parallel optimization method based on a two-dimensional linearly time-varying sea surface in any embodiment of the present invention is implemented.

[0114] The embodiments of the present invention further provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute a parallel optimization method based on a two-dimensional linearly time-varying sea surface in any embodiment of the present invention.

[0115] Specifically, a system or device equipped with a storage medium can be provided. Software program codes for implementing the functions in any of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium.

[0116] In this case, the program codes read from the storage medium itself can implement the functions in any one of the above embodiments. Therefore, the program codes and the storage medium storing the program codes constitute a part of the present invention.

[0117] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program codes can be downloaded from a server computer via a communication network.

[0118] Furthermore, it should be clear that not only can the actual operations in part or in whole be completed by executing the program codes read by the computer, but also by instructions based on the program codes to cause an operating system or the like operating on the computer, thereby implementing the functions in any one of the above embodiments.

[0119] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion module connected to the computer, and then based on the instructions of the program code, the CPU etc. installed on the expansion board or the expansion module are made to execute part or all of the actual operations, thereby implementing the functions of any of the above embodiments.

[0120] An embodiment of the present application also provides a computer-readable storage medium, on which at least one instruction, at least one segment of program, code set or instruction set is stored, and the at least one instruction, at least one segment of program, code set or instruction set is loaded and executed by a processor to implement a parallel optimization method based on two-dimensional linear time-varying sea surface generation provided by the above method embodiments.

[0121] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0122] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk or optical disc that can store program code.

[0123] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A parallel optimization method based on two-dimensional linear time-varying sea surface generation, characterized in that: include: Obtain various algorithm modules for generating sea surface using the standard sea spectrum method; wherein the algorithm module for generating sea surface using the standard sea spectrum method includes a random number generation module, a sea spectrum superposition module, and a fast Fourier transform module; Perform performance analysis on the serial program of each algorithm module and determine the communication overhead of each algorithm module; According to the characteristics of each algorithm module, a single-machine multi-threaded parallel optimization strategy and / or a cluster multi-node parallel optimization strategy are combined to optimize the design of the algorithm module with high communication overhead to complete the parallel optimization of the two-dimensional linear time-varying sea surface generation.

2. The method according to claim 1, characterized in that In the random number generation module, a parallel library or a centralized random number generation method is used to generate random numbers.

3. The method according to claim 1, characterized in that The performance analysis of the serial program of each algorithm module and the determination of the communication overhead of each algorithm module include: Get the time for the serial program of each algorithm module to generate the time-varying sea surface respectively; The time of each algorithm module is compared and sorted, and the algorithm module corresponding to the generation time exceeding the preset threshold is determined as the algorithm module with large communication overhead; wherein the algorithm module with large communication overhead includes a spectrum superposition module and a fast Fourier transform module.

4. The method according to claim 3, characterized in that The method of combining the single-machine multi-threaded parallel optimization strategy with the cluster multi-node parallel optimization strategy to optimize the algorithm module with high communication overhead includes: Based on the OpenMP compilation guidance command statement, the Haipu overlay module is optimized for single-machine multi-thread parallel operation; MPI is used to perform cluster multi-node parallel optimization on the sea spectrum overlay module and the fast Fourier transform module respectively, so as to complete the parallel optimization of the two-dimensional linear time-varying sea surface generation.

5. The method according to claim 4, characterized in that The Haipu superposition module optimizes and enhances the data output by the random number generation module by means of two-dimensional matrix filling; The use of MPI to perform cluster multi-node parallel optimization on the Haipu overlay module includes: The data of the two-dimensional matrix is ​​evenly divided along the X-axis direction according to a preset number of nodes to obtain a 1 / 2 sea surface space spectrum; wherein each node in the 1 / 2 sea surface space spectrum is used to process one of the corresponding pieces of data in the two-dimensional matrix; Each node in the 1 / 2 sea surface spatial spectrum is conjugate expanded along the Y-axis direction to obtain a sea surface mapping spectrum.

6. The method according to claim 5, characterized in that The fast Fourier transform module is optimized in parallel using MPI, including: The two-dimensional matrix output by the Haipu superposition module is split to obtain four solution modules; wherein the first solution module corresponds to the values ​​of the odd numbers of the rows and columns of the two-dimensional matrix, the second solution module corresponds to the values ​​of the odd numbers of the rows and columns of the two-dimensional matrix, the third solution module corresponds to the values ​​of the even numbers of the rows and columns of the two-dimensional matrix, and the fourth solution module corresponds to the values ​​of the even numbers of the rows and columns of the two-dimensional matrix; The four solution modules are respectively assigned to corresponding nodes for OpenMP task parallelism, so that each solution module can independently complete the two-dimensional fast Fourier transform; The sea surfaces calculated by the four solution modules are merged using the MPI collective communication function to generate a two-dimensional linear time-varying sea surface.

7. A parallel optimization device based on two-dimensional linear time-varying sea surface generation, characterized in that: include: An acquisition unit, used for acquiring various algorithm modules for generating sea surface using the standard sea spectrum method; wherein the algorithm modules for generating sea surface using the standard sea spectrum method include a random number generation module, a sea spectrum superposition module and a fast Fourier transform module; An analysis unit, for performing performance analysis on the serial program of each algorithm module and determining the communication overhead of each algorithm module; The optimization unit is used to optimize the design of algorithm modules with high communication overhead by combining single-machine multi-threaded parallel optimization strategy and / or cluster multi-node parallel optimization strategy according to the characteristics of each algorithm module, so as to complete the parallel optimization of the two-dimensional linear time-varying sea surface generation.

8. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.