Marine Data Assimilation Method and System Based on High-Performance Parallel Optimization

By generating a computational topology map for load balancing, optimizing reading with OST, designing three-layer communication and writeback strategies, the problems of load imbalance, slow reading, inefficient communication and unstable writeback in marine data assimilation are solved, and the performance and scalability of marine data assimilation programs are improved.

CN114546638BActive Publication Date: 2025-07-04INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210100983.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-07-04
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

There are problems in the existing marine data assimilation technology such as unbalanced load balancing, long reading time, low communication efficiency, time-consuming and unstable write back and poor scalability, especially in supercomputing cluster environments.

Method used

By generating a load balancing strategy based on computation topology graphs, OST-based read optimization, three-layer communication algorithm and write back optimization, combined with overlapping processing of calculation, read and write back, the ocean data assimilation process is optimized.

Benefits of technology

It realizes load balancing, shortened read time, improved communication efficiency, and stable write back time in supercomputing cluster environment, and improves program scalability and overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114546638B_ABST
    Figure CN114546638B_ABST
Patent Text Reader

Abstract

The present invention provides an ocean data assimilation method and system based on high-performance parallel optimization, including: obtaining ocean data to be assimilated and a mathematical model, calculating the data assimilation complexity of each ocean grid point respectively according to the background field data and the ocean grid distribution map in the ocean data, and generating a calculation topology map based on the data assimilation complexity; grouping the grid points in the ocean grid distribution map according to a preset longitude range, and according to the calculation topology map, statistically calculating the overall assimilation complexity corresponding to each group, so as to evenly allocate multiple computing nodes to each group, and evenly dividing the computing workload responsible for each computing node to the computing processes within the computing node; after the computing processes within the computing node complete their respective data assimilation tasks, obtaining the assimilated result data, and writing the assimilated result data back to the mathematical model as the ocean data assimilation result. The present invention uses the calculation topology map as the basis for load balancing division, and realizes the load balancing of each process in the distributed computing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of high-performance computing and ocean data assimilation, and particularly relates to an ocean data assimilation method and system based on high-performance parallel optimization. Background Art

[0002] Ocean data assimilation is a prediction correction method that combines physical models with observational data and has evolved from the field of climate numerical prediction. It is widely used in many fields such as the atmosphere, ocean, and land surface. Currently, the most common data assimilation methods are the Ensemble Kalman Filter algorithm (EnKF), the Ensemble Optimal Interpolation algorithm, and the three-dimensional and four-dimensional variational methods, etc. Among them, the Ensemble Optimal Interpolation is developed from the Ensemble Kalman Filter.

[0003] Since ocean data assimilation uses high-resolution ocean observation data, there are problems such as large computational volume, high memory requirements, and long I / O time during assimilation. In the prior art, some optimization work has been done on the data assimilation program in this mode, and good results have been achieved. The flowchart of the prior art program is as Figure 1 shown.

[0004] Regarding the problem of load balancing division in the assimilation calculation process, currently, there are mainly Figure 2 and 3 shown load balancing divisions based on uniform domain decomposition and load balancing divisions based on non-uniform domain decomposition. The load balancing based on uniform domain decomposition has the problem of uneven computational load. Although the load balancing division based on non-uniform domain decomposition can ensure computational load balance to a certain extent, due to the overly irregular divided regions, the reading efficiency is extremely low. The prior art combines the advantages of the above two methods and proposes a method for unevenly dividing the grid according to the computational tasks, as Figure 4 shown. This method only changes the longitude range while keeping the latitude unchanged, thus ensuring equal computational amounts in each geographical region to a certain extent. This method ensures computational load balance to a certain extent and also ensures that the grid division is not overly complex.

[0005] Regarding the problem of low reading efficiency in the assimilation calculation process, the prior art has proposed a reading method based on an I / O proxy. As Figure 5 shown, this strategy designates several processes as I / O proxies. Each proxy is responsible for reading one strip of data, and the remaining processes do not perform I / O operations but obtain data by communicating with the proxy through MPI. This can ensure that the program does not cause disk contention due to the expansion of the parallel scale, thereby affecting the I / O efficiency.

[0006] The above prior art has the following problems:

[0007] Load balancing problem. In the prior art, the computational loads of each process are extremely unbalanced. A large number of processes have no computational tasks, and most of the computational tasks are assigned to a small number of processes. Moreover, the loads within these processes are also very unbalanced, resulting in a great waste of processes. In the prior art, the assimilation program divides the computational tasks of processes based on the number of ocean grid points in the assimilated data for load balancing. Such a division method only understands ocean data assimilation from the perspective of the assimilation process and does not deeply understand the assimilation algorithm and the computational process of the processes. Therefore, under this load balancing method, although the loads of each process seemingly appear balanced on the surface, there is actually a large load imbalance during execution. Although the number of assimilable grid points obtained by each process is the same, not every ocean grid point can be assimilated. When the number of valid observation data points nobs <= 2 around the assimilable grid point, it is not possible to assimilate because the data is too little to accurately predict the ocean climate conditions in the geographical area. In the worst case, all the ocean grid points obtained by some processes need to be assimilated, while the processes obtained by some processes do not need to be assimilated at all. This leads to a great load imbalance and a waste of a large number of computing nodes.

[0008] Reading problem. In the prior art, not only is the reading time very long when reading data, but the reading time also increases as the number of cores increases. The prior art uses the I / O proxy method for file reading. Such a design method only optimizes the reading of the assimilation program at the upper algorithm level and does not consider the file system and communication architecture of the supercomputer cluster. Therefore, the parallel reading ability of the Lustre parallel file system is not fully exploited, the parallel throughput utilization rate is very low, and the reading time gradually increases as the number of cores increases.

[0009] Communication problem. The communication in the prior art not only takes a long time but also increases significantly as the number of cores increases. After the I / O proxy process in the prior art reads the corresponding data, it centrally distributes the data. This distribution process is a one-to-many process, and one I / O proxy process is responsible for the data preparation and data distribution of all group members. Such a design method results in a very large communication pressure on the I / O proxy process. At the same time, the prior art does not consider the node communication architecture of the supercomputer cluster and simply calls the MPI_Scatter function for data distribution, resulting in a large communication load imbalance and a very low utilization rate of the communication channels in the supercomputer cluster.

[0010] Write-back problem. The write-back of the prior art is very time-consuming and becomes extremely unstable as the number of cores increases. When the existing program performs write-back, it adopts the strategy of each process directly writing back to the corresponding data block. Such a strategy will cause a large number of processes to densely access the same OST (Object Storage Target) simultaneously, and the OST cannot handle the requests of so many I / O processes at the same time, resulting in great conflicts, especially when the number of cores reaches more than a thousand. In addition, a large number of processes accessing the same disk simultaneously will also lead to serious head contention and a large number of addressings.

[0011] Scalability problem. When the program of the prior art is tested on a scale of more than a thousand cores, when the number of cores is greater than 2000, as the number of cores increases, the overall running time does not decrease significantly or even increases. Summary of the Invention

[0012] In response to the above five technical problems existing in the prior art, the present invention proposes corresponding technical solutions respectively:

[0013] In response to the load balancing problem. The present invention deeply analyzes the algorithm of ocean data assimilation and the actual calculation process of the process, and observes that nobs is the key to achieving load balancing for each process. Moreover, the most time-consuming part in the actual assimilation process is the SVD decomposition process, and its computational complexity is nobs, where the value range of nobs can be 3 to 1000. It can also be seen from this that in the load balancing method based on ocean grid points, some ocean grid points have a very large nobs, while some ocean grid points have a very small nobs, and the maximum difference in computational complexity is 1000^3 / 27, which is more than 3.7 million times. Such a load balancing method is very unreasonable. In response to this characteristic in the actual execution process, the present invention adds a pre-assimilation step before actual assimilation, calculates the computational complexity of data assimilation for each ocean grid point respectively, generates a computational topology map based on the computational complexity on the basis of the original ocean grid point distribution map, and uses this as the basis for load balancing division. The load balance of each process is achieved to the greatest extent. Not only that, the implemented computational topology map can also largely reflect the distribution characteristics of the current data assimilation calculation task, and show whether the corresponding ocean grid points in the entire ocean assimilation data need to be assimilated. Therefore, the present invention realizes the read-write and communication optimization for the calculation distribution characteristics of ocean assimilation data, and this part will be introduced below. Furthermore, in order to combine with the three-layer communication in the following text and for more fine-grained load balancing, the present invention introduces a two-layer load balancing strategy.

[0014] Regarding the reading problem. Data assimilation based on the LICOM mode itself involves a large amount of file reading. When the assimilation time is not fully optimized, the reading time can be ignored. For example, in the Fortran version of the prior art, the running time of the program with thousands of cores is about 8000 seconds, and the proportion of the reading time in the total time is very small and can be ignored. However, after the assimilation time is fully optimized, the reading time will account for a large proportion of the overall running time of the program. Especially after the load balancing optimization of the present invention, for example, the execution time of other parts of the program with thousands of cores is only dozens of seconds, while the program needs 97 seconds to read 64.8G of data. Therefore, the optimization of the reading time is also the key to improving the performance of the parallel assimilation program.

[0015] In the optimization of traditional reading algorithms, most of the optimization algorithms reorganize the data to ensure the continuous reading of data in chunks, reduce the number of file addressing and I / O requests, and thus reduce the file reading time. This is the idea adopted by MPI-IO, and many reading-intensive algorithms use this as the reading optimization strategy and optimize it based on MPI-IO.

[0016] Traditional reading optimizations mostly optimize I / O at the algorithm level without considering how the reading process interacts with the underlying file storage system or the characteristics of the supercomputer cluster. Therefore, these optimization algorithms can only optimize the reading time of the program to a certain extent. If the old reading strategy already involves continuous reading in chunks, the performance improvement of the new optimization algorithm will be very limited and vary greatly with different platform network environments. The present invention designs a reading optimization algorithm based on OST for the above problems. This algorithm combines the file storage architecture of the supercomputer platform and the characteristics of the supercomputer cluster, fully exploiting the I / O bandwidth of the Lustre file storage architecture and the allocation characteristics of the node communication channels of the supercomputer cluster. The reading time for the program to read a 64.8G file is reduced to about 10 seconds, and the I / O performance is improved by 90%.

[0017] Regarding the communication problem. In the existing program, after the reading process reads the corresponding data, it directly distributes the data scatter to all processes within the group. Although this implementation method is simple, its efficiency is very low. In actual execution, a reading process may need to distribute data to hundreds or thousands of processes simultaneously. Although MPI optimizes scatter using a binary method at the underlying level, it does not pay attention to Figure 6In this case, there are problems with the communication channels within and between nodes on the supercomputer cluster. There may be a situation where, in the MPI scatter optimization algorithm, the distribution processes are concentrated on one or several nodes, and the communication channels of one node are used to distribute data outward at the same time, which will cause a large network congestion and the distribution efficiency is very low. MPI_Scatter is just a general method, and the actual execution effect is not good.

[0018] Based on the Figure 6 configured characteristics of the communication channels between and within nodes on the supercomputer cluster shown in the figure, a three-layer communication algorithm is designed. And a ring-based distribution strategy is designed during distribution, which well optimizes the communication efficiency.

[0019] Regarding the write-back problem. In the existing program, the write-back time accounts for 26% of the running time of the existing program. Its time is between 200 and 300. And it fluctuates greatly with the number of supercomputer cluster users and the network congestion situation. Sometimes the write-back time can even reach more than 1500 seconds, which is not only very time-consuming but also very unstable. This is mainly because the existing technology uses each computing process to write back the assimilated data alone. Such a write-back strategy will initiate millions of write-back requests to the OSS. At the same time, the OSS needs to process thousands of file write-back requests, which puts a very large pressure on the scheduling of the OSS and the network communication of the supercomputer cluster. When the number of supercomputer cluster users increases and the communication network becomes congested, the write-back performance will be greatly affected. According to the characteristics of the Lustre file storage architecture, the present invention designs a write-back optimization strategy based on OST (Object Storage Target), and the write-back performance is greatly improved.

[0020] Regarding the scalability problem. In the existing program, due to the I / O optimization problem, the overall I / O time of the program will gradually increase as the number of cores increases, affecting the scalability of the program. After optimizing the assimilation calculation time in the present invention, as the number of cores increases, the main influencing factor affecting the program in the program becomes the I / O time. In order to improve the scalability of the program, the present invention designs a three-layer overlap of calculation, reading communication, and write-back, which well improves the scalability of the program.

[0021] Specifically, the present invention provides a method for ocean data assimilation based on high-performance parallel optimization, which includes:

[0022] A pre-assimilation step of obtaining the ocean data to be assimilated and a mathematical model, and respectively calculating the data assimilation complexity of each ocean grid according to the background field data and the ocean grid distribution map in the ocean data, so as to generate a calculation topology map based on the data assimilation complexity on the basis of the ocean grid distribution map;

[0023] Load balancing step: Group the grid points in the ocean grid distribution map according to a preset longitude range, and based on the calculated topology map, calculate the overall assimilation complexity corresponding to each group, so as to evenly allocate multiple computing nodes to each group, and evenly divide the computing workload responsible for each computing node among the computing processes within the computing node;

[0024] Data assimilation step: After the computing processes within the computing node complete their respective data assimilation tasks, they obtain the assimilated result data and write the assimilated result data back to the mathematical model as the ocean data assimilation result.

[0025] The described ocean data assimilation method based on high-performance parallel optimization, wherein the pre-assimilation step includes:

[0026] Background field data reading step: Divide the background field data equally into multiple object storage OSTs, set multiple reading groups group composed of computing nodes, and set a reading process for each reading group to perform parallel reading on the object storage OST it is responsible for. After the reading is completed, distribute the data to each reading group group, and each group performs subsequent calculation work on the data assimilation complexity after receiving the data.

[0027] The described ocean data assimilation method based on high-performance parallel optimization, wherein the load balancing step includes:

[0028] Set multiple reading processes to communicate with the object storage OST to obtain the corresponding data; each reading process is located in a different computing node respectively, independently uses its own communication channel to interact with its corresponding OST, and does not affect other reading processes to achieve parallel OST interaction;

[0029] After the reading process obtains the corresponding data, divide the computing processes into split groups in units of nodes, where the 0th process of each computing node is the master of this spilt group; the 0th process is responsible for sending and receiving data between nodes and exclusively occupies the communication channel of the current node;

[0030] Group the grid points in the ocean grid distribution map according to a preset longitude range

[0031] Horizontally cut the ocean data to be assimilated according to the longitude range, and divide it into n groups group in total; according to the number of groups, divide all master processes into n master groups, each group corresponding master group corresponds to a data block, configure multiple reading processes for each group, and have all reading processes parallelly read the corresponding data block of the ocean data and parallelly distribute it to the corresponding master process.

[0032] The described ocean data assimilation method based on high-performance parallel optimization, in which the master processes within each group are further evenly grouped to generate scatter_master groups. Initially, each read process corresponds to a scatter_matser group, which is used to distribute data to the master processes within the scatter_matser group. After one round of distribution, the corresponding relationship between the read processes and the scatter_matser groups is adjusted until all the data is distributed.

[0033] The described ocean data assimilation method based on high-performance parallel optimization, which further includes using the ocean data assimilation result to generate an ocean meteorological prediction map at a specified time point.

[0034] The present invention also proposes an ocean data assimilation system based on high-performance parallel optimization, which includes:

[0035] A pre-assimilation module that acquires the ocean data to be assimilated and a mathematical model, calculates the data assimilation complexity of each ocean grid point respectively according to the background field data and the ocean grid point distribution map in the ocean data, so as to generate a calculation topology map based on the data assimilation complexity on the basis of the ocean grid point distribution map;

[0036] A load balancing module that groups the grid points in the ocean grid point distribution map according to a preset longitude range, and statistically calculates the overall assimilation complexity corresponding to each group according to the calculation topology map, so as to evenly allocate multiple computing nodes to each group, and evenly divide the computing workloads responsible for by the computing nodes to the computing processes within the computing nodes;

[0037] A data assimilation module that, after the computing processes within the computing nodes complete their respective data assimilation tasks, obtains the assimilated result data and writes the assimilated result data back to the mathematical model as the ocean data assimilation result.

[0038] The described ocean data assimilation system based on high-performance parallel optimization, in which the pre-assimilation module includes:

[0039] A background field data reading module that evenly distributes the background field data to multiple object storage OSTs, sets multiple reading groups group composed of computing nodes, and sets a read process for each reading group to perform parallel reading on the object storage OST it is responsible for. After the reading is completed, the data is distributed to each reading group group, and each group performs subsequent data assimilation complexity calculation work after receiving the data.

[0040] The described ocean data assimilation system based on high-performance parallel optimization, in which the load balancing module includes:

[0041] Multiple read processes are set up to communicate with the object storage OST to obtain corresponding data; each read process is located on a different computing node and independently uses its own communication channel to interact with its corresponding OST, achieving parallel OST interaction without affecting other read processes.

[0042] After the read processes obtain the corresponding data, the computing processes are divided into split groups based on nodes, where the 0th process of each computing node is the master of this spilt group; the 0th process is responsible for sending and receiving data between nodes and exclusively occupies the communication channel of the current node.

[0043] Group the grid points in the ocean grid distribution map according to a preset longitude range.

[0044] Horizontally cut the ocean data to be assimilated according to this longitude range, and divide it into n groups group; according to the number of groups, divide all master processes into n master groups, each group corresponding master group corresponds to a data block, configure multiple read processes for each group, and let all read processes parallelly read the corresponding data blocks of the ocean data and distribute them to the corresponding master processes in parallel.

[0045] In the described ocean data assimilation system based on high-performance parallel optimization, the master processes within each group group are further evenly grouped to generate scatter_master groups. Initially, each read process corresponds to a scatter_matser group, which is used to distribute data to the master processes within the scatter_matser group. After one round of distribution, adjust the corresponding relationship between the read processes and the scatter_matser groups until all data is distributed.

[0046] The described ocean data assimilation system based on high-performance parallel optimization further includes using the assimilation result of the ocean data to generate an ocean meteorological prediction map at a specified time point.

[0047] The present invention also proposes a storage medium for storing a program for executing any one of the ocean data assimilation methods based on high-performance parallel optimization.

[0048] The present invention also proposes a client for any one of the ocean data assimilation systems based on high-performance parallel optimization.

[0049] From the above solutions, the advantages of the present invention are as follows:

[0050] Next, the present invention will show the comparison of the running times before and after the optimization of the overall parallel assimilation program. As Figure 11As shown, it respectively shows the comparison of the overall running time of the existing program, the non-computation overlapping version of the present invention, and the final optimized version of the present invention. It can be seen from the figure that the calculation, reading, and write-back times of the existing program are relatively large. Although the calculation time significantly decreases as the number of cores increases, the file reading time gradually increases as the number of cores increases. At the same time, the write-back time of the existing program is also very unstable as the number of cores increases. In the non-computation overlapping algorithm, it can be seen that the write-back time and pre-assimilation time of the program still account for a certain proportion, and the pre-assimilation time increases as the number of cores increases. In the final optimized version of the present invention, except for the running time of overlapping calculation, reading communication, and write-back, the running time of other parts accounts for a small proportion of the total time. Moreover, by comparing the two optimization algorithms before and after, it can be clearly seen that the reading communication and write-back times achieve complete overlap when the calculation time is long enough. Description of the Drawings

[0051] Figure 1 It is a flowchart of the existing program for LICOM data assimilation;

[0052] Figure 2 It is a schematic diagram of uniform domain decomposition;

[0053] Figure 3 It is a schematic diagram of non-uniform domain decomposition;

[0054] Figure 4 It is a schematic diagram of non-uniform grid splitting according to the calculation task;

[0055] Figure 5 It is a schematic diagram of read optimization based on the I / O agent;

[0056] Figure 6 It is a schematic diagram of the computing node layout and communication mode of the supercomputer cluster;

[0057] Figure 7 It is a comparison chart of the assimilation time of two algorithms;

[0058] Figure 8 It is a comparison chart of the pre-assimilation time before and after optimization;

[0059] Figure 9 : A comparison chart of the reading time of two algorithms;

[0060] Figure 10 It is a comparison chart of the communication time of two algorithms;

[0061] Figure 11 It is a schematic diagram of the overall running time of the Tianhe thousands-of-core program;

[0062] Figure 12 It is a schematic diagram of two-layer load balancing;

[0063] Figure 13Schematic diagram of the read policy for the I / O agent. This read policy only guarantees algorithmically sequential reads, but in actual execution, I / O communication is very busy, disorderly, and mutually blocking, and the OST utilization rate is not high (n > 11).

[0064] Figure 14 Schematic diagram of the read policy based on OST. The I / O communication of this read policy is orderly and does not affect each other, and the OST is fully utilized to achieve fully parallel reads.

[0065] Figure 15 Schematic diagram of three-layer communication

[0066] Figure 16 Schematic diagram of the distribution policy based on the ring

[0067] Figure 17 Schematic diagram of the overlap between read communication and computing

[0068] Figure 18 Schematic diagram of the strategy to avoid redundant reads

[0069] Figure 19 Schematic diagram of the technical effect of the write-back optimization technology based on OST in the present invention Detailed implementation manners

[0070] The purpose of the present invention is to solve the problems of load balancing, reading, communication, write-back, and scalability existing in the above-mentioned prior art, and a high-performance parallel optimization technology for ocean data assimilation is proposed. To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and are described in detail in conjunction with the accompanying drawings of the specification as follows.

[0071] The present invention includes the following key technical points:

[0072] Key point 1, load balancing optimization based on the computational topology graph, two-layer load balancing, and pre-assimilation optimization based on OST; the technical effects are as shown in Figure 7 and 8 shown. It can be seen from Figure 7 that the load balancing algorithm based on the computational topology graph makes full use of the available number of cores, avoids the situation where a small number of cores are assigned a large amount of computing and most cores are only responsible for a small part or even no assigned metrics, and makes full use of the computing performance of the supercomputer cluster, greatly shortening the assimilation time. It can be seen from Figure 8 that before optimization, the pre-assimilation time increases linearly with the increase in the number of cores. After optimization, the pre-assimilation time basically remains stable with the increase in the number of cores.;

[0073] Key point 2, read optimization based on OST; the technical effects are as shown in Figure 9As shown, the reading time of the existing program increases with the increase in the number of cores. The reading time of the program of the present invention is not only much less than that of the existing program, but also basically remains at about 10 seconds with the increase in the number of cores;

[0074] Key point 3, three-layer communication algorithm and ring-based distribution strategy; The technical effect shows that Figure 10 it can be seen that the communication time of the existing program increases linearly with the increase in the number of cores, while the communication time of the program of the present invention basically remains unchanged with the increase in the number of cores;

[0075] Key point 4, write-back optimization based on OST; The technical effect is as Figure 19 shown in the table. The write-back time of the existing program is very long and is very unstable during actual testing. The write-back time of the program of the present invention is not only much less than that of the existing program but also basically remains stable with the increase in the number of cores;

[0076] Key point 5, overlap of three layers of calculation, reading communication, and write-back; The technical effect is as Figure 11 shown.

[0077] Load balancing optimization:

[0078] (1) Load balancing optimization based on the calculation topology map

[0079] Before assimilation execution, the present invention first introduces a pre-assimilation step. This step first reads a background field file, calculates the calculation complexity of each ocean grid point, and generates a global calculation topology map based on this, which is used as the basis for load balancing. The background field belongs to a basic map of marine meteorology. Marine data assimilation needs to select grid points on the background field for assimilation and cover the data above after assimilation is completed. After assimilation is completed, a new global marine meteorology prediction map is generated covering all areas that need to be assimilated.

[0080] (2) Two-layer load balancing

[0081] As Figure 12As shown in the figure, assume that the computational workload of assimilating data for the entire ocean is 90, with 5 cores in each node (split group), and core 0 being the master of the split group. After generating the computational topology graph, for the convenience of display, assume that the entire assimilated data is divided into group = 4 groups, and the computational workload of each group is calculated. The present invention first performs one - layer load balancing in the group dimension according to longitude. Taking group0 as an example, its computational workload is 25. The present invention assigns 5 masters to this group. Since the distribution of observation points in the geographical area is uneven, the geographical area size responsible for each master may be large or small, and the computational workload responsible for each master is 5. Inside each split group, the algorithm of the present invention performs another layer of load balancing, and evenly divides the computational workload again among the computing processes within the node. Here, the split groups belong to different nodes and use the inter - node and intra - node communication channels respectively. Such a two - layer load - balancing method can not only maximize the load balancing of computing processes, but also pave the way for the reading and communication grouping in Section 3. The reading and communication optimization strategies in Section 3 also fully consider the inter - node and intra - node communication channel problems, and implement the corresponding reading and communication optimization strategies according to the load - balancing grouping. Through such a grouping method, the data distribution from the group to the master and from the master to the split group use the inter - node and intra - node communication channels respectively, without affecting each other, which can greatly improve the communication efficiency.

[0082] (3) Pre - assimilation optimization based on OST

[0083] In the load - balancing strategy based on the computational topology graph mentioned above, a pre - assimilation operation needs to be performed. This part needs to first read the data of the background field, and then perform pre - assimilation based on the background field, the observation field, and the grid data to calculate the computational complexity of each ocean grid point, so as to generate the computational topology graph. This process involves reading a background - field file, and the size of this background - field file is 5.4G. Under 1000 cores, the reading and communication time for this part is about 6 seconds, and it grows linearly with the increase in the number of cores. The algorithm of the present invention is optimized to an overall running time of about 34 seconds under 4000 cores. Therefore, the optimization of the pre - assimilation part is also very important. Among them, the observation - field data is the ocean meteorological data of past observations. The grid data is the specific geographical information and water - depth information of each ocean grid point.

[0084] The present invention evenly distributes the 5.4G file into group (group < 12) OSTs, and then sets a reader for each group to perform parallel reading of the OSTs it is responsible for. After the reading is completed, the data is distributed to different groups by imitating the three - layer communication mentioned above. After each group receives the data, it performs the pre - assimilation process.

[0085] Read optimization:

[0086] (1) Read optimization based on OST

[0087] The file system used by the Tianhe-2 supercomputer platform used in the program of the present invention is the Lustre file system. When storing files, the system stores the files in different OSTs in a sharded manner, where the OSS (Object Storage Server) is responsible for processing the upper-layer I / O read and write requests and scheduling the OSTs. Theoretically, if the OSS and OST are in a one-to-one correspondence, the number of OSTs is m, and the I / O bandwidth of the OST is n GB / s, then the bandwidth of the entire supercomputer platform can reach mn GB / s in theory. On the Tianhe-2 supercomputer platform, the number of OSTs is 12 and the number of OSSs is 2. If the scheduling time of the OSS is ignored, the overall bandwidth of the Tianhe system can reach 12n GB / s. During actual measurement, the throughput of the existing program is about 0.5 GB / s. Although there are certain user contention and scheduling problems in the Tianhe system, such throughput far from reaches the peak throughput of Tianhe.

[0088] First, as Figure 13 、 14 shown, there is a situation where a large number of read processes simultaneously access a small number of OSTs in the I / O read algorithm of the existing program. And since the file storage method is not set, by default, the same file may exist in several different OSTs disorderly, which leads to the situation that although different files are read, different processes may access the same OST. Such a read method makes only a small number of OSTs overloaded at the same time, while other OSTs are always idle, without making full use of the OSTs. Moreover, there are a large number of I / O request conflicts in the intensive access to a small number of OSTs at the same time, increasing the scheduling difficulty of the OSS and greatly increasing the scheduling time of the OSS. In the algorithm of the present invention, for 64.8 G, that is, 12 files with 5.4 G of data for each file, the present invention stores them in different OSTs respectively through the user interface of Lustre. When reading files, the present invention only designates 12 processes as readers to read files. For these 12 processes, the present invention also makes special settings. The present invention binds these 12 processes to different Tianhe nodes respectively. The nodes on Tianhe communicate through the inter-node and intra-node communication channels respectively. As Figure 6As shown in the figure, each node on Tianhe has 24 cores, occupying a node - to - node communication channel independently, and the internal communication within the node uses the intra - node communication channel. Such a design can make full use of the bandwidth of each node's communication channel, and maximize the avoidance of conflicts in OST requests. When the number of OSSs is sufficient, the I / O requests of each reading process and the OST are completely parallel, not affected by other reading processes, achieving fully parallel communication. However, as mentioned before, there are only two OSSs in Tianhe, so it is impossible to achieve the situation where there are no I / O request conflicts in theory. But this design can also maximize the avoidance of I / O request conflicts and make full use of the available OSTs in Tianhe at the same time, achieving fully parallel reading.

[0089] Communication optimization:

[0090] (1) Three - layer communication algorithm

[0091] First, as Figure 15 shown, set 12 processes as readers to communicate with the OST first to obtain the corresponding data. Each reader is allocated to a different node, independently using its own communication channel to interact with its corresponding OST, and interacting with other readers without interference to achieve parallel OST interaction. After the readers obtain the corresponding data, the computing processes are divided into split groups based on nodes. The 0 - numbered process of each node is the master of this split group. The 0 - numbered process is responsible for sending and receiving data between nodes, monopolizing the communication channel of the current node.

[0092] For parallel processing, the present invention horizontally slices the assimilated data in the ny direction of the latitude, dividing it into group groups. The present invention divides the masters into group master groups according to the group. Each master group of each group needs a corresponding data block in 12 files. Therefore, the present invention configures 12 readers for each group. In each round, 12 readers are used to parallelly read the corresponding data blocks of 12 files and distribute them to the corresponding masters in parallel. Such a design can concentrate the communication between readers and masters on the node - to - node communication channel, make full use of the communication channels of each node, and at the same time isolate the intra - node communication described below. After the masters obtain the corresponding data, each master distributes the data to the split group it is responsible for through the intra - node communication channel. Such a design can achieve parallel communication between nodes and within nodes on different communication channels.

[0093] (2) Ring - based distribution strategy

[0094] When the 12 readers in each group distribute data to their respective masters, if the scatter method is directly adopted and the 12 readers initiate communication with all masters simultaneously at the same time, "hotspots" may occur, resulting in serious communication congestion. Therefore, the present invention designs Figure 16 the ring-based distribution strategy shown below.

[0095] First, the present invention further evenly divides the masters within each group to generate scatter_master groups. Initially, each reader corresponds to a scatter_matser group and is responsible for distributing data to these masters. After one round of distribution, each reader rotates, as Figure 16 shown below, and cycles in sequence until all data is distributed. This design method can achieve fully parallel data distribution among readers, without generating "hotspots" and communication congestion among readers, greatly improving the communication efficiency.

[0096] Write-back problem:

[0097] (1) Write-back optimization based on OST

[0098] According to the characteristics of the Lustre file storage architecture, the present invention evenly distributes the written-back files to different OSTs according to groups, and then, through three-layer communication, realizes three-layer gather operations. The data calculated by each computing process is collected to reader0 in each group via split -> master -> reader, and reader0 in each group is responsible for centralized write-back, writing the corresponding data back to the corresponding OSTs, achieving fully parallel write-back. The final number of write-back requests is the number of groups, and the write-back time is stable at about 2 - 3 seconds and does not increase with the increase in the number of cores. It is very stable and does not fluctuate with the number of users in the supercomputer cluster and the fluctuations in the network environment.

[0099] Scalability problem:

[0100] (1) Overlapping of read communication and calculation

[0101] Since in the specific data assimilation process, as long as each computing process obtains the corresponding data, the assimilation process can proceed without other communication operations, and the computing processes are completely parallel. For this assimilation process, the present invention realizes the overlap of read communication and computing. To achieve the overlap of read communication and computing, the reader slices the data to be read in the ny direction of the latitude to generate an overlap slice data block. Each time the reader only reads a small piece of data. After reading the corresponding data, it then conducts communication from the reader to the master and from the master to the computing processes within the node. Each computing process conducts assimilation calculations after obtaining the data. While other processes are conducting assimilation calculations, the reader reads the next small piece of data and communicates with the master. In this way, the overlap of read communication and computing is realized.

[0102] Since the assimilated data all have edges, where prep_rx in the nx direction of longitude is 105 and prep_ry in the ny direction of latitude is 19. Assuming that the length of the nx direction of the data block required by each computing process is sub_nx_len and the length of the ny direction is ny_len, then the data that each reader needs to read is ((nx + 2 * prep_rx) * (ny_len + 2 * prep_ry)) * sizeof(type), and the size of the communication data required by each computing process is ((sub_nx_len + 2 * prep_rx * (ny_len + 2 * prep_ry)) * sizeof(type)). Therefore, the finer the data is sliced in the ny direction, the more edge data is required, and the greater the redundancy of reading and communication generated.

[0103] To solve the problem of redundancy in read and communication data, the present invention adopts the method of pre-allocating a memory space of ((nx + 2 * prep_rx) * (ny_len + 2 * prep_ry)) * sizeof(type) during reading. Its size is the memory size of the overall data block without slicing plus the memory size of the edge data, where ny_len is the length of the total data block in the ny direction. Then, psi_read_index is set as the data offset that the reader has read. When the reader reads, it calculates the amount of data to be read next according to psi_read_index, that is, the logically required data offset minus the data offset of the data that has been read into the memory represented by psi_read_index. In actual execution, the amount of data read after slicing is equal to the amount of data read in the non-slicing manner, completely solving the problem of redundant read data caused by edge data and realizing efficient data reuse.

[0104] After solving the problem of redundant data reading, the solution to the communication data redundancy problem will be introduced below. Before communication, the algorithm of the present invention, similar to data reading, first pre-allocates a memory space of ((sub_nx_len + 2 * prep_rx * (ny_len + 2 * prep_ry)) * sizeof(type) in advance, and sets psi_index as the data offset of the assimilated data that has been obtained through communication. During communication, the amount of data to be communicated is calculated using psi_index, that is, subtracting the data offset that has been obtained through communication from the data offset to be communicated. In this way, the final amount of communication data is equal to the amount of unsharded data, completely solving the problem of communication data redundancy caused by border data.

[0105] The specific overlapping algorithms for reading, communication and calculation and the strategy for avoiding redundant reading and communication are as follows Figure 17 、 18 shown.

[0106] (2) Overlapping of calculation, reading communication, and write-back

[0107] As mentioned above, the calculation topology graph can reflect the distribution characteristics of the current data assimilation calculation task. The assimilated data shows the global ocean climate information, so it contains a large amount of land and a large number of ocean areas without observation points. Moreover, these areas are mostly in blocks and continuous. Therefore, according to the distribution characteristics of the calculation amount of ocean assimilated data, a strategy for avoiding reading and communication is designed.

[0108] After generating the calculation topology graph, the entire ocean assimilated data is classified into areas that need to be assimilated and areas that do not need to be assimilated. For the areas that need to be assimilated, task division based on load balancing and reading, distribution, and calculation of assimilated data are performed according to the above algorithm. For the areas that do not need to be assimilated, some readers can be set for reading and write-back. Although these areas do not need to be calculated, in order to generate the analysis field of the global ocean climate, these data should also be written back to the final analysis field.

[0109] In the experiment of the present invention, the areas that do not need to be assimilated account for about 40% of the total ocean assimilated data. The readers responsible for this part of reading and write-back are independent of other processes and can be executed independently. Therefore, after generating the calculation topology graph, these processes can perform write-back of the data that does not need to be assimilated in parallel while other processes are performing calculation and reading communication. Such data reading and write-back can be completely masked by the assimilation calculation of other processes, reducing the time for a large amount of data reading, communication, and write-back in the overall program.

[0110] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0111] The present invention also proposes an ocean data assimilation system based on high-performance parallel optimization, which includes:

[0112] A pre-assimilation module, which acquires the ocean data to be assimilated and a mathematical model, and calculates the data assimilation complexity of each ocean grid point respectively according to the background field data and the ocean grid distribution map in the ocean data, so as to generate a calculation topology map based on the data assimilation complexity on the basis of the ocean grid distribution map;

[0113] A load balancing module, which groups the grid points in the ocean grid distribution map according to a preset longitude range, and statistically calculates the overall assimilation complexity corresponding to each group according to the calculation topology map, so as to evenly allocate multiple computing nodes to each group, and evenly divide the computing workloads responsible for by each computing node among the computing processes within the computing node;

[0114] A data assimilation module, after the computing processes within the computing node complete their respective data assimilation tasks, obtains the assimilated result data, and writes the assimilated result data back to the mathematical model as the ocean data assimilation result.

[0115] In the above-mentioned ocean data assimilation system based on high-performance parallel optimization, the pre-assimilation module includes:

[0116] A background field data reading module, which equally distributes the background field data to multiple object storages OST, sets multiple reading groups group composed of computing nodes, and sets a reading process for each reading group to perform parallel reading on the object storage OST responsible for it. After the reading is completed, the data is distributed to each reading group group, and each group performs subsequent data assimilation complexity calculation work after receiving the data.

[0117] In the above-mentioned ocean data assimilation system based on high-performance parallel optimization, the load balancing module includes:

[0118] Set multiple reading processes to communicate with the object storage OST to obtain corresponding data; each reading process is located in a different computing node respectively, independently uses its own communication channel to interact with its corresponding OST, and does not affect each other with other reading processes to achieve parallel OST interaction;

[0119] After the reading process obtains the corresponding data, the computing processes are divided into split groups based on nodes, where the 0th process of each computing node is the master of this spilt group; the 0th process is responsible for sending and receiving data between nodes and exclusively occupies the communication channel of the current node.

[0120] Group the grid points in the ocean grid distribution map according to a preset longitude range.

[0121] Horizontally cut the ocean data to be assimilated according to the longitude range, and divide it into n groups named group; according to the number of groups, divide all master processes into n master groups. Each corresponding master group corresponds to a data block. Configure multiple reading processes for each group, and let all reading processes read the corresponding data blocks of the ocean data in parallel and distribute them to the corresponding master processes in parallel.

[0122] In the described ocean data assimilation system based on high-performance parallel optimization, the master processes within each group named group are further evenly grouped to generate scatter_master groups. Initially, each reading process corresponds to a scatter_matser group, which is used to distribute data to the master processes within the scatter_matser group. After one round of distribution, adjust the corresponding relationship between the reading processes and the scatter_matser groups until all data is distributed.

[0123] The described ocean data assimilation system based on high-performance parallel optimization further includes using the ocean data assimilation result to generate an ocean meteorological prediction map at a specified time point.

[0124] The present invention also proposes a storage medium for storing a program for executing any one of the ocean data assimilation methods based on high-performance parallel optimization.

[0125] The present invention also proposes a client for any one of the ocean data assimilation systems based on high-performance parallel optimization.

Claims

1. An ocean data assimilation method based on high-performance parallel optimization, characterized in that, Including: A pre-assimilation step, obtaining the ocean data to be assimilated and a mathematical model, calculating the data assimilation complexity of each ocean grid point respectively according to the background field data and the ocean grid distribution map in the ocean data, so as to generate a calculation topology map based on the data assimilation complexity on the basis of the ocean grid distribution map; A load balancing step, grouping the grid points in the ocean grid distribution map according to a preset longitude range, and counting the overall assimilation complexity corresponding to each group according to the calculation topology map, so as to evenly allocate multiple computing nodes to each group, and evenly divide the computing amount responsible for by each computing node to the computing processes in the computing node; A data assimilation step, after the computing processes in the computing node complete their respective data assimilation tasks, obtaining the assimilated result data, writing the assimilated result data back to the mathematical model as the ocean data assimilation result, and generating an ocean meteorological prediction map at a specified time point using the ocean data assimilation result.

2. The ocean data assimilation method based on high-performance parallel optimization according to claim 1, characterized in that The pre-assimilation step includes: A background field data reading step, equally dividing the background field data into multiple object storage OSTs, setting multiple reading groups group composed of computing nodes, and setting a reading process for each reading group to perform parallel reading on the object storage OST responsible for it, and distributing the data to each reading group group after the reading is completed. After each group group receives the data, it performs subsequent data assimilation complexity calculation work.

3. The ocean data assimilation method based on high-performance parallel optimization as described in claim 1, characterized in that The load balancing step includes: Setting multiple reading processes to communicate with the object storage OST to obtain corresponding data; each reading process is located in a different computing node respectively, independently uses its own communication channel to interact with its corresponding OST, and does not affect each other with other reading processes to achieve parallel OST interaction; After the reading process obtains the corresponding data, dividing the computing processes into split groups by node, where the 0th process of each computing node is the master of this spilt group; the 0th process is responsible for sending and receiving data between nodes and monopolizes the communication channel of the current node; Grouping the grid points in the ocean grid distribution map according to a preset longitude range Horizontally cutting the ocean data to be assimilated according to the longitude range, and dividing it into n groups group in total; according to the number of groups, dividing all master processes into n master groups, each group corresponding master group corresponds to a data block, configuring multiple reading processes for each group, and parallelly reading the corresponding data block of the ocean data by all reading processes and parallelly distributing it to the corresponding master process.

4. The ocean data assimilation method based on high-performance parallel optimization according to claim 3, wherein Dividing the master processes in each group group into scatter_master groups again on average, initially each reading process corresponds to a scatter_matser group for distributing data to the master processes in the scatter_matser group. After one round of distribution is completed, adjusting the corresponding relationship between the reading process and the scatter_matser group until all data is distributed.

5. An ocean data assimilation system based on high-performance parallel optimization, characterized in that, Including: A pre-assimilation module that obtains the ocean data and mathematical model to be assimilated, and calculates the data assimilation complexity of each ocean grid point respectively according to the background field data and the ocean grid distribution map in the ocean data, so as to generate a calculation topology map based on the data assimilation complexity on the basis of the ocean grid distribution map; A load balancing module that groups the grid points in the ocean grid distribution map according to a preset longitude range, and statistically calculates the overall assimilation complexity corresponding to each group according to the calculation topology map, so as to evenly allocate multiple computing nodes to each group, and evenly divide the computing workloads responsible for by each computing node to the computing processes within the computing node; A data assimilation module that, after the computing processes within the computing node complete their respective data assimilation tasks, obtains the assimilated result data, writes the assimilated result data back to the mathematical model as the ocean data assimilation result, and generates an ocean meteorological prediction map for a specified time point using the ocean data assimilation result.

6. The ocean data assimilation system based on high-performance parallel optimization according to claim 5, characterized in that The pre-assimilation module includes: A background field data reading module that equally distributes the background field data to multiple object storage OSTs, sets multiple reading groups group composed of computing nodes, and sets a reading process for each reading group to perform parallel reading on the object storage OST it is responsible for. After the reading is completed, the data is distributed to each reading group group, and each group performs subsequent data assimilation complexity calculation work after receiving the data.

7. The ocean data assimilation system based on high-performance parallel optimization according to claim 5, characterized in that, The load balancing module includes: Setting multiple reading processes to communicate with the object storage OST to obtain corresponding data; each reading process is located in a different computing node respectively, independently uses its own communication channel to interact with its corresponding OST, and does not affect each other with other reading processes to achieve parallel OST interaction; After the reading process obtains the corresponding data, the computing processes are divided into split groups in units of nodes, and the 0th process of each computing node is the master of this spilt group; the 0th process is responsible for sending and receiving data between nodes and exclusively occupies the communication channel of the current node; Group the grid points in the ocean grid distribution map according to a preset longitude range Horizontally cut the ocean data to be assimilated according to the longitude range, and divide it into n groups group in total; according to the number of groups, divide all master processes into n master groups, and each master group corresponding to each group corresponds to a data block. Configure multiple reading processes for each group, and have all reading processes parallelly read the corresponding data block of the ocean data and parallelly distribute it to the corresponding master process.

8. The ocean data assimilation system based on high-performance parallel optimization according to claim 7, characterized in that Average the master processes within each group group again to generate scatter_master groups. Initially, each reading process corresponds to a scatter_matser group, which is used to distribute data to the master processes within the scatter_matser group. After one round of distribution is completed, adjust the corresponding relationship between the reading process and the scatter_matser group until all data is distributed.

9. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the ocean data assimilation method based on high-performance parallel optimization described in any one of claims 1-4 are implemented.

10. A client for an ocean data assimilation system based on high-performance parallel optimization described in any one of claims 5 to 8.

Citation Information

Patent Citations

  • Population elite distribution cloud collaboration equilibrium method used for feature extraction of electronic medical record

    CN104462853A

  • Unmanned ship global safety path planning method

    CN111412918A