Data assimilation method and device, equipment and storage medium
By calculating the location information of the data to be processed by multiple processes in the ocean numerical mode, reading and optimizing communication in parallel, and improving the algorithm structure, solving the problems of inefficiency and performance waste in the prior art, and achieving an efficient data assimilation process.
Patent Information
- Application Number
- CN202510688901.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing data assimilation technology has problems such as inefficiency, high computational cost, difficult parallelization and waste of performance in the ocean numerical model, especially in high-precision and high-resolution computing environments.
By calculating the location information of observation data and status data to be processed by multiple processes, and using multiple processes to read data in parallel, optimize communication methods and improve algorithm structure, an efficient data assimilation process is achieved.
It significantly improves the execution efficiency and resource utilization of data assimilation, reduces performance waste, and improves the computing parallelism and memory access efficiency.
Smart Images

Figure CN120216524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a data assimilation method, apparatus, device, and storage medium. Background Art
[0002] During the operation of ocean numerical models, data assimilation, as a core link, corrects the simulated state data by integrating observational data, effectively suppressing the accumulation of simulation errors over time iterations. Currently adopted data assimilation techniques include schemes such as Ensemble Kalman Filter (EnKF). Among them, EnKF combines ensemble prediction and Kalman filter techniques. Although it significantly improves the simulation accuracy, the computational cost increases substantially.
[0003] As ocean models develop towards higher accuracy and resolution, the performance of data assimilation has gradually become a bottleneck restricting computational efficiency and scalability. Existing data assimilation software faces multiple challenges: First, it involves large-scale data I / O operations and complex multi-process communication mechanisms, resulting in a sharp increase in storage and transmission overheads, further increasing the burden on the system; Second, the nested structure of weakly dependent calculations and strongly dependent calculations increases the difficulty of parallelization; In addition, the complex algorithm code structure limits the scalability and execution efficiency of the program, leading to significant performance waste and low execution efficiency. Therefore, it is particularly necessary to optimize the performance of key algorithm programs for data assimilation based on domestic E-class supercomputer systems, aiming to solve the problem of low efficiency in the existing technology and make full use of the advantages of new computing architectures to improve the overall performance. However, there are still significant performance waste and inefficient execution problems in the current technical implementation, and it is necessary to further explore effective solutions to overcome these challenges. Ultimately, how to efficiently utilize resources and reduce performance waste has become an urgent technical problem to be solved. Summary of the Invention
[0004] To solve the above technical problems, embodiments of the present disclosure provide a data assimilation method, apparatus, device, and storage medium.
[0005] In a first aspect, embodiments of the present disclosure provide a data assimilation method, including: Calculating first position information of observation data to be processed by multiple processes in a current observation file; Parallelly reading the observation data by multiple processes according to the first position information; Determining second position information of state data to be processed by multiple processes in a current state file, where the current state file is a state estimation file obtained based on a historical observation file; Reading the state data by multiple processes according to the second position information; Performing data assimilation on the state data and the observation data to update the current state file.
[0006] Optionally, after multiple processes read the observation data in parallel according to the first position information, the method further includes: Storing the read observation data into multiple buffers corresponding to multiple processes, where the buffers are used to perform read parsing on the observation data.
[0007] Optionally, reading the status data by multiple processes according to the second position information, including: Determining target observation data to be processed by multiple processes among the multiple observation data included in the current observation file; Determining target status data corresponding to the target observation data among the multiple status data included in the current status file; Reading the target status data by multiple processes using a preset communication method according to the second position information of the target status data, where the preset communication method refers to a method of directly performing data communication between multiple processes.
[0008] Optionally, determining target observation data to be processed by multiple processes among the multiple observation data included in the current observation file, including: Calculating the number of processing rounds of multiple processes according to a preset quantity and the quantity of the multiple observation data included in the current observation file, where the preset quantity is determined based on the number of multiple processes; Determining multiple target observation data to be processed by multiple processes in the current processing round among the multiple observation data.
[0009] Optionally, determining target observation data to be processed by multiple processes among the multiple observation data included in the current observation file, including: Statistical processing time of multiple processes for processing different observation data; Calculating the sum value of the processing time of a preset quantity of observation data, where the sum value represents the total duration of multiple processes for processing a preset quantity of observation data in one round; In the case where the sum value is less than the time threshold, determining the preset quantity of observation data as the target observation data.
[0010] Optionally, after calculating the sum value of the processing time of a preset quantity of observation data, the method further includes: In the case where the sum value is greater than or equal to the time threshold, adjusting the preset quantity until the recalculated sum value is less than the recalculated time threshold.
[0011] Optionally, after determining the second position information of the status data to be processed by multiple processes in the current status file, the method further includes: Constructing a dynamic array according to the second position information, where the dynamic array refers to an index file of the status data in the current status file; Dynamically allocate storage space for the index file according to the set grid information; Read status data by multiple processes according to the second location information, including: Read status data from the second location information in the storage space by multiple processes.
[0012] In a second aspect, an embodiment of the present disclosure provides a data assimilation device, including: A calculation unit for calculating the first location information of the observation data to be processed by multiple processes in the current observation file; A first reading unit for parallelly reading observation data by multiple processes according to the first location information; A determination unit for determining the second location information of the status data to be processed by multiple processes in the current status file, where the current status file is a state estimation file obtained based on a historical observation file; A second reading unit for reading status data by multiple processes according to the second location information; A data assimilation unit for assimilating the status data and the observation data to update the current status file.
[0013] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: A memory; A processor; and A computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method of the first aspect as described above.
[0014] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method of the first aspect as described above are implemented.
[0015] The data assimilation method provided by the present disclosure includes: calculating the first location information of the observation data to be processed by multiple processes in the current observation file; parallelly reading observation data by multiple processes according to the first location information; determining the second location information of the status data to be processed by multiple processes in the current status file, where the current status file is a state estimation file obtained based on a historical observation file; reading status data by multiple processes according to the second location information; assimilating the status data and the observation data to update the current status file. The method provided by the present application can efficiently utilize resources during the data assimilation process and reduce performance waste. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0017] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 Flow chart of an EnKF method provided by an embodiment of the present disclosure; Figure 2 Flow chart of a data assimilation scheme provided by an embodiment of the present disclosure; Figure 3 Schematic diagram of multiple processes parallelly reading data provided by an embodiment of the present disclosure; Figure 4 Schematic diagram of synchronous waiting time provided by an embodiment of the present disclosure; Figure 5 Flow chart of data update provided by an embodiment of the present disclosure; Figure 6 Schematic diagram of the structure of a data assimilation device provided by an embodiment of the present disclosure; Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0019] In order to be able to more clearly understand the above objects, features and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0020] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0021] In the operation of ocean numerical models, data assimilation plays a crucial role. By integrating observational data to correct the model results, it effectively inhibits the trend of model errors increasing with time iteration. Among them, ocean numerical models are tools that use mathematical models and computer technology to simulate and predict ocean physical, chemical, biological and other processes. Data Assimilation is a method that combines observational data with the state data output by numerical models, aiming to generate a state data that can reflect the real situation as accurately as possible. The main data assimilation techniques currently adopted include Ensemble Kalman Filter and Improved Local weighted Ensemble Kalman Filter (LwEnKF). Among them, the EnKF scheme combines the advantages of ensemble prediction and Kalman filter, realizing the quantification of uncertainty; LwEnKF further enhances the processing ability of local outliers by introducing localization technology and weighted algorithm on this basis. Although these two methods have significantly improved the simulation accuracy, they have also greatly increased the computational cost.
[0022] As ocean models develop towards higher accuracy and resolution, the performance of data assimilation has gradually become a bottleneck restricting computational efficiency and scalability. Existing data assimilation software faces multiple challenges: First, it involves large-scale data I / O operations and complex multi-process communication mechanisms, which increase the burden on the system; Second, due to the uneven grid density, the load of Central Processing Unit (CPU) cores is unbalanced, affecting the overall computational efficiency; Third, the phenomenon of nested weak-dependent calculations and strong-dependent calculations is widespread, which makes parallel optimization more complex. In addition, the complex algorithm code structure limits the scalability and execution efficiency of the program. These problems work together, resulting in significant performance waste and low execution efficiency. Therefore, it is particularly necessary to optimize the performance of the key algorithm programs of data assimilation based on domestic E-class supercomputer systems, aiming to solve the problem of low efficiency in existing technologies and make full use of the advantages of new computing architectures to improve the overall performance. However, there are still significant performance waste and inefficient execution problems in the current technical implementation, and it is necessary to further explore effective solutions to overcome these challenges. Finally, how to efficiently utilize resources and reduce performance waste has become an urgent technical problem to be solved.
[0023] Example 1: In view of the above technical problems, the embodiments of the present disclosure provide a data assimilation method, which systematically optimizes data assimilation from multiple dimensions such as I / O performance optimization, communication efficiency improvement, and algorithm improvement, significantly improving the memory access efficiency and computational parallelism of processes, and effectively reducing the synchronization waiting time between processes. The specific details are described in one or more of the following embodiments.
[0024] The data assimilation method provided by the embodiments of the present disclosure is applicable to the data assimilation scenario. This method can be executed by a data assimilation device, which can be implemented in software and / or hardware, and can be integrated into an electronic device. Among them, the electronic device may include, but is not limited to, mobile terminals such as smart phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), portable multimedia players (PMPs), vehicle terminals (such as in-vehicle navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc.
[0025] Figure 1 It is a schematic flowchart of an EnKF method provided by the embodiments of the present disclosure, specifically the code flow of data assimilation software using the ensemble Kalman filter scheme. Among them, the main time-consuming sub-processes are sending and / or receiving observation data, and obtaining the neighboring state points and observation points of this process, resulting in problems of system performance waste and low execution efficiency.
[0026] Figure 2 It is a schematic flowchart of a data assimilation scheme provided by the embodiments of the present disclosure, specifically including the following steps: S201. Calculate the first position information of the observation data to be processed by multiple processes in the current observation file.
[0027] It is understandable that, according to the set distribution logic, the first position information of the observation data to be processed by multiple processes in the current observation file is calculated, that is, the position information of each observation data to be processed by each process in the entire observation file is determined, so as to ensure that each process can read the required observation data. Among them, the set distribution logic refers to the logic of a specific process performing multiple distribution loops and distributing observation data to multiple processes at one time. For example, the logic of extracting a part of the data required by a certain receiving process from the observation file and packing and sending it, and the logic of the certain receiving process assigning the received part of the data to the array it controls. A process is the basic unit for executing computing tasks, with an independent address space and system resources. In a distributed computing environment, multiple processes working in parallel can accelerate data processing. Observation data is a record of the actual environmental parameters obtained by sensors or other measurement devices, which is used to correct the model prediction results and improve the simulation accuracy. The current observation file refers to a file containing all the observation data collected at a specific time point or time period, which is an important part of the input in the data assimilation process. The first position information refers to the starting byte offset or index value of the observation data to be processed in each process in the current observation file, which is used to locate the starting point of reading the observation data and ensure the accuracy of data distribution. It is understandable that the multiple observation data to be processed by the process may be discontinuous in the observation file.
[0028] S202. Multiple processes parallelly read the observation data according to the first position information.
[0029] It is understandable that, on the basis of S201 above, after determining the first position information of at least one observation data to be processed by each process in the entire observation file, multiple processes parallelly read at least one observation data based on this first position information. This way of multiple processes parallelly reading multiple observation data based on the first position information optimizes the I / O operation efficiency, reduces data redundancy and repeated reading, thereby improving the execution efficiency and resource utilization rate of the overall data assimilation. Specifically, the MPI_FILE_read_at interface can be used to let each process simultaneously read the observation data at different positions in the observation file according to the index (the first position information). Among them, the Message Passing Interface (MPI) is a standard protocol for writing parallel programs, which is mainly used to realize communication between processes on a distributed memory system.
[0030] It is understandable that, compared with the prior art where all observed data is serially read through a specific process, the read data is stored in an intermediate array, and then the data is extracted from the intermediate array and sent to other processes sequentially and cyclically. During the reading process of a specific process, other processes need to wait and are in an idle state. This situation not only has a complex process but also has a large performance waste problem. The method provided by the present disclosure for enabling all processes to read data in parallel omits the processes of intermediate array storage, data packaging and distribution, and parsing, simplifies the reading process, and effectively improves the system performance.
[0031] Optionally, after multiple processes parallelly read the observed data according to the first position information, the method further includes: Storing the read observed data into multiple buffers corresponding to the multiple processes, where the buffers are used to perform read parsing on the observed data.
[0032] It is understandable that at least one buffer is set for at least one process. For example, one buffer is set for each process, or one buffer is set for a certain number of processes, or multiple buffers are set for multiple processes. That is, the buffer corresponding to each process and the number of buffers are not limited and can be determined according to the actual storage situation of the buffer and / or the data volume of the observed data to be processed by the process, which is not limited herein. In the following embodiments, taking one buffer set for each process as an example, after a process reads all the observed data of one observed object at a time, it stores the data into the corresponding buffer. The number of observed data read by the process each time is not limited. During the reading and processing process, the process will also perform read parsing on the observed data in the buffer to complete the final acquisition of the observed data. The number of observed data for which read parsing is performed by the process each time is not limited. For example, partial observed data in the buffer is subjected to read parsing each time.
[0033] Exemplarily, refer to Figure 3 , Figure 3 which is a schematic diagram of multiple processes parallelly reading data provided by an embodiment of the present disclosure. The multiple processes are denoted as process 0 to process n. After determining the index of the observed data to be processed by each thread in the observed file or completing the offset calculation, process 0 to process n parallelly read the required observed data based on the index.
[0034] S203. Determine the second position information of the state data to be processed by multiple processes in the current state file.
[0035] Wherein, the current state file is a state estimation file obtained based on the historical observed file.
[0036] Understandably, based on the above S202, the current status file corresponding to the current observation file is obtained. Here, the current observation file and the current status file can be understood as the actual data set and the simulated data set of the same observation object. The current status file includes multiple status data, and the status data can be understood as the simulated data at each status point. The current observation file includes multiple observation data, and the observation data can be understood as the actual data at each observation point. Here, the observation points and the status points can be the same or different. Specifically, the status data represents the status information of the system or model at a certain moment, usually data obtained through a series of processing and estimation based on historical observation data. The current status file can be understood as a file generated after state estimation based on the historical observation file, containing the latest system status information and also serving as a basis for summarizing and predicting the past behavior of the system. The second position information refers to the position information of the status data processed by each process in the entire status file, that is, the starting byte offset or index value in the current status file. For example, in an ocean numerical simulation application, if we want to predict the future ocean state, we first need to perform state estimation based on the historical observation file to generate the current status file. Then, to improve the computing efficiency, this status file will be divided into several parts, and each part is processed by a different process. Determining the second position information of the status data to be processed by each process in this status file helps to optimize I / O operations, reduce unnecessary data transmission, and thus improve the overall processing speed and resource utilization efficiency.
[0037] S204. Read the status data by multiple processes according to the second position information.
[0038] It is understandable that, based on the above S203, for an MPI program, increasing the parallelism is crucial for the running efficiency of the program. Before each process loops to process at least one observed data read, that is, before executing the loop calculation of S205, at least one state data corresponding to at least one observed data is preferentially obtained according to the second position information. That is, the position information is transmitted for the first time, so that the process preferentially reads all the state data and all the observed data to be processed in a processing round, and the inflation information is transmitted for the second time. Since the position information is independent, during the process of finding adjacent points, the position information of the required state points and observed points remains unchanged throughout the calculation process and there is no data dependency. Therefore, the communication method is changed for the first communication, allowing all processes to simultaneously package the position information of the observed points to be processed and the position information of the state points, so as to read all the observed data and all the state data, and all processes communicate with each other in an all-to-all manner. This method does not significantly increase the additional communication overhead. In addition, before looping to process the data, preferentially reading the position index of at least one data can significantly increase the parallelism of the algorithm and indirectly reduce the idle time of the process in synchronous waiting, thus improving the overall computing efficiency. The following is a detailed description through the following steps.
[0039] Optionally, multiple processes read the state data according to the second position information, which can be specifically implemented through the following steps: Among the multiple observed data included in the current observed file, determine the target observed data to be processed by multiple processes; among the multiple state data included in the current state file, determine the target state data corresponding to the target observed data; multiple processes use a preset communication method to read the target state data according to the second position information of the target state data, and the preset communication method refers to a method of directly performing data communication between multiple processes.
[0040] Understandably, after the process stores multiple observed data read in parallel into the buffer, it determines at least one target observed data to be processed among the multiple observed data. That is, the process can process at least a part of the data in each round. The specific determination method is not limited and can be determined based on a certain selection criterion or allocation strategy. For example, a certain number of observed data can be processed each time. Among them, the target observed data refers to at least one observed data selected from the current observed file that is relevant to a specific task or calculation requirement. Subsequently, the target status data related to the target observed data is determined among the multiple status data. That is, for each target observed data point, there is one or more associated status data points, and these status data reflect the status estimation results of the system at the same time point or a certain time period. Subsequently, multiple processes use a preset communication method to read the target status data according to the second position information of the target status data. The preset communication method refers to a data exchange mechanism directly carried out among multiple processes. For example, the all-to-all method mentioned above. In this mode, a process can directly communicate with other processes without passing through other nodes. This method not only improves the data processing efficiency but also ensures the consistency and accuracy of the data, providing strong support for large-scale parallel computing.
[0041] Optionally, among the multiple observed data included in the current observed file, the target observed data to be processed by multiple processes is determined, which can be specifically implemented through the following steps: According to the preset quantity and the quantity of the multiple observed data included in the current observed file, calculate the number of processing rounds of multiple processes. The preset quantity is determined based on the number of multiple processes; among the multiple observed data, determine the multiple target observed data to be processed by multiple processes in the current processing round.
[0042] Understandably, considering that the scale of the observed data may be very large, processing all the data at once may lead to excessive memory overhead (equivalent to each process needing to store all the observed data). Based on this, according to the preset quantity and the actual quantity of the observed data, calculate the number of processing rounds of multiple processes, that is, divide the processing task into multiple rounds. For example, a process processes 1000 observed data at a time. Specifically, the preset quantity can be an integer multiple of the total number of processes. Among them, the total number of processes is n, and 2n observed data can be processed in each round. The 2n observed data can be evenly divided among n processes or dynamically allocated according to the processing efficiency of each process. The specific allocation method is not limited. This multi-round processing method controls the additional memory overhead within a small range while ensuring the parallel efficiency. Subsequently, when determining the target observed data, the total number of processing rounds can also be considered. For example, the observed data to be processed is evenly divided in each round.
[0043] Optionally, among the multiple pieces of observation data included in the current observation file, determine the target observation data to be processed by multiple processes. Specifically, this can be achieved through the following steps: Statistically analyze the processing times of multiple processes for different pieces of observation data; calculate the sum of the processing times of a preset number of pieces of observation data, where the sum represents the total duration for multiple processes to complete the processing of the preset number of pieces of observation data in one round; in the case where the sum is less than the time threshold, determine the preset number of pieces of observation data as the target observation data; or, in the case where the sum is greater than or equal to the time threshold, adjust the preset number until the recalculated sum is less than the recalculated time threshold.
[0044] It can be understood that for the problem that there may be at least one process with hidden waiting time in the above-mentioned round-by-round processing mechanism, when determining the target observation data round by round, it is necessary to consider the waiting times of different processes. Among them, the length of the process waiting time is closely related to the amount of data processed in each round. Since there are differences in the workload of each process for processing each piece of observation data, the total processing duration of one round is equivalent to the time for the slowest process to complete all the observation data. Other processes need to wait for the slowest process to complete the processing of all its corresponding observation data, that is, other processes have synchronous waiting times. Therefore, when a process processes multiple pieces of observation data at a time, the synchronous waiting time depends on the slowest process that completes the processing of all the observation data in one round. Specifically, when the memory resources are sufficient, statistically analyze the time required for each process to process different pieces of observation data, that is, record the specific time-consuming situation of each process for processing a specific piece of observation data. Subsequently, based on the above-mentioned statistically analyzed processing times, calculate the total time required for all processes to process a preset number of pieces of observation data (i.e., the "sum"), where the sum represents the total duration for all processes to complete the processing task of the specified number of pieces of observation data in one round.
[0045] Exemplarily, see Figure 4 , Figure 4A schematic diagram of the synchronization waiting time provided by an embodiment of the present disclosure. The processing time of process 0 for processing a single observation data A (Observation A) is 10 s, the processing time of process 1 for processing a single observation data A is 4 s, the processing time of process 0 for processing a single observation data B (Observation B) is 5 s, and the processing time of process 1 for processing a single observation data B is 9 s. If each process processes only one observation data in one round, when processing Observation A in the first round, process 0 is the slowest process, and process 1 needs to wait for 6 s. The processing duration of the first round is 10 s. When processing Observation B in the second round, process 1 is the slowest process, and process 0 needs to wait for 4 s. The processing duration of the first round is 9 s. The total processing duration of the two rounds is 19 s. Not only does each process need to wait for a relatively long time, but the total processing duration of multiple rounds is also relatively long. In this case, the processing duration of each observation data is preferentially counted, and processing multiple observation data at one time is considered to reduce the synchronization waiting duration. For example, Figure 5 as shown, the total duration of process 0 for processing 2 observation data (Observation A and Observation B) is 15 s (i.e., "10 s + 5 s"), the total duration of process 1 for processing 2 observation data is 13 s (i.e., "4 s + 9 s"), and process 0 is the slowest process. When completing this processing round, process 1 only needs to wait for 2 s, that is, wait for process 0 to finish processing the data, and the synchronization waiting time is 2 s. The way that each process processes Observation A and Observation B simultaneously is 4 s faster than the way that each process processes one observation data in turn. Therefore, when the memory resources are sufficient, processing more data in each round can hide more synchronization waiting time, ensure that the system can efficiently complete tasks within the specified time, maximize the resource utilization rate at the same time, and thus significantly improve the overall performance. This method achieves a balance between memory overhead and computing efficiency, avoiding excessive occupation of memory resources and reducing the synchronization waiting time by increasing the amount of data processed in each round, and finally achieving higher parallel efficiency and better performance.
[0046] Understandably, after processing more data in each round to reduce the hidden waiting time, the calculated sum value can also be compared with the time threshold, that is, the maximum processing time acceptable to the system can be limited. The time threshold is the maximum processing time that the system can accept under an ideal state and can be determined according to user requirements. If the sum value is less than the time threshold, it is considered that the currently preset number of observation data can be used as target observation data, that is, these data can be effectively processed by all processes without exceeding the expected time. If the sum value is greater than or equal to the time threshold, it indicates that the processing efficiency of all processes under the current configuration fails to meet the requirements, and the preset number needs to be adjusted. That is, when the sum value does not meet the conditions, the total processing duration can be attempted to be reduced by reducing the preset number of observation data until the recalculated sum value is lower than the time threshold. This process may need to be iterated repeatedly until the optimal amount of observation data is found, so that the processing time not only meets the efficiency requirements but also is as close as possible to but does not exceed the time threshold.
[0047] Optionally, after determining the second position information of the status data to be processed by multiple processes in the current status file, the method further includes: Construct a dynamic array according to the second position information, where the dynamic array refers to the index file of the status data in the current status file; dynamically allocate storage space for the index file according to the set grid information; read the status data by multiple processes according to the second position information, including: reading the status data from the second position information in the storage space by multiple processes.
[0048] It is understandable that the arrays for storing the neighbor point indices and distances are declared to be the size of the total number of in-process state elements. However, this approach results in significant space waste. Due to the limitation of the neighbor radius, an observation point only has neighbor points in a limited number of grids. Therefore, such a large intermediate array is not necessary at all. To address this issue, the present disclosure constructs a dynamic array based on the second position information. The dynamic array records the specific position information of the state data in the current state file, facilitating quick positioning and access. That is to say, the intermediate array is changed to a dynamic array, and when initializing the grid, the storage space is dynamically allocated for the index file according to the set grid information. That is, the memory resources are reasonably arranged to ensure that the index file can efficiently support subsequent data access operations. Herein, the set grid information refers to the spatial partitioning information of the model. For example, the grid partitioning in an ocean numerical model determines how to organize and manage the state data. Subsequently, the process can read the state data based on the second position information in this storage space, thereby avoiding unnecessary I / O operations and potential data redundancy problems. This way of dynamically allocating space for the neighbor point indices not only significantly reduces the space consumption but also improves the memory access efficiency. Additionally, since the size of the intermediate array decreases exponentially, there is no need to frequently replace the cache when accessing the array storing the index file, thus greatly improving the cache hit rate and further achieving a faster memory access speed.
[0049] S205. Perform data assimilation on the state data and the observation data to update the current state file.
[0050] It is understandable that, based on the above S204, data assimilation is performed according to the first position information, the second position information, the state data, and the observation data. Herein, data assimilation is a method of integrating the observation data into the numerical model, aiming to improve the accuracy of the model prediction. By considering the error between the observation data and its uncertainty and the state data predicted by the model, an optimized state estimate is provided. Specifically, the difference and its uncertainty between the model prediction and the actual observation data can be evaluated. Based on the observation data and its error, the state data of the model is adjusted to obtain a state estimate closer to the real situation. Subsequently, the updated state data is written back to the current state file to replace the original state estimate, ensuring that subsequent simulations and predictions are based on the latest and most accurate state information.
[0051] Exemplarily, refer to Figure 5 , Figure 5 which is a schematic flowchart of a data update provided by an embodiment of the present disclosure, specifically including the following steps: 1) After reading the observation data and status data, each process exchanges location information in an all-to-all manner, where the location information refers to the location information of the observation points and status points. 2) Based on the location information, find the neighboring observation objects (observation points) and neighboring status points. 3) Determine whether the current observation point (observation element) is local. 4) If the current observation point is local, calculate the inflation information and broadcast it to all processes through MPI_BCAST. 5) If the current observation point is not local, wait to receive the inflation information of the current observation point. 6) After all processes have obtained the inflation information, update the status points and observation points, and loop to process the next observation point.
[0052] It can be understood that each process will loop through multiple observation points in one processing round. However, before looping through the observation points, it will first obtain the location information of all observation points and all status points. Each loop will only process one observation point (i.e., the current observation point), and there is no need to obtain the location information of this current observation point and its neighboring status points again.
[0053] A data assimilation method provided by an embodiment of the present disclosure, in terms of I / O optimization, by calculating the location information of the observation data in the observation file, realizes a parallel optimization scheme based on MPI-IO for multiple processes, significantly improves the running efficiency of the observation data input module, and basically solves the performance bottleneck problem in the data reading process. In terms of communication optimization, by adopting strategies such as data pre-packaging, optimizing the MPI communication method, and reducing the synchronous waiting time, the communication overhead is greatly reduced. In terms of algorithm optimization, by separating the calculation parts without data dependencies and improving the neighboring array reuse mechanism and other measures, the calculation efficiency is significantly improved. In addition, according to the data characteristics and task division characteristics, dynamically adjust the size of the intermediate array, which not only reduces the space complexity but also further improves the memory access efficiency. It not only brings significant performance improvement to the data assimilation software, but also lays an important foundation for realizing a higher-precision and higher-resolution ocean prediction model.
[0054] Embodiment 2: Figure 6 It is a schematic structural diagram of a data assimilation device provided by an embodiment of the present disclosure. The data assimilation device provided by an embodiment of the present disclosure can execute the processing flow provided by the embodiment of the data assimilation method, as Figure 6 shown, the device 600 includes: A calculation unit 601, configured to calculate the first location information of the observation data to be processed by multiple processes in the current observation file; A first reading unit 602, configured to parallelly read the observation data by multiple processes according to the first location information; A determination unit 603, configured to determine second position information of state data to be processed by multiple processes in a current state file, where the current state file is a state estimation file obtained based on a historical observation file; A second reading unit 604, configured to read state data by multiple processes according to the second position information; A data assimilation unit 605, configured to perform data assimilation on the state data and the observation data to update the current state file.
[0055] Optionally, the apparatus 600 is further configured to: Store the read observation data into multiple buffers corresponding to multiple processes, where the buffers are used to perform read parsing on the observation data.
[0056] Optionally, the second reading unit 604 is configured to: Determine target observation data to be processed by multiple processes from multiple observation data included in the current observation file; Determine target state data corresponding to the target observation data from multiple state data included in the current state file; Read the target state data by multiple processes using a preset communication method according to the second position information of the target state data, where the preset communication method refers to a method of directly performing data communication between multiple processes.
[0057] Optionally, the second reading unit 604 is configured to: Calculate the number of processing rounds of multiple processes according to a preset quantity and the number of multiple observation data included in the current observation file, where the preset quantity is determined based on the number of multiple processes; Determine multiple target observation data to be processed by multiple processes in the current processing round from multiple observation data.
[0058] Optionally, the second reading unit 604 is configured to: Statistically analyze the processing time of multiple processes for processing different observation data; Calculate the sum value of the processing times of a preset quantity of observation data, where the sum value represents the total duration of multiple processes for processing a preset quantity of observation data in one round; In a case where the sum value is less than a time threshold, determine the preset quantity of observation data as the target observation data.
[0059] Optionally, the second reading unit 604 is configured to: In a case where the sum value is greater than or equal to the time threshold, adjust the preset quantity until the recalculated sum value is less than the recalculated time threshold.
[0060] Optionally, the apparatus 600 is further configured to: Construct a dynamic array according to the second position information, where the dynamic array refers to the index file of the state data in the current state file; Dynamically allocate storage space for the index file according to the set grid information; Optionally, the second reading unit 604 is further configured to: Read the state data from the second position information in the storage space through multiple processes.
[0061] Figure 6 The data assimilation device of the illustrated embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0062] Embodiment 3: Figure 7 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specifically refer to Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the present disclosure. The electronic device 700 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 7 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0063] As Figure 7 shown, the electronic device 700 may include a processing device 701 (such as a central processing unit, a graphics processing unit, etc.), which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703 to implement the data assimilation method of the embodiments as described in the present disclosure. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0064] Typically, the following devices may be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.
[0065] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart, so as to implement the data assimilation method as described above. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0066] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0067] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0068] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device.
[0069] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also perform the other steps described in the above embodiments.
[0070] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0071] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0072] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.
[0073] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0074] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0075] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "including a data assimilation" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0076] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data assimilation method, characterized in that, Including: Calculating the first position information of the observation data to be processed by multiple processes in the current observation file; Parallelly reading the observation data by the multiple processes according to the first position information; Determining the second position information of the status data to be processed by the multiple processes in the current status file, where the current status file is a status estimation file obtained based on a historical observation file; Reading the status data by the multiple processes according to the second position information; Performing data assimilation on the status data and the observation data to update the current status file.
2. The method according to claim 1, characterized in that, After the step of parallelly reading the observation data by the multiple processes according to the first position information, the method further includes: Storing the read observation data into multiple buffers corresponding to the multiple processes, where the buffers are used for reading and parsing the observation data.
3. The method according to claim 1, characterized in that, The step of reading the status data by the multiple processes according to the second position information includes: Determining, among the multiple observation data included in the current observation file, the target observation data to be processed by the multiple processes; Determining, among the multiple status data included in the current status file, the target status data corresponding to the target observation data; Reading the target status data by the multiple processes in a preset communication manner according to the second position information of the target status data, where the preset communication manner refers to a manner of directly performing data communication between the multiple processes.
4. The method according to claim 3, characterized in that, The step of determining, among the multiple observation data included in the current observation file, the target observation data to be processed by the multiple processes includes: Calculating the processing rounds of the multiple processes according to a preset quantity and the quantity of the multiple observation data included in the current observation file, where the preset quantity is determined based on the quantity of the multiple processes; Determining, among the multiple observation data, the multiple target observation data to be processed by the multiple processes in the current processing round.
5. The method according to claim 3, characterized in that, The step of determining, among the multiple observation data included in the current observation file, the target observation data to be processed by the multiple processes includes: Statistical processing times of the multiple processes for processing different observation data; Calculating the sum value of the processing times of a preset quantity of observation data, where the sum value represents the total duration for the multiple processes to process the preset quantity of observation data in one round; When the sum value is less than a time threshold, determining the preset quantity of observation data as the target observation data.
6. The method according to claim 5, characterized in that, After calculating the sum value of the processing times of the preset quantity of observation data, the method further includes: When the sum value is greater than or equal to the time threshold, adjusting the preset quantity until the recalculated sum value is less than the recalculated time threshold.
7. The method according to claim 3, characterized in that After determining the second position information of the status data to be processed by the multiple processes in the current status file, the method further includes: Constructing a dynamic array according to the second position information, where the dynamic array is an index file of the status data in the current status file; Dynamically allocating storage space for the index file according to set grid information; The step of reading the status data by the multiple processes according to the second position information includes: Read status data from the second location information in the storage space through the multiple processes.
8. A data assimilation device, characterized in that, Comprising: A calculation unit for calculating the first location information of the observation data to be processed by the multiple processes in the current observation file; A first reading unit for parallelly reading the observation data according to the first location information through the multiple processes; A determination unit for determining the second location information of the status data to be processed by the multiple processes in the current status file, where the current status file is a state estimation file obtained based on historical observation files; A second reading unit for reading the status data according to the second location information through the multiple processes; A data assimilation unit for assimilating the status data and the observation data to update the current status file.
9. An electronic device, characterized in that, Comprising: A memory; A processor; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the data assimilation method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the data assimilation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Large-scale video monitoring storage method
CN106231252A
DNA data file reading method and computer readable storage medium
CN107169313A
Parallel sequence alignment method and device based on load balancing and computer equipment
CN112764922A
Ocean data assimilation method and system based on high-performance parallel optimization
CN114546638A
Land surface data assimilation method and device for numerical weather forecast and application
CN118195343A