Chlorophyll-a concentration fusion calculation system and method based on deep learning

CN122549301BActive Publication Date: 2026-09-08OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611038735.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-08
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种基于深度学习的叶绿素a浓度融合计算系统及计算方法,用于解决现有叶绿素a浓度计算算法存在的无法形成物理模式预报→深度学习修正→同化反馈→模式再预报的闭环迭代,EnKF同化方法计算效率低、参数敏感、非线性适应能力不足的问题

Benefits of technology

[0045] The beneficial effects of this invention are as follows: This invention achieves closed-loop online fusion of numerical models and deep learning models. Compared to the existing offline post-processing correction mode of HBGC-CNN, it feeds back the inference results of the deep learning model to the state variable update of the FVCOM model in the form of analysis increments through a data assimilation framework, forming a closed-loop iteration of FVCOM prediction → CNN inference → assimilation calculation → incremental feedback → model re-prediction. This makes the deep learning model no longer an external corrector of the numerical model, but an online assimilator embedded in the model's running loop, which can continuously improve the initial field and subsequent prediction performance of the model. Compared to the complex implementation of traditional EnKF assimilation, which requires generating a large number of set members and modifying the model's source code, this invention directly accesses the state variables through the NetCDF average data linked list pointer, requiring only one additional independent module and about 20-30 lines of main program interface code, significantly reducing the development cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549301B_ABST
    Figure CN122549301B_ABST
Patent Text Reader

Abstract

The present application discloses a chlorophyll-a concentration fusion calculation system and method based on deep learning, and relates to the field of chlorophyll-a concentration calculation. The present application solves the problem that the physical model and deep learning are mutually isolated in the existing chlorophyll-a concentration prediction, and the assimilation operation cannot continuously constrain the model state. The calculation system of the present application comprises a numerical model process cluster module, an artificial intelligence reasoning module and a message queue communication module. The calculation method comprises the following steps: FVCOM model initialization and assimilation parameter configuration, FVCOM main loop and assimilation trigger judgment, model state variable reading and cross-process data transmission, data processing and assimilation calculation of the reasoning process, and assimilation increment return and model state update. The present application realizes the closed-loop online fusion of the numerical model and the deep learning model, and forms a closed-loop iteration of prediction, reasoning, assimilation calculation, increment feedback and model re-prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chlorophyll a concentration calculation technology, specifically a chlorophyll a concentration fusion calculation system and calculation method based on deep learning. Background Technology

[0002] With the rapid development of artificial intelligence technology, deep learning models are widely used in ocean parameter forecasting and inversion. Convolutional Neural Networks (CNNs) can automatically extract spatial features from multi-channel remote sensing data, while Long Short-Term Memory Networks (LSTMs) can capture long-term dependencies in time series data; both have achieved good results in ocean chlorophyll a concentration forecasting. However, these purely data-driven models lack physical consistency guarantees, exhibit poor generalization performance in sparse data regions and extreme environmental conditions, and cannot provide feedback corrections to the state variables of the numerical model itself.

[0003] To overcome the limitations of purely data-driven methods, patent application number 2022101644231, entitled "A Method for Improving the Accuracy of Marine Chlorophyll a Concentration Forecasting Based on Machine Learning," proposes an HBGC-CNN hybrid model. This model uses environmental variables such as temperature, salinity, dissolved inorganic nitrogen, dissolved organic phosphorus, and zooplankton output from the FVCOM-HBGC model as input features, and satellite remote sensing chlorophyll a products as the target variable. A CNN network learns the nonlinear mapping relationship between these two variables, achieving effective prediction of the daily spatial distribution of chlorophyll a in the Bohai Sea. However, this approach is essentially an offline post-processing operation. The CNN model only corrects errors after the numerical model completes its integration output; the correction results are neither fed back to the initial field of the numerical model nor participate in the next iteration of the numerical model's forecast. This open-loop architecture isolates the physical processes of the numerical model from the data-driven learning process, preventing the formation of a closed-loop iteration of model forecasting → deep learning correction → assimilation feedback → model re-forecasting.

[0004] Data assimilation (DA) technology organically integrates observational data with the simulation results of numerical models through mathematical optimization methods, serving as a crucial bridge connecting observational data and numerical models. However, traditional sequential assimilation methods, represented by ensemble Kalman filtering (EnKF), essentially perform optimal estimation and correction of the model's initial field. After the model enters the forecast integration process, it lacks the ability to continuously assimilate newly arriving observational data into the model state. Li Xianghao et al. (2024) applied EnKF to the numerical simulation of low-frequency fluctuations in the Bohai Sea level (Oceania and Limnology, 55(5), 1070-1081). Although they successfully reproduced the subtidal level changes caused by wind relaxation, they also revealed the inherent engineering bottlenecks of traditional EnKF: it requires the generation and propagation of a large number of ensemble members, resulting in high computational costs. It is also highly sensitive to assimilation parameters, and parameter tuning is time-consuming and prone to filter divergence. Furthermore, its background error covariance estimation relies on the linear approximation assumption, making it difficult to accurately capture the strong nonlinear biogeochemical processes in marine ecosystems. More importantly, the fact that assimilation only applies to the initial field means that once the model begins to integrate forward, subsequent observation data can no longer be incorporated into the system, and the forecast quality can only depend on the assimilation effect at the initial moment.

[0005] In summary, the core problem that the existing technology has not yet solved is: how to embed the prediction results of deep learning models into the running loop of numerical models in an online manner, so as to realize the real-time closed-loop fusion between physical models and data-driven models, thereby improving the prediction ability of numerical models for chlorophyll a and related ecological variables during the continuous assimilation process. Summary of the Invention

[0006] The purpose of this invention is to provide a deep learning-based chlorophyll a concentration fusion calculation system and method to solve the problems of existing chlorophyll a concentration calculation algorithms, such as the inability to form a closed-loop iteration of physical model prediction → deep learning correction → assimilation feedback → model re-prediction, and the low computational efficiency, parameter sensitivity, and insufficient nonlinear adaptability of the EnKF assimilation method.

[0007] The technical solution adopted by this invention to solve its technical problem is: a deep learning-based chlorophyll a concentration fusion calculation system, including a numerical model process cluster module, an artificial intelligence inference module, and a message queue communication module;

[0008] The numerical model process cluster module is based on the hydrodynamic-water quality coupled model FVCOM-HBGC, and includes a main process and multiple sub-processes; wherein, FVCOM is a finite volume ocean model, and HBGC is a hydrodynamic-biogeochemical coupled model; the main process is responsible for managing the control of the assimilation process, which includes assimilation time determination, state variable reading, data transmission, incremental reception, and broadcasting; the sub-processes are responsible for the hydrodynamic-water quality integral calculation of the FVCOM model.

[0009] The artificial intelligence inference module includes a data preprocessing module, a deep learning inference module, a data assimilation calculation module, and a data post-processing module. The data preprocessing module receives state variables sent by the FVCOM main process and performs standardization processing and patch extraction. The deep learning inference module performs chlorophyll a concentration inference prediction based on a pre-trained HBGC-CNN model, where CNN is a convolutional neural network. The data assimilation calculation module calculates and analyzes incremental vectors according to the configured assimilation algorithm. The data post-processing module serializes the analyzed incremental vectors and returns them to the FVCOM main process.

[0010] The message queue communication module is based on the ZeroMQ message middleware library for message processing, enabling bidirectional data exchange between the numerical model process cluster module and the artificial intelligence inference module.

[0011] Furthermore, the main process traverses the NetCDF average data linked list structure within the FVCOM pattern and locks the memory pointer of the target variable based on the variable name, achieving zero-intrusion state variable reading. After the read state variable is serialized, it is sent to the AI ​​inference process through the ZeroMQ request / response channel. After the AI ​​inference process completes the inference and assimilation calculations, it serializes the analysis increment vector and returns it to the main process through the same channel. The main process broadcasts the increment vector to all child processes through the broadcast function MPI_Bcast, and the child processes apply the increment to their respective grid cells to update the model state variables.

[0012] The present invention also provides a deep learning-based method for calculating chlorophyll a concentration fusion, comprising the following steps.

[0013] S1. FVCOM mode initialization and assimilation parameter configuration.

[0014] S2.FVCOM main loop and assimilation trigger judgment.

[0015] In S2.1.FVCOM mode, the main time loop is entered, and hydrodynamic and water quality calculations are performed at each time step.

[0016] S2.2. At the end of each time step, the main process checks whether the current simulation time has reached the predetermined assimilation time.

[0017] S3. Reading mode state variables and cross-process data transfer.

[0018] S3.1. When the assimilation condition is met, the main process calls the data post-processing module, which locks the memory pointer of the target variable by traversing the NetCDF average data linked list structure inside the FVCOM mode and according to the variable name.

[0019] S3.2. The main process serializes the read state variable data into a binary stream.

[0020] S3.3. The main process sends serialized data to the AI ​​inference process through the ZeroMQ request channel and then blocks to wait for an acknowledgment response.

[0021] S4. Data processing and assimilation computation in the AI ​​inference process.

[0022] S4.1. The AI ​​inference process receives serialized data sent by the main process, deserializes it, and restores it to the NumPy array format in Python.

[0023] S4.2. Data standardization processing.

[0024] Standardize all state variables;

[0025] S4.3. Deep learning model inference.

[0026] The standardized state variables are input into a pre-trained deep learning model. The deep learning model extracts the spatial pattern of multi-channel input features through convolutional layers, and outputs the estimated value of the target variable after mapping through fully connected layers.

[0027] S4.4. Data assimilation computation.

[0028] Incremental analysis is used to update the IAU as an assimilation scheme, and the optimal magnitude of the analysis increment is calculated.

[0029] S4.5. Incremental sustained-release application.

[0030] The IAU method is used to gradually inject the analytical increments in a linear, sustained-release manner across multiple time steps.

[0031] S5. Assimilation Increment Return and Pattern State Update.

[0032] S5.1. The AI ​​inference process serializes the calculated analysis increment vector and returns it to the main process through the ZeroMQ response channel.

[0033] S5.2. After receiving the analysis increment vector, the main process broadcasts the analysis increment vector to all child processes through the broadcast function MPI_Bcast.

[0034] S5.3. Each subprocess applies the analysis increment vector to the grid cell it is responsible for.

[0035] S5.4. The updated state variables are used as the initial conditions for the next time step, and the physical-biochemical integral calculation of the FVCOM mode continues; this completes a full online assimilation loop.

[0036] Further, step S1 includes the following sub-steps: S1.1. Read the mesh file, initial field file, and boundary condition file of the FVCOM-HBGC mode to complete mode initialization; S1.2. Read the assimilation configuration file and parse the parameters assimilation switch DA_ON, assimilation time interval string DA_INTERVAL_STRING, and AI inference service address DA_SERVER_ADDR; S1.3. Read the background error standard deviation from the static file. and observation error standard deviation .

[0037] Furthermore, in step S1.3, The RMSE was calculated using historical simulation data from the HBGC model and observation data from the OC-CCI satellite. The RMSE was calculated using the data reconstructed by the HBGC-CNN model and the test set data of OC-CCI.

[0038] Furthermore, in step S2.1, the hydrodynamic calculation includes a three-dimensional transport solution for temperature, salinity, and flow velocity, and the water quality calculation includes phytoplankton growth, nutrient cycling, and zooplankton predation biochemical processes.

[0039] Furthermore, in step S2.2, the method for determining whether the current simulation time has reached the predetermined assimilation time is as follows: parse DA_INTERVAL_STRING to obtain the assimilation interval in seconds. If the previous simulation time - the last assimilation time ≥ the interval, then the assimilation process is triggered.

[0040] Furthermore, the variables involved in step S3.1 include: temperature T, salinity S, total phytoplankton PPT, dissolved inorganic nitrogen DIN, total zooplankton ZPT, and dissolved organic nitrogen DON.

[0041] Further, in step S4.2, the standardized formula is: ; where X i This represents the original data; μ is the mean of the state variable on the training set; σ is the standard deviation of the state variable on the training set; X i * This represents the standardized data.

[0042] Further, in step S4.4, the calculation of the analysis increment includes: S4.4.1. Calculating the deviation between the background field and the observation field. ;in, For the observation field; For observation operators; For the mode background field; S4.4.2. Calculate the Kalman gain matrix. ;in, The background error covariance matrix; The observation error covariance matrix; matrix From the standard deviation of background error Spatial correlation function structure: Where B(i,j) are elements in matrix B; For grid points With observation point The distance between them; For spatially relevant scales; The expression is: ;in, S4.4.3. Calculate the IAU analysis increment. ;in, These are the coefficients of the constraint terms.

[0043] Furthermore, in step S4.5, the incremental component applied at each time step ;in, For a time window of delayed release; This is the time step of the pattern.

[0044] Furthermore, in step S5, the application of the analysis increment vector to the respective grid cell by each subprocess refers to: calculating the forecast state of the FVCOM model at the current moment. ;in, This represents the current forecast status of the FVCOM mode. The updated state variables are used as initial conditions for the next time step to continue the integration process.

[0045] The beneficial effects of this invention are as follows: This invention achieves closed-loop online fusion of numerical models and deep learning models. Compared to the existing offline post-processing correction mode of HBGC-CNN, it feeds back the inference results of the deep learning model to the state variable update of the FVCOM model in the form of analysis increments through a data assimilation framework, forming a closed-loop iteration of FVCOM prediction → CNN inference → assimilation calculation → incremental feedback → model re-prediction. This makes the deep learning model no longer an external corrector of the numerical model, but an online assimilator embedded in the model's running loop, which can continuously improve the initial field and subsequent prediction performance of the model. Compared to the complex implementation of traditional EnKF assimilation, which requires generating a large number of set members and modifying the model's source code, this invention directly accesses the state variables through the NetCDF average data linked list pointer, requiring only one additional independent module and about 20-30 lines of main program interface code, significantly reducing the development cycle. Attached Figure Description

[0046] Figure 1 This is a system architecture diagram of the present invention;

[0047] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0048] like Figure 1 As shown, the computing system of the present invention includes a numerical model process cluster module, an artificial intelligence inference module, and a message queue communication module. The numerical model process cluster module is based on the hydrodynamic-water quality coupling model FVCOM-HBGC and includes a main process and multiple sub-processes. Among them, FVCOM is a finite volume ocean model, and HBGC is a hydrodynamic-biogeochemical coupling model. The main process of the numerical model process cluster module is responsible for managing the control of the assimilation process, which includes assimilation time determination, state variable reading, data sending, incremental reception, and broadcasting. The sub-processes of the numerical model process cluster module are responsible for the hydrodynamic-water quality integral calculation of the FVCOM model. The artificial intelligence inference module is independent of the Python programming language service process running in the numerical model process and is deployed on the GPU server. The artificial intelligence inference module includes the following sub-modules: (1) Data preprocessing module: receives the state variables sent by the FVCOM main process, performs standardization processing and patch extraction; (2) Deep learning inference module: performs chlorophyll a concentration inference prediction based on the pre-trained HBGC-CNN model; where CNN is a convolutional neural network. (3) Data assimilation calculation module: Calculates the analysis increment vector according to the configured assimilation algorithm; the assimilation algorithm is such as optimal interpolation OI, Newton relaxation Nudging or incremental analysis to update IAU; (4) Data post-processing module: Serializes the analysis increment vector and returns it to the FVCOM main process. The message queue communication module is based on the message processing queue library ZeroMQ message middleware, and adopts the request / response communication mode to realize bidirectional data exchange between the numerical model process cluster module and the artificial intelligence inference module. The FVCOM main process is the client and the AI ​​inference service is the server. Data is transmitted across languages ​​using binary stream serialization. The connection relationship of each part of the computing system is as follows: The main process of the numerical model process cluster module locks the memory pointer of the target variable according to the variable name by traversing the NetCDF average data linked list structure inside the FVCOM mode to realize zero-intrusion state variable reading. After the read state variable is serialized, it is sent to the AI ​​inference process through the request / response channel of the message processing queue library ZeroMQ. After the AI ​​inference process completes the inference and assimilation calculation, it serializes the analysis increment vector and returns it to the FVCOM main process through the same channel. The main process broadcasts the increment vector to all child processes via the broadcast function MPI_Bcast. The child processes then apply the increment to their respective grid cells to update the model state variables.

[0049] like Figure 2 As shown, based on the computing system of the present invention, the computing method of the present invention includes the following steps.

[0050] S1. FVCOM mode initialization and assimilation parameter configuration.

[0051] S1.1. Read the mesh file, initial field file and boundary condition file of the FVCOM-HBGC mode to complete the mode initialization; the basic configuration of the FVCOM-HBGC mode is: 28388 mesh nodes, 54332 elements, 21 vertical layers, and a horizontal resolution of 300m-3km.

[0052] S1.2. Read the assimilation configuration file and parse the parameters assimilation switch DA_ON, assimilation time interval string DA_INTERVAL_STRING, and AI inference service address DA_SERVER_ADDR.

[0053] The assimilation algorithm type is uniformly controlled by the Python side through the assim_type parameter: 0 = no assimilation, 1 = OI, 2 = Nudging, 3 = IAU. In addition, the deep learning model path and other assimilation algorithm parameters are configured on the Python side, and are managed independently from the NML configuration file on the Fortran programming language side.

[0054] S1.3. Read the background error standard deviation from the static file and observation error standard deviation . The RMSE was calculated using historical simulation data from the HBGC model and observation data from the OC-CCI satellite. The RMSE was calculated using the data reconstructed by the HBGC-CNN model and the test set data of OC-CCI.

[0055] S2.FVCOM main loop and assimilation trigger judgment.

[0056] In S2.1.FVCOM mode, the main time loop begins, and hydrodynamic calculations are performed at each time step, including three-dimensional transport solutions for temperature, salinity, and flow velocity. Water quality calculations are also performed at each time step, including phytoplankton growth, nutrient cycling, and zooplankton predation biochemical processes.

[0057] S2.2. At the end of each time step, the main process checks whether the current simulation time has reached the predetermined assimilation time. The determination method is as follows: parse DA_INTERVAL_STRING to obtain the assimilation interval in seconds. If the previous simulation time - the last assimilation time ≥ the interval, then the assimilation process is triggered. The assimilation frequency can be configured according to actual needs. For example, assimilation can be performed once per hour.

[0058] S3. Reading mode state variables and cross-process data transfer.

[0059] S3.1. When the assimilation conditions are met, the main process calls the data post-processing module, which traverses the NetCDF average data linked list structure within the FVCOM mode and locks the memory pointer of the target variable based on the variable name. The variables involved include: temperature T, salinity S, total phytoplankton PPT, dissolved inorganic nitrogen DIN, total zooplankton ZPT, and dissolved organic nitrogen DON.

[0060] S3.2. The main process serializes the read state variable data into a binary stream.

[0061] S3.3. The main process sends serialized data to the AI ​​inference process through the ZeroMQ request channel and blocks after sending, waiting for an acknowledgment response. Actual testing shows that a single communication session takes less than 0.3 seconds.

[0062] S4. Data processing and assimilation computation in the AI ​​inference process.

[0063] S4.1. The AI ​​inference process receives serialized data sent by the main process, deserializes it, and restores it to the NumPy array format in Python.

[0064] S4.2. Data standardization processing.

[0065] The state variables are standardized to eliminate the influence of different variable units. The standardization formula is as follows: Among them, X i This represents the original data; μ is the mean of the state variable on the training set; σ is the standard deviation of the state variable on the training set; X i * This represents the standardized data.

[0066] S4.3. Deep learning model inference.

[0067] The standardized state variables are input into a pre-trained deep learning model. The deep learning model extracts the spatial pattern of multi-channel input features through convolutional layers, and outputs the estimated value of the target variable after mapping through fully connected layers. This invention does not limit the specific architecture of the deep learning model, only requiring that the model input and the variable list sent from the Fortran terminal are aligned in data type and dimension.

[0068] S4.4. Data assimilation computation.

[0069] Incremental analysis is used to update the IAU as an assimilation scheme. The IAU uses the optimal interpolation formula to calculate the optimal magnitude of the analysis increment, and a constraint term α is introduced to scale the increment magnitude to ensure numerical stability. The calculation of incremental analysis includes the following three steps: S4.4.1. Calculate the deviation between the background field and the observed field: ;in, This refers to the observation field, i.e., the inference values ​​of the HBGC-CNN model; This is the background field of the model, i.e., the FVCOM forecast value; For observation operators. Since the chlorophyll a concentration output by the CNN corresponds one-to-one with the FVCOM grid, and... and All have been defined in the model chlorophyll a space, therefore It is the identity matrix. S4.4.2. Calculate the Kalman gain matrix: ;in, The background error covariance matrix; The observation error covariance matrix is ​​given by... structure; This is the Kalman gain matrix. The matrix is ​​composed of the background error standard deviation Spatial correlation function structure: Where B(i,j) are elements in matrix B; For grid points With observation point The distance between them For spatially relevant scales; For Gaspari-Cohn tight support correlation functions, ,in Function C in Take the maximum value at the location , Increase smoothly and gradually decrease, in The cutoff value is set to zero, ensuring that Matrix sparsity and computational efficiency. S4.4.3. Calculate the IAU analysis increment and apply... constraint: ;in, These are the constraint coefficients. Sensitivity experiments show that... and The time-average calculation diverges. It can operate stably for a period of time, with typical values ​​taken as follows: .

[0070] S4.5 incremental sustained-release application.

[0071] The IAU protocol administers the incremental analysis in a linear, gradual-release manner across multiple time steps, avoiding the impact of a single injection on the FVCOM physicochemical processes. The incremental component applied at each time step... ;in, The time window for slow release is usually equal to the assimilation interval; This is the time step of the pattern.

[0072] S5. Assimilation Increment Return and Pattern State Update.

[0073] S5.1. The AI ​​inference process serializes the calculated analysis increment vector and returns it to the main process through the ZeroMQ response channel.

[0074] S5.2. After receiving the analysis increment vector, the main process broadcasts the analysis increment vector to all child processes through the broadcast function MPI_Bcast.

[0075] S5.3. Each subprocess applies the analysis increment vector to the mesh cells it is responsible for: ;in, This represents the current forecast status of the FVCOM mode. The updated state variables are used as initial conditions for the next time step to continue the integration process.

[0076] S5.4. The updated state variables are used as the initial conditions for the next time step, and the physical-biochemical integral calculation in FVCOM mode continues. This completes one full online assimilation loop.

[0077] The online assimilation method in steps S2-S5 not only improved the simulation accuracy of chlorophyll a itself, but also produced a synergistic optimization effect on unobserved variables through the inherent biochemical coupling mechanism of the HBGC model. After assimilation, phytoplankton photosynthetic consumption of DIN increased, and the increase in total phytoplankton population enhanced sedimentation flux and biological feeding pressure, altering the nutrient redistribution pattern. Diagnostic analysis of material fluxes before and after assimilation was performed, including biological growth and physical consumption. Taking the PPT flux during a red tide event in a certain area as an example, the results are shown in Table 1.

[0078] Table 1 Comparison of PPT flux before and after assimilation

[0079]

[0080] After assimilation, the contribution of phytoplankton to biological growth increased significantly, with an average increase of about 16%, while the proportion of physical consumption decreased. This indicates that the increase in chlorophyll a introduced by assimilation reshaped the material distribution pattern of the ecosystem through the biochemical coupling chain.

[0081] The innovations of this invention are: (1) Heterogeneous process coupling architecture based on the message processing queue library ZeroMQ. Decoupling communication between the Fortran numerical model FVCOM and the Python artificial intelligence inference process is achieved through the request / response communication mode of ZeroMQ, solving the real-time interaction problem of online assimilation in heterogeneous language environments. The numerical model process and the AI ​​inference process run independently and do not interfere with each other. The numerical model does not need to be recompiled when the AI ​​model is updated or the algorithm is replaced. (2) Zero-intrusive data access method: By directly traversing the NetCDF average data linked list structure inside the FVCOM model, the target variable pointer is locked according to the variable name, realizing the access of state variables without modifying the model data read and write script. Only one independent module and about 20-30 lines of main program interface code need to be added. (3) Model-independent inference engine interface: The deep learning model is only used as the inference engine for assimilation incremental calculation. This invention does not limit the specific model architecture. Users only need to configure the list of variables that need to be communicated on the Fortran side and align the byte length of the data type on the Python side. Theoretically, any deep learning or traditional machine learning model can be accessed. (4) Utilization of the synergistic improvement effect of chlorophyll assimilation on nutrients: Utilizing the inherent biochemical coupling mechanism of the HBGC model, namely phytoplankton photosynthesis → DIN consumption → zooplankton feeding → nutrient regeneration, the non-observed variable DIN is indirectly constrained and improved by assimilating PPT.

[0082] The technical advantages of the method of this invention are as follows: (1) It realizes the closed-loop online fusion of numerical models and deep learning models. Compared with the existing offline post-processing correction mode of HBGC-CNN, this invention feeds back the inference results of deep learning models to the state variable update of FVCOM model in the form of analysis increments through the data assimilation framework, forming a closed-loop iteration of FVCOM prediction → CNN inference → assimilation calculation → incremental feedback → model re-prediction. This makes the deep learning model no longer an external corrector of numerical models, but an online assimilator embedded in the model running loop, which can continuously improve the initial field and subsequent prediction performance of the model. (2) Low invasiveness and low development and maintenance costs. Compared with the complex implementation of traditional EnKF assimilation which requires generating a large number of set members and modifying the model source code, this invention directly accesses the state variables through the NetCDF average data linked list pointer, which only requires adding one independent module and about 20-30 lines of main program interface code, significantly reducing the development cycle. (3) Low latency communication to meet the real-time assimilation requirements. Compared to the serious cache pollution problem caused by direct C++ calls within the process, this invention uses ZeroMQ for inter-process communication. The total time for a single ZMQ communication is about 0.002 seconds, and the overall FVCOM calculation time only increases by about 10%, avoiding cache pollution and ensuring that the core numerical calculation always runs at the highest efficiency. (4) Supports flexible switching of multiple assimilation algorithms. Compared to existing online assimilation systems that mostly use a single fixed algorithm, this invention supports flexible configuration of multiple assimilation algorithms without recompiling the FVCOM mode; and the assimilation method can be designed and extended by oneself. Data verification using the IAU scheme as an example shows that the root mean square error of chlorophyll a decreased from 2.64 to 1.92, and the Willmott skill score increased from 0.37 to 0.54. (5) Decoupling of heterogeneous languages ​​facilitates technology iteration. Compared to the in-process call scheme that requires mixed compilation of Fortran, C++ and Libtorch, this invention isolates the Fortran and Python processes through the ZeroMQ middleware, and both can upgrade their technology stacks independently. (6) Assimilation of chlorophyll a synergistically improves the simulation accuracy of nutrient DIN. Utilizing the inherent ecodynamic coupling mechanism of the HBGC model, this invention assimilates only the variable PPT, which indirectly improves the simulation of the non-observed variable DIN through the biochemical chain of phytoplankton photosynthesis → DIN consumption. Flux diagnostic analysis further reveals that the phytoplankton biological growth term increases after assimilation, providing high-quality teacher data for adjusting the parameterization scheme of nutrient simulation in numerical simulation.

Claims

1. A deep learning-based chlorophyll a concentration fusion calculation system, characterized in that, It includes a numerical model process cluster module, an artificial intelligence inference module, and a message queue communication module; The numerical model process cluster module includes a main process and multiple sub-processes. The main process is responsible for managing the control of the assimilation process, which includes assimilation time determination, state variable reading, data transmission, incremental reception, and broadcasting. The sub-processes are responsible for the hydrodynamic-water quality integral calculation of the finite volume ocean model (FVCOM). The artificial intelligence inference module includes a data preprocessing module, a deep learning inference module, a data assimilation calculation module, and a data postprocessing module. The data preprocessing module receives state variables sent by the FVCOM main process and performs standardization processing. The deep learning inference module performs chlorophyll a concentration inference prediction. The data assimilation calculation module calculates and analyzes the incremental vector according to the configured assimilation algorithm. The data postprocessing module serializes the analyzed incremental vector and returns it to the FVCOM main process. The message queue communication module is based on the ZeroMQ message middleware library for message processing, enabling bidirectional data exchange between the numerical model process cluster module and the artificial intelligence inference module. The main process traverses the NetCDF average data linked list structure within the FVCOM pattern, locks the memory pointer of the target variable based on the variable name, and achieves zero-intrusion state variable reading. After the read state variable is serialized, it is sent to the AI ​​inference process through the ZeroMQ request / response channel. After the AI ​​inference process completes the inference and assimilation calculations, it serializes the analysis increment vector and returns it to the main process through the same channel. The main process broadcasts the increment vector to all child processes through the broadcast function MPI_Bcast, and the child processes apply the increment to their respective grid cells to update the model state variables. The calculation of the incremental analysis includes: calculating the deviation between the background field and the observed field. ;in, For the observation field; For observation operators; Background field for the mode; calculate the Kalman gain matrix. ;in, The background error covariance matrix; The observation error covariance matrix; matrix From the standard deviation of background error Spatial correlation function structure: Where B(i,j) are elements in matrix B; For grid points With observation point The distance between them; For spatially relevant scales; The expression is: ;in, ; Calculate the IAU analysis increment ;in, For constraint term coefficients, This is the Kalman gain matrix.

2. A deep learning-based chlorophyll a concentration fusion calculation method, using the deep learning-based chlorophyll a concentration fusion calculation system described in claim 1, characterized in that... Includes the following steps: S1. FVCOM mode initialization and assimilation parameter configuration; S2. In FVCOM mode, the main time loop is entered. Hydrodynamic calculations and water quality calculations are performed at each time step. At the end of each time step, the main process checks whether the current simulation time has reached the predetermined assimilation time. S3. When the assimilation condition is met, the main process calls the data post-processing module, which traverses the NetCDF average data linked list structure inside the FVCOM mode, locks the memory pointer of the target variable according to the variable name, serializes the read state variable data into a binary stream, and then sends the serialized data to the AI ​​inference process through the ZeroMQ request channel, and blocks after sending to wait for the confirmation response. The S4.AI inference process receives serialized data sent by the main process, deserializes it, and restores it to the NumPy array format in Python. It then standardizes the various state variables and inputs the standardized state variables into a pre-trained deep learning model. The deep learning model extracts the spatial pattern of multi-channel input features through convolutional layers, maps it through fully connected layers, and outputs the estimated value of the target variable. It uses incremental analysis to update the IAU as an assimilation scheme, calculates the optimal magnitude of the analysis increment, and uses the IAU to gradually inject the analysis increment into the pattern in a linear and slow manner in multiple time steps. S5. After the AI ​​inference process serializes the calculated analysis increment vector, it returns it to the main process through the ZeroMQ response channel. After receiving the analysis increment vector, the main process broadcasts the analysis increment vector to all child processes through the broadcast function MPI_Bcast. Each child process applies the analysis increment vector to the grid cell it is responsible for, and the updated state variables are used as the initial conditions for the next time step to continue the physical-biochemical integral calculation in FVCOM mode. This completes a full online assimilation loop; In step S4, the calculation of the analysis increment includes: calculating the deviation between the background field and the observed field. ;in, For the observation field; For observation operators; Background field for the mode; calculate the Kalman gain matrix. ;in, The background error covariance matrix; The observation error covariance matrix; matrix From the standard deviation of background error Spatial correlation function structure: Where B(i,j) are elements in matrix B; For grid points With observation point The distance between them; For spatially relevant scales; The expression is: ;in, ; Calculate the IAU analysis increment ;in, These are the coefficients of the constraint terms; This is the Kalman gain matrix.

3. The deep learning-based chlorophyll a concentration fusion calculation method according to claim 2, characterized in that, Step S1 includes the following sub-steps: S1.1 Read the mesh file, initial field file, and boundary condition file of the FVCOM-HBGC mode to complete mode initialization; S1.2 Read the assimilation configuration file and parse the parameters assimilation switch DA_ON, assimilation time interval string DA_INTERVAL_STRING, and AI inference service address DA_SERVER_ADDR; S1.3 Read the background error standard deviation from the static file. and observation error standard deviation .

4. The deep learning-based chlorophyll a concentration fusion calculation method according to claim 2, characterized in that, In step S2, the method for determining whether the current simulation time has reached the predetermined assimilation time is as follows: parse DA_INTERVAL_STRING to obtain the assimilation interval in seconds. If the previous simulation time - the last assimilation time ≥ the interval, then the assimilation process is triggered.

5. The deep learning-based chlorophyll a concentration fusion calculation method according to claim 2, characterized in that, The variables involved in step S3 include: temperature T, salinity S, total phytoplankton PPT, dissolved inorganic nitrogen DIN, total zooplankton ZPT, and dissolved organic nitrogen DON.

6. The deep learning-based chlorophyll a concentration fusion calculation method according to claim 2, characterized in that, In step S4, the standardized formula is: ; where X i This represents the original data; μ is the mean of the state variable on the training set; σ is the standard deviation of the state variable on the training set; X i * This represents the standardized data.

7. The deep learning-based chlorophyll a concentration fusion calculation method according to claim 2, characterized in that, In step S4, the incremental component applied at each time step ;in, For a time window of slow release; This is the time step of the pattern.

8. The deep learning-based chlorophyll a concentration fusion calculation method according to claim 7, characterized in that, In step S5, each subprocess applying the analysis increment vector to its respective grid cell means calculating the forecast state of the FVCOM model at the current moment. ; Among them, among them, This represents the current forecast status of the FVCOM mode. The updated state variables are used as initial conditions for the next time step to continue the integration process.

Citation Information

Patent Citations

  • Sea surface temperature forecasting method and system based on deep learning and storage medium

    CN120974064A

  • Ai-enhanced simulation and modeling experimentation and control

    US20240348663A1