Variable grid atmospheric chemical data assimilation method and system based on deep learning

By using a deep learning-based approach, the current-moment feature information of the variable grid is obtained and an assimilation model is constructed, which solves the problems of assimilation accuracy and efficiency under dynamic grid structure, and realizes efficient and high-precision pollutant concentration assimilation, adapting to the assimilation needs of high-density observation stations.

CN120950962APending Publication Date: 2025-11-14INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510934802.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies are difficult to adapt to dynamically changing variable grid structures, resulting in low accuracy and poor computational efficiency in atmospheric pollutant concentration assimilation, especially under conditions of high density and massive observation stations, making it difficult to achieve efficient assimilation.

Method used

By employing a deep learning-based approach, a concentration assimilation model is constructed by acquiring the current-time feature information of the initial grid in a variable grid. The model is then iteratively trained using training samples from the previous time step, and the grid is selected and optimized. Clusters are then constructed and a comprehensive score is calculated to achieve efficient and high-precision pollutant concentration assimilation.

Benefits of technology

It improves the accuracy and efficiency of atmospheric pollutant concentration assimilation, can adapt to the assimilation needs of high-density and massive observation stations, reduces computational errors and overhead, and enhances the model's time-series generalization ability and adaptability to short-term changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950962A_ABST
    Figure CN120950962A_ABST
Patent Text Reader

Abstract

The invention provides a variable grid atmospheric chemical data assimilation method and system based on deep learning, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining the grid feature information of all initial grids in a variable grid at the current moment; inputting all the grid feature information at the current moment into an assimilation concentration model at the current moment to obtain an actual assimilation concentration at the current moment; wherein the assimilation concentration model at the current moment is obtained by training the assimilation concentration model at the previous moment according to a plurality of training samples; the training sample comprises grid feature information of a target moment for optimizing the grid; the target moment comprises a current moment and a historical moment; the optimized grid is obtained by screening a plurality of actual grids, and the actual grids are initial grids including observation stations at the current moment. According to the method, fusion training is carried out on observation and simulation information under the variable grid through the artificial intelligence model, so that efficient and high-precision assimilation reasoning of the atmospheric pollutant concentration is realized under the variable grid model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for assimilating variable grid atmospheric chemical data based on deep learning. Background Technology

[0002] In recent years, air pollution has received increasing attention. Atmospheric chemical data assimilation technology, as an important means to improve the accuracy of pollutant concentration simulation, has been widely used in the field of air quality forecasting and analysis.

[0003] With the development of numerical simulation technology, variable mesh structures have emerged, where the model mesh can be dynamically adjusted over time to adapt to changes in local pollutant concentration gradients and computational needs. This dynamically adjustable variable mesh structure presents a significant contradiction with the fixed mesh structure required by traditional data assimilation methods. Existing methods cannot effectively handle dynamically changing mesh characteristics, leading to frequent spatial interpolation operations and introducing uncontrollable interpolation errors. Furthermore, traditional methods require high-dimensional matrix inversion operations, and the matrix dimension is directly related to the number of observations. When the mesh changes dynamically, the matrix dimension changes frequently, significantly reducing the stability and computational efficiency of matrix inversion. This results in decreased accuracy of the assimilation results and increased computational overhead, making it unsuitable for assimilation of high-density, massive observation stations. Summary of the Invention

[0004] This invention provides a method, system, electronic device, storage medium, and computer program product for assimilating atmospheric chemical data based on a variable grid using deep learning. This addresses the shortcomings of existing technologies, which struggle to adapt to dynamically changing grid structures, resulting in low assimilation accuracy and poor computational efficiency. The invention achieves efficient and high-precision assimilation inference of atmospheric pollutant concentrations under a variable grid model.

[0005] This invention provides a method for assimilating variable grid atmospheric chemical data based on deep learning, comprising the following steps: Obtain the grid feature information of all initial grids in the variable grid at the current time; the number of observation stations in each initial grid is continuously updated as time changes; Input all the grid feature information at the current time into the assimilation concentration model at the current time to obtain the actual assimilation concentration at the current time output by the assimilation concentration model at the current time; The assimilation concentration model at the current moment is obtained by training the assimilation concentration model at the previous moment based on multiple training samples; the training samples include grid feature information of the target moment of at least one optimized grid; the target moment includes the current moment and historical moments; the optimized grid is selected from multiple actual grids, and the actual grid is the initial grid that includes at least one of the observation stations at the current moment.

[0006] The method for assimilating variable grid atmospheric chemical data based on deep learning provided by the present invention further includes: For any of the actual grids, an observation quality feature vector is constructed for each of the observation stations therein, the observation quality feature vector including at least two information quality indicators for characterizing the integrity and reliability of the observation data; Obtain the grid density corresponding to each of the actual grids; Based on all the information quality indicators and the grid density corresponding to each actual grid, all the actual grids are divided into several clusters; Calculate the overall score for each actual grid in each cluster; The actual grid with the highest comprehensive score is selected from each cluster and used as the optimized grid.

[0007] According to a deep learning-based variable grid atmospheric chemical data assimilation method provided by the present invention, the method involves dividing all the actual grids into several clusters based on all the information quality indicators and the grid density corresponding to each actual grid, including: Calculate the comprehensive fuzzy quality score for each actual grid based on all the information quality indicators corresponding to each actual grid. The comprehensive fuzzy quality score and the grid density of each actual grid are concatenated to construct the clustering input feature vector of each actual grid; A clustering algorithm based on the clustering input feature vector is performed on all the actual grids to divide all the actual grids into several clusters.

[0008] According to a deep learning-based variable grid atmospheric chemical data assimilation method provided by the present invention, the step of calculating the comprehensive fuzzy quality score of each actual grid based on all the information quality indices corresponding to each actual grid includes: For any observation station in each of the actual grids, each information quality index in the corresponding observation quality feature vector is input into the corresponding fuzzy membership function to obtain the fuzzy membership degree corresponding to each information quality index. The initial fuzzy quality score of the observation station is obtained by weighting all the fuzzy membership degrees corresponding to the observation station according to the preset fuzzy weight vector. The average of the initial fuzzy quality scores of all observation stations in each actual grid is used to obtain the comprehensive fuzzy quality score of each actual grid.

[0009] According to a deep learning-based variable grid atmospheric chemical data assimilation method provided by the present invention, the calculation of the comprehensive score of each actual grid in each cluster includes: For any of the clusters, the comprehensive fuzzy quality score of each of the actual grids is normalized to obtain a first normalized score for each of the actual grids, and the grid density of each of the actual grids is normalized to obtain a second normalized score for each of the actual grids. The first normalized score and the second normalized score of each actual grid are weighted and summed to obtain the comprehensive score of each actual grid.

[0010] The method for assimilating variable grid atmospheric chemical data based on deep learning provided by the present invention further includes: For any of the optimized grids, obtain its grid feature information at each of the target times, and obtain its actual observation concentration at each of the target times; the actual observation concentration is calculated based on the observation data of all observation stations in the optimized grid at the corresponding target times; The grid feature information of each target time is input into the assimilation concentration model of the previous time to obtain the assimilation concentration sample of each target time output by the assimilation concentration model of the previous time; wherein, the assimilation concentration model of the starting time is an artificial intelligence model; Based on the error between the assimilated concentration sample and the corresponding actual observed concentration at each target time, calculate the first loss function value at each target time. Based on the first loss function values ​​of all target times in each optimization grid, the second loss function values ​​of each optimization grid are obtained; Calculate the comprehensive loss function value based on the second loss function values ​​of all the optimized grids; The assimilation concentration model at the previous time step is trained with the minimum comprehensive loss function value as the training objective to obtain the assimilation concentration model at the current time step.

[0011] According to a deep learning-based variable grid atmospheric chemical data assimilation method provided by the present invention, the step of obtaining a second loss function value for each of the optimized grids based on the first loss function values ​​of all the target times in each optimized grid includes: Assign influence weights to the first loss function value at each target time, wherein the influence weight of historical times that are closer to the current time is greater; The second loss function value is obtained by weighting and summing each of the first loss function values ​​with its corresponding influence weight.

[0012] According to the present invention, a method for assimilating atmospheric chemical data based on a variable grid using deep learning includes obtaining the grid feature information of all initial grids in the variable grid at the current time, comprising: Extract the simulation information of all initial grids at the current time from the model simulation results of the variable grid, and obtain the environmental information of the unit region where all the initial grids are located at the current time; The simulation information includes one or more of the following: grid area, wind speed, temperature, humidity, air pressure, and emissions. The environmental information includes one or more of the following: topography, population density, nighttime light intensity, and road network density.

[0013] This invention also provides a deep learning-based variable grid atmospheric chemical data assimilation system, comprising the following modules: The acquisition module is used to acquire the grid feature information of all initial grids in the variable grid at the current time; the number of observation stations in each initial grid is continuously updated as time changes; The first processing module is used to input all the grid feature information of the current time into the assimilation concentration model of the current time to obtain the actual assimilation concentration of the current time output by the assimilation concentration model of the current time; wherein, the assimilation concentration model of the current time is trained on the assimilation concentration model of the previous time based on multiple training samples; the training samples include grid feature information of at least one optimized grid at the target time; the target time includes the current time and historical time; the optimized grid is selected from multiple actual grids, and the actual grid is the initial grid that includes at least one of the observation stations at the current time.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the deep learning-based variable grid atmospheric chemical data assimilation method as described above.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the deep learning-based variable grid atmospheric chemical data assimilation method as described above.

[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the deep learning-based variable grid atmospheric chemical data assimilation method as described above.

[0017] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: By acquiring the grid feature information of all initial grids in the variable grid at the current time and updating the number of observation stations within each initial grid in real time, the spatiotemporal dynamics of the atmospheric field can be accurately reflected. Since the number and distribution of grid points differ at different times in the variable grid structure, the system needs to reacquire the corresponding grid feature information at each time step to ensure the consistency of the input data with the current grid state, thereby improving the input effectiveness of the concentration model. The assimilation concentration model at the current time step is iteratively trained from the model at the previous time step, possessing continuous rolling optimization capabilities. It can gradually learn new observation information over time, achieving dynamic adjustment of model parameters. By using multiple training samples to train the assimilation concentration model at the previous time step, the model's temporal generalization ability and adaptability to short-term changes are improved. The training samples contain grid feature information of at least one optimized grid at multiple target times, covering both the current time and historical times, enabling the construction of a stable time-series training structure. Each optimized grid is selected from the initial grids containing observation stations at the current time step, possessing strong representativeness and high data quality. By summarizing and learning the feature information of these optimized grids, the model can maintain its ability to learn the overall pollution field distribution characteristics while effectively reducing the amount of high-density, massive observation data. By inputting the grid feature information of all optimized grids at the current moment into the assimilation concentration model at the current moment, the model can achieve efficient and high-precision inference of pollutant concentrations based on actual grid feature information, thereby achieving assimilation of high-density, massive observation stations. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the variable grid atmospheric chemical data assimilation method based on deep learning provided by the present invention.

[0020] Figure 2 This is a schematic diagram of the training method for the assimilation concentration model at the current moment provided by the present invention.

[0021] Figure 3 This is a schematic diagram of the artificial intelligence model structure provided in this application.

[0022] Figure 4 This is a schematic diagram of the structure of the variable grid atmospheric chemical data assimilation system based on deep learning provided by the present invention.

[0023] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0025] It should be noted that in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships according to the accompanying drawings, are only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the system or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0026] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] The following is combined with Figures 1 to 5 This invention describes the deep learning-based variable grid atmospheric chemical data assimilation method, system, electronic device, storage medium, and computer program product provided by this invention.

[0028] Reference Figure 1 , Figure 1 This is one of the flowcharts illustrating the deep learning-based variable grid atmospheric chemical data assimilation method provided by this invention, such as... Figure 1 As shown, steps 101 to 102 are included: Step 101: Obtain the grid feature information of all initial grids in the variable grid at the current time; the number of observation stations in each initial grid is constantly updated as time changes.

[0029] In variable-grid atmospheric chemistry simulations, the distribution of the initial grid is dynamically adjusted based on factors such as geographical characteristics, pollutant distribution, and meteorological evolution. Therefore, the number and structure of grids vary at different time points. In particular, the number of observation stations covered within each grid changes over time, exhibiting a non-uniform and dynamic distribution. Therefore, dynamically sensing the physical and environmental characteristics of each initial grid at the current moment is fundamental to achieving high-precision data assimilation. Traditional assimilation methods typically set observation mappings based on a fixed grid structure. This invention, however, obtains the current grid characteristic information of all initial grids in the variable grid through step 101, thereby providing sufficient and representative input features for the subsequent assimilation concentration model, enabling the model to accurately express the spatial variation patterns under the current atmospheric conditions.

[0030] In a preferred embodiment, step 101 is specifically implemented through the following steps: The simulation information of all initial grids at the current moment is extracted from the model simulation results of the variable grid, and the environmental information of the unit area where all initial grids are located at the current moment is obtained. The simulation information includes one or more of the following: grid area, wind speed, temperature, humidity, air pressure and emissions. The environmental information includes one or more of the following: terrain, population density, nighttime light intensity and road network density.

[0031] The initial grid refers to the set of grid cells determined by the variable grid model at the current moment, with each initial grid corresponding to a spatial cell in the simulation region. Due to the dynamic adjustment capability of the variable grid, this set may have different quantities and distributions at different times. In this step, the grid characteristic information includes two parts: simulation information and environmental information. The simulation information reflects the numerical simulation results of the atmospheric model for each initial grid at the current moment, specifically including one or more of the following: grid area, wind speed, temperature, humidity, air pressure, and emissions. These elements are usually directly derived from the output of atmospheric chemistry models. The environmental information consists of static or semi-static geographic features related to the location of the grid, including one or more of the following: topography, population density, nighttime light intensity, and road network density. This information often comes from external high-resolution geographic information datasets and can provide prior information such as urban morphology and human activities related to pollutant distribution.

[0032] In practice, the system first calls upon the current output of the variable grid simulation system to traverse and extract data from the initial grid across the entire region, recording the simulated meteorological and pollution values ​​for each grid, including grid area, wind speed, temperature, humidity, air pressure, and emissions. Next, based on the geographic location index of the initial grid, it retrieves topographic elevation data, nighttime light remote sensing inversion values, population density, or road network density from a pre-set high-resolution environmental factor database. For each initial grid, this process constructs its complete current-moment grid feature information vector, which serves as crucial foundational data for subsequent AI model assimilation.

[0033] Step 102: Input all the grid feature information of the current time into the assimilation concentration model of the current time to obtain the actual assimilation concentration of the current time output by the assimilation concentration model of the current time; wherein, the assimilation concentration model of the current time is trained on the assimilation concentration model of the previous time based on multiple training samples; the training samples include the grid feature information of the target time of at least one optimized grid; the target time includes the current time and historical time; the optimized grid is selected from multiple actual grids, and the actual grid is the initial grid that includes at least one observation station at the current time.

[0034] Step 102 involves inputting all current-time grid feature information into the current-time assimilation concentration model and outputting the actual assimilation concentration at the current time. This step primarily addresses the problems of dynamic grid inconsistency, interpolation error accumulation, and heavy matrix computation burden inherent in traditional data assimilation methods under variable grid conditions. It replaces traditional assimilation methods with an artificial intelligence model, achieving efficient and nonlinear correction of the variable grid simulation results. Because the AI ​​model possesses strong characterization capabilities for complex spatiotemporal features and multi-source data, it can achieve assimilation optimization of the atmospheric pollution concentration field without relying on an explicit error covariance matrix.

[0035] The current-moment assimilation concentration model refers to an inference model built using artificial intelligence methods that outputs the pollutant assimilation concentration at the current moment. Unlike traditional assimilation methods, this model is not built directly from scratch, but rather iteratively updated based on the previous assimilation concentration model through training samples, exhibiting good temporal continuity and rolling evolution capabilities. The model structure is typically based on a fully connected neural network, and can incorporate mechanisms such as residual connections, regularization, and Dropout to improve training stability and inference accuracy.

[0036] During training, the training samples used by the model are derived from the target time feature information of multiple optimized grids. Here, the optimized grid refers to the representative grid selected from multiple actual grids at the current time, where the actual grid refers to the initial grid that includes at least one observation station at the current time. The selection process of the optimized grid comprehensively considers the completeness, reliability, and spatial distribution density of the observation data to ensure its representativeness and data quality. For the specific method of selecting the optimized grid from multiple actual grids, please refer to steps 201 to 205 in the following implementation.

[0037] The target time in the training samples includes the current time t and at least one historical time t-i, to construct a data sequence suitable for short-term time series modeling, where i∈[1,n] and n is a positive integer. The time span between each adjacent time can be selected according to the actual task scenario. For example, for a conventional assimilation concentration inference task, the time span can be 1 hour, that is, the time span between the current time t and the previous time t-1 is 1 hour; if it is a task made for an assimilation dataset, the time span can be 7 days, that is, the time span between the current time t and the previous time t-1 is 7 days. The training samples for each target time consist of grid feature information corresponding to the optimized grid, including simulation information (such as temperature, humidity, wind speed, emissions, etc.) and environmental information (such as terrain, population density, nighttime lights, road network density, etc.), and are equipped with the observed concentration of the grid at the target time as the training target value.

[0038] During the training of the assimilation concentration model at the current time, the system first loads the assimilation concentration model from the previous time step, then inputs the feature information of each optimized grid at each target time step into the model to obtain the concentration value predicted by the model (i.e., the assimilation concentration sample). The error is then calculated by comparing this predicted concentration with the actual observed concentration, forming the target loss. The loss values ​​of multiple optimized grids are weighted by time and space to obtain the total loss value, which is used to guide the model parameter update, thereby generating the assimilation concentration model at the current time step. After the update is completed, the model is used to process the current time step feature input of the entire grid, outputting the actual assimilation concentration of each initial grid. The specific method for selecting optimized grids from multiple actual grids can be found in steps 601 to 606 of the subsequent implementation methods.

[0039] This step significantly improves the accuracy and resolution of atmospheric pollution assimilation results. Compared to traditional assimilation algorithms that rely on spatial interpolation and matrix inversion, this embodiment directly inputs grid features into the AI ​​model and outputs assimilated concentrations, eliminating intermediate interpolation and matrix calculation processes. This reduces numerical errors and significantly improves computational efficiency. Furthermore, the model possesses continuous adaptive updating capabilities. While maintaining a constant model structure, it continuously adjusts parameters based on input samples to better adapt to variable grid structures and changes in observational data, thus ensuring the dynamic consistency and physical rationality of the concentration field results. Particularly under large-scale station observation conditions, this approach effectively avoids the numerical instability and storage overhead caused by the high matrix dimension of traditional methods, providing reliable support for high-frequency, large-scale atmospheric chemical data assimilation.

[0040] In one possible implementation, the method for selecting an optimized mesh from multiple actual meshes specifically includes steps 201 to 205: Step 201: For any actual grid, construct the observation quality feature vector for each observation station. The observation quality feature vector includes at least two information quality indicators used to characterize the integrity and reliability of the observation data.

[0041] In this step, the actual grid refers to the initial grid containing at least one observation station at the current moment; that is, the set of grid cells that have a data support foundation before the current model simulation results are fused with the observation data. Each actual grid may contain one or more observation stations, and the quality of observation data from different stations varies. Therefore, to achieve grid-level data quality assessment, this step first extracts quality features at the station level, and then performs subsequent aggregation processing at the grid level.

[0042] An observation quality feature vector is a vector-based data structure used to characterize the stability, completeness, and reliability of observation data from a single observation station within a target timeframe. It contains multiple information quality indicators, each describing a dimension of observation data quality. For example, selectable quality indicators include: the missing observation rate, the proportion of outliers, historical stability indicators, instrument calibration frequency, and data release delay time. These indicators comprehensively reflect whether an observation point possesses the ability to continuously, stably, and reliably output data.

[0043] In practice, the system iterates through all observation stations within each actual grid, performing a quality assessment on the observation data of each station within a preset target time window (e.g., the first 24 hours). For each station, the system calculates its missing observation rate, outlier detection results, and deviation from historical averages according to various indicators, and combines these indicators into an observation quality feature vector corresponding to that station. This vector can be normalized or standardized to unify its numerical range for subsequent cluster analysis and fuzzy scoring.

[0044] Step 201 achieves two objectives: firstly, by meticulously constructing quality feature vectors at the site level, the credibility of observation data can be identified from the bottom up, providing a solid foundation for subsequent grid-level quality assessment; secondly, the multi-dimensional information quality indicator system avoids a single criterion for judging the "good" or "bad" of observation data, making the selection of optimized grids more objective.

[0045] Step 202: Obtain the grid density corresponding to each actual grid.

[0046] In this step, grid density is a metric used to measure the spatial coverage of observation data within the actual grid. It is typically expressed as the number of observation stations per unit area, but can also be adjusted by considering the distribution of station numbers in neighboring grids to more accurately reflect the local spatial observation density. Grid density, as a key indicator of spatial sample uniformity, will be directly used in subsequent clustering and scoring processes.

[0047] In practice, the system first traverses all actual grids at the current moment and counts the number of observation stations contained in each actual grid. Combining this with the actual area information of the grid in the simulation model, the number of stations per unit area is calculated, and this value is used as the basic density index for that grid. To further enhance the sensitivity to changes in spatial neighborhood, the system can choose to introduce a weighted average of the station densities in neighboring grids to form a smoothed grid density. If some stations are located at the intersection of multiple grids, an interpolation-based allocation method can be used to assign them to multiple grids.

[0048] Step 203: Based on all information quality indicators and grid density corresponding to each actual grid, divide all actual grids into several clusters.

[0049] Step 203 divides all actual grids into several clusters based on all information quality indicators and grid density corresponding to each actual grid. This step aims to incorporate a comprehensive consideration of two key dimensions—information quality and spatial distribution—in the grid selection process, ensuring both the quality of the observation data and the spatial diversity and representativeness of the optimized grids. Since the observation quality and spatial density characteristics of actual grids are often highly heterogeneous, directly ranking and selecting the best grids could easily lead to an over-concentration of the selected grids in a particular region, resulting in insufficient training sample coverage and ultimately affecting the assimilation model's generalization ability across the entire grid field. Therefore, this step uses clustering to divide the actual grids into several structurally distinct clusters, providing a structured basis for subsequent selection of the best points within each cluster.

[0050] In the specific implementation process, the system first constructs a corresponding clustering input feature vector for each actual grid. This process includes extracting the observation quality feature vectors of all observation stations in the grid, synthesizing the comprehensive fuzzy quality score of the grid using the mean method, and then concatenating the comprehensive fuzzy quality score with the corresponding grid density to form a complete input vector. Afterwards, based on the clustering input feature vectors of all actual grids, the system executes a preset clustering algorithm for clustering. Specific clustering methods can include K-means clustering, fuzzy C-means (FCM), Gaussian mixture model (GMM), or density-based clustering methods (such as DBSCAN). The number of clusters can be set based on empirical values, or the optimal number of clusters can be automatically determined using evaluation indicators such as the silhouette coefficient and Davies-Bouldin index.

[0051] Step 204: Calculate the overall score of each actual grid in each cluster; Step 204 calculates the comprehensive score of each actual grid within each cluster. This step aims to further refine the evaluation and ranking of the actual grids within each cluster, based on the completed clustering, thus providing an objective and quantifiable basis for subsequent grid selection. Since different actual grids exhibit significant differences in observation data quality and spatial density, even within the same cluster, there may be variations in quality. Therefore, it is necessary to introduce comprehensive indicators within the cluster to score grid quality, ensuring that the final selected optimized grids are representative and highly reliable within their respective clusters.

[0052] In the implementation process, the system first extracts the corresponding comprehensive fuzzy quality score and grid density index for all actual grids within any given cluster. To eliminate the influence of different dimensions between the indices, the system normalizes these two indices separately, forming a first normalized score and a second normalized score. The normalization method can employ min-max normalization or Z-score normalization. Subsequently, the system performs a weighted sum of the two normalized scores according to preset weighting coefficients, ultimately obtaining the comprehensive score for the actual grid. This score reflects the grid's comprehensive performance in terms of both observation reliability and spatial balance; a higher score indicates a higher priority for optimization.

[0053] Step 205: Select the actual grid with the highest comprehensive score from each cluster as the optimized grid.

[0054] Step 205 involves selecting the actual grid cell with the highest overall score from each cluster as the optimized grid cell. This step is implemented to select the most representative samples from each group of actual grid cells after clustering and scoring. The selected samples will be used to train the assimilation concentration model. This improves the spatial representativeness of the samples and reduces the impact of data noise.

[0055] In practice, the system iterates through each cluster sequentially. Within each cluster, the system obtains the comprehensive score of all actual grids. The system selects the actual grid with the highest score as the optimized grid for that cluster. If multiple actual grids have the same highest score, the system can select the grid with the lower missing detection rate. A secondary ranking rule can also be introduced for this determination. The system can also set a scoring threshold. If a grid's score is below the threshold, it will not be selected as an optimized grid, even if it has the highest score within the cluster.

[0056] Using the above method, the system selects an optimized grid from each cluster. The selected optimized grid is used for subsequent assimilation concentration model training. Each optimized grid has high representativeness and observational reliability within its cluster.

[0057] In one possible implementation, step 203 specifically includes steps 301 to 303: Step 301: Calculate the comprehensive fuzzy quality score of each actual grid based on all information quality indicators corresponding to each actual grid.

[0058] Step 301 calculates the comprehensive fuzzy quality score for each actual grid based on all information quality indicators corresponding to each actual grid. The purpose of this step is to convert multiple quality indicators at the observation station level into a single score value at the actual grid level, providing a unified input format for the subsequent construction of clustering input feature vectors. Since an actual grid typically contains multiple observation stations, and the data quality of each station varies, a fuzzy method is needed to uniformly measure the overall observation quality under multi-dimensional indicators.

[0059] In one alternative implementation, step 301 specifically calculates the overall fuzzy quality score for each actual grid in the following manner: Step 401: For any observation station in each actual grid, input each information quality index in its corresponding observation quality feature vector into the corresponding fuzzy membership function to obtain the fuzzy membership degree corresponding to each information quality index.

[0060] Step 402: Calculate the weighted average of all fuzzy membership degrees corresponding to the observation station according to the preset fuzzy weight vector to obtain the initial fuzzy quality score of the observation station.

[0061] Step 403: Average the initial fuzzy quality scores of all observation stations in each actual grid to obtain the comprehensive fuzzy quality score of each actual grid.

[0062] Specifically, step 401 involves inputting each information quality index from the corresponding observation quality feature vector of any observation station in each actual grid into the corresponding fuzzy membership function to obtain the fuzzy membership degree corresponding to each information quality index. The system first traverses all observation stations in the current actual grid and extracts the observation quality feature vector for each station. The system then inputs each information quality index in the vector into a preset fuzzy membership function. Each membership function maps the original index value to a fuzzy membership degree between 0 and 1 according to a predefined membership interval.

[0063] Step 402 involves calculating a weighted average of all fuzzy membership degrees corresponding to the observation station based on a preset fuzzy weight vector to obtain the initial fuzzy quality score for the observation station. The system sequentially reads the fuzzy membership degrees of all information quality indicators for the station. The system then calls the fuzzy weight vector to weight each membership degree and sums the weighted results to obtain the initial fuzzy quality score for the observation station. This score, a value between 0 and 1, is used to measure the overall quality level of the observation data at the station.

[0064] Step 403 is used to average the initial fuzzy quality scores of all observation stations in each actual grid to obtain the comprehensive fuzzy quality score of each actual grid. The system collects the initial fuzzy quality scores of all observation stations within the current actual grid. The system performs an arithmetic mean on these scores, and the result is the comprehensive fuzzy quality score of that actual grid. The system uses this score as an input indicator for the quality of the observation data in subsequent cluster analysis.

[0065] Step 302: Concatenate the comprehensive fuzzy quality score and grid density of each actual grid to construct the clustering input feature vector for each actual grid.

[0066] In this step, the comprehensive fuzzy quality score is the numerical value obtained in step 403 used to characterize the overall quality of the observation data in the actual grid, and the grid density is the numerical value calculated in step 202 used to reflect the spatial distribution density of the observation stations. The clustering input feature vector is a one-dimensional vector composed of these two values, which serves as the input unit of the actual grid in the clustering operation.

[0067] In practice, the system iterates through all actual grids at the current time. For each actual grid, the system retrieves its comprehensive fuzzy quality score and corresponding grid density value. These two values ​​are then concatenated in a fixed order to form a one-dimensional feature vector containing two elements. This clustering input feature vector is subsequently labeled as a unique cluster representation for the corresponding actual grid and stored in a unified data structure for use in subsequent steps.

[0068] Step 303: Perform a clustering algorithm based on the clustering input feature vector on all actual grids to divide all actual grids into several clusters.

[0069] In the specific implementation process, the system first collects the clustering input feature vectors of all actual grids at the current moment. The system then calls a preset clustering algorithm to process this feature set. The clustering algorithm can be K-means, fuzzy C-means, hierarchical clustering, or density-based methods. Algorithm parameters include the number of clusters, distance metric, and number of iterations. The system performs clustering operations according to the set parameters, dividing all actual grids into several clusters and assigning a cluster identifier to each actual grid.

[0070] In one possible implementation, step 204 specifically includes steps 501 to 502: Step 501: For any cluster, normalize the comprehensive fuzzy quality score of each actual grid to obtain the first normalized score of each actual grid, and normalize the grid density of each actual grid to obtain the second normalized score of each actual grid.

[0071] Step 502: Perform a weighted summation of the first and second normalized scores of each actual grid to obtain the comprehensive score of each actual grid.

[0072] Specifically, step 501 involves normalizing the overall fuzzy quality score of each actual grid cell within any given cluster to obtain a first normalized score for each actual grid cell, and then normalizing the grid density of each actual grid cell to obtain a second normalized score. The system extracts the overall fuzzy quality score of all actual grid cells in the current cluster, selecting the maximum and minimum values ​​as boundary values ​​and using linear normalization to map all scores to the interval between 0 and 1 to obtain the first normalized score. The system similarly processes the grid density to obtain the second normalized score. Each actual grid cell simultaneously possesses a pair of normalized scores, representing its relative position within the cluster in terms of both observed data quality and spatial distribution density.

[0073] Step 502 involves a weighted sum of the first and second normalized scores of each actual grid to obtain a comprehensive score for each actual grid. The system assigns weights to the two normalized scores, with the weight values ​​configured by the system (e.g., observation quality weight α, density weight 1-α), reflecting the degree of emphasis on observation quality and spatial equilibrium. The system multiplies the first normalized score of each actual grid by the observation quality weight and the second normalized score by the density weight, then adds the two products to obtain the comprehensive score for that actual grid. The comprehensive score is a single numerical value representing the overall preference of that actual grid within the current cluster. The system records the comprehensive score of each actual grid, providing a basis for subsequent grid selection optimization.

[0074] In one possible implementation, a training method for the assimilation concentration model at the current time is also included. For a better understanding of the training method for the assimilation concentration model at the current time in this scheme, please refer to... Figure 2 , Figure 2 This is a schematic diagram of the training method for the assimilation concentration model at the current moment provided by the present invention. The training method specifically includes steps 601 to 606: Step 601: For any optimized grid, obtain its grid feature information at each target time and obtain its actual observed concentration at each target time; the actual observed concentration is calculated based on the observation data of all observation stations in the optimized grid at the corresponding target time.

[0075] In the specific implementation process, the system traverses the optimized grid set to determine the state of each optimized grid at multiple preset target times. The system retrieves the simulation results of the variable grid and extracts simulation information such as grid area, wind speed, temperature, humidity, air pressure, and emissions for the corresponding optimized grid at each target time. The system also extracts environmental information such as terrain, population density, nighttime light intensity, and road network density corresponding to the grid location. The system merges the above information to form the grid feature information for that target time. Furthermore, the system processes the observation data of all observation stations in the same grid at that target time, and calculates the actual observed concentration at that time using an averaging or weighted method.

[0076] Step 602: Input the grid feature information of each target time into the assimilation concentration model of the previous time to obtain the assimilation concentration sample of each target time output by the assimilation concentration model of the previous time; wherein, the assimilation concentration model of the initial time is an artificial intelligence model.

[0077] Step 602 involves inputting the grid feature information of each target time step into the assimilation concentration model of the previous time step, obtaining the assimilation concentration sample for each target time step output by the assimilation concentration model of the previous time step. This step is implemented because it requires inference calculations based on the optimized grid feature information of multiple historical target time steps using the currently trained assimilation concentration model of the previous time step, thereby forming a predicted output. The system can compare this predicted output with the actual observed concentration obtained in step 601, which is used to subsequently construct the loss function and complete the model update.

[0078] Combination Figure 2 It can be seen that the system uses the initial artificial intelligence model at the initial time (t-n). The model structure of the artificial intelligence model can be referred to... Figure 3 , Figure 3 This is a schematic diagram of the artificial intelligence model structure provided in this application. The artificial intelligence model adopts a feedforward neural network structure based on a fully connected network. From left to right, the model structure includes an input layer, several feature transformation regions (blocks 1 to 3), and an output layer. The model structure has a clear hierarchical organization and modular design, possesses strong nonlinear expression capabilities and anti-overfitting ability, and is suitable for atmospheric pollutant concentration assimilation prediction tasks under complex variable conditions.

[0079] The input layer receives the mesh feature information of the optimized mesh at the target time. The input data is first normalized by a regularization module to improve the numerical stability of the model. The normalized result then enters the fully connected layer to complete the transformation of the input features into a high-dimensional feature space.

[0080] Blocks 1 through 3 are used for feature extraction and representation learning, respectively. Each block includes a regularization layer and a fully connected layer. Block 2 includes a dropout layer after the fully connected layer, with a dropout rate of 0.5. This dropout layer randomly discards half of the activated nodes during training to reduce the risk of overfitting. The feature dimension of each block's output is n=512, meaning that the dimension of each layer's output vector is 512.

[0081] The output layer receives the output features of the last block and generates the final assimilation concentration value through a set of fully connected transformations.

[0082] In this overall structure, regularization layers stabilize the model training process, fully connected layers implement feature mapping and combination, and temporary deprecation layers enhance the model's generalization ability. This structure, while ensuring model capacity, improves training stability and simulation accuracy through modular design and regularization mechanisms. The model can be trained using the standard backpropagation algorithm, adapting to iterative updates of input samples at multiple time steps.

[0083] This AI model, serving as the first version of the assimilation concentration model, takes grid feature information from time t-n as input. The system calculates the difference between the output and the actual observed concentration at the corresponding time, forming training data to obtain the assimilation concentration model at time t-n+1. The system then iteratively advances, gradually inputting grid feature information from multiple target times, from t-n+1 to t-1, into the assimilation concentration model corresponding to the previous time. Each inference process is based on the previous model version, inputting the current features and outputting the assimilation concentration sample for the corresponding target time. The system iteratively trains the model until the model training at time t-1 is complete.

[0084] After the system completes the assimilation concentration model training at time t-1, the grid feature information of all optimized grids at each target time (including t-n to t-1) is input into the assimilation concentration model at time t-1 to uniformly generate assimilation concentration samples for each target time. These samples constitute the prediction output part required for the current model training. The system retains the assimilation concentration sample of each optimized grid at each target time for error calculation with the corresponding observed concentration in subsequent step 603.

[0085] Step 603: Calculate the first loss function value for each target time based on the error between the assimilation concentration sample and the corresponding actual observed concentration at each target time.

[0086] It is necessary to quantify the difference between the model's predicted values ​​and the actual observed values, so as to provide a clear direction for subsequent model training. Therefore, step 603 calculates the error between the model output and the observed data at each target time, and the system can construct a time-level training loss, which provides a basis for subsequent time-weighted processing and model parameter updates.

[0087] In this step, the assimilated concentration sample is the simulated concentration value output by the assimilated concentration model from the previous time step in step 602, and the actual observed concentration is the real pollutant concentration value obtained from the observation data in step 601. The first loss function value is the quantified result calculated by the system for the error between the assimilated concentration sample and the corresponding observed concentration at a certain target time, reflecting the accuracy of the model prediction at that time.

[0088] In the implementation process, the system sequentially traverses each target time point corresponding to each optimization grid. For each target time point, the system extracts the assimilation concentration sample and the actual observed concentration value at that time. The system compares the two using a preset error calculation function, which may include mean squared error (MSE), absolute error (MAE), or Huber loss function. The system calculates the first loss function value for that target time point based on the error function and records this value for subsequent steps. A first loss function value is generated independently for each target time point.

[0089] Step 604: Based on the first loss function values ​​of all target times in each optimization grid, obtain the second loss function values ​​of each optimization grid.

[0090] Since it is necessary to integrate the error information from multiple time points into a single index after completing the error assessment at each target time point, to reflect the overall prediction bias of each optimization grid across the entire time series, step 604 provides a grid-level unified metric for subsequent integrated loss calculation and model training by constructing a second loss function value.

[0091] In an optional implementation, step 604 specifically calculates the second loss function value for each optimization grid in the following manner: Step 701: Assign influence weights to the first loss function value at each target time, where the influence weight of historical times that are closer to the current time is greater.

[0092] Step 702: Sum the first loss function values ​​with their corresponding influence weights to obtain the second loss function values.

[0093] Steps 701 and 702 are used to perform time-weighted processing on the first loss function values ​​at each target time point to calculate the second loss function values ​​for each optimization grid. This step is implemented because the model training process not only needs to consider the prediction errors of historical target times, but also needs to dynamically adjust the weight of each time point's error in the overall loss based on the time distance between different target times and the current time point. By emphasizing the prediction errors closer to the current time point, the model's ability to learn the features of the current time point can be improved, thus enhancing the model's timeliness and adaptability.

[0094] Step 701 assigns influence weights to the first loss function value at each target time, where the influence weight of historical times closer to the current time is greater. The system calculates the influence weight of the target time based on the time difference between the target time t-i and the current time t. The system calls a preset time weight function, such as an exponential decay function or a linear decay function, to generate corresponding weight values ​​based on the time difference. The system performs this operation for each target time separately, ensuring that the closer the target time is to the current time, the higher its corresponding weight, reflecting the greater contribution of that time to the current model update.

[0095] Step 702 involves weighted summation of each first loss function value with its corresponding influence weight to obtain the second loss function value. The system multiplies the first loss function value at each target time step by the influence weight determined in step 701, and then sums all the weighted loss function values. The system may optionally normalize the result to control the numerical range of the loss value. The final generated second loss function value represents the overall prediction error of the optimized grid after considering the time distance factor, and is used to guide the loss minimization objective of model training.

[0096] Step 605: Calculate the comprehensive loss function value based on the second loss function values ​​of all optimized grids.

[0097] After calculating the time-weighted error for each optimization grid, these error information needs to be aggregated to obtain a unified numerical index, which serves as the optimization objective for the current round of assimilation concentration model training. This comprehensive loss function value guides the direction of model parameter updates and serves as an important basis for determining whether the training has converged.

[0098] In this step, the second loss function value refers to the value obtained by weighting and summing the first loss function values ​​at multiple target times and the corresponding time weights for each optimization grid, representing the overall prediction error of that optimization grid in the time dimension. The comprehensive loss function value refers to the total loss obtained by integrating the second loss function values ​​of all optimization grids in the current training cycle according to a set method, and is used as the unified training objective for all optimization samples.

[0099] In the specific implementation process, the system first retrieves the second loss function values ​​of all optimized grids from the storage module. Assuming there are currently N optimized grids, the system labels their second loss function values ​​as L1, L2, ..., LN. The system then merges these N loss values ​​using a weighted average or arithmetic average to obtain a comprehensive loss function value. This comprehensive loss function value is then used as the objective function value for training the current assimilation concentration model and fed back to the training module to perform backpropagation and update the model parameters.

[0100] Step 606: Train the assimilation concentration model of the previous time step with the minimum integrated loss function value as the training objective to obtain the assimilation concentration model of the current time step.

[0101] Step 606 involves performing a supervised learning process on the assimilation concentration model from the previous time step, using the comprehensive loss function value as the optimization objective. This enables continuous updates to the model parameters, resulting in a model with stronger time adaptability and higher prediction accuracy.

[0102] In the specific implementation process, the system calls the assimilation concentration model from the previous time step and loads its parameters. The system uses the training samples constructed in steps 601 to 604, including the grid feature information of each optimized grid at each target time step as input, and the actual observed concentration as the output label. The system inputs these training samples into the model for forward propagation to obtain the assimilation concentration sample for each sample. The system calls the comprehensive loss function value as the objective function, executes the backpropagation algorithm, and adjusts the values ​​of each parameter in the model based on gradient descent. The system performs multiple iterations using a preset learning rate and number of training rounds until the comprehensive loss function value converges to a minimum or meets the preset training termination condition. The system writes the final parameters into the model to generate the assimilation concentration model for the current time step.

[0103] Reference Figure 4 , Figure 4 This is a schematic diagram of the structure of the deep learning-based variable grid atmospheric chemical data assimilation system provided by the present invention. The system includes: The acquisition module is used to acquire the grid feature information of all initial grids in the variable grid at the current time; the number of observation stations in each initial grid is continuously updated as time changes; The first processing module is used to input all the grid feature information at the current time into the assimilation concentration model at the current time to obtain the actual assimilation concentration at the current time output by the assimilation concentration model at the current time. The assimilation concentration model at the current time is trained on the assimilation concentration model at the previous time based on multiple training samples. The training samples include the grid feature information of the target time of at least one optimized grid. The target time includes the current time and historical times. The optimized grid is selected from multiple actual grids. The actual grid is the initial grid that includes at least one observation station at the current time.

[0104] In one possible implementation, the system further includes a second processing module for: For any real grid, construct the observation quality feature vector for each observation station. The observation quality feature vector includes at least two information quality indicators used to characterize the integrity and reliability of the observation data. Obtain the grid density corresponding to each actual grid; Based on all information quality indicators and grid density corresponding to each actual grid, all actual grids are divided into several clusters; Calculate the overall score for each actual grid cell in each cluster; The actual grid with the highest overall score is selected from each cluster and used as the optimized grid.

[0105] In one possible implementation, the second processing module is further configured to: Calculate the comprehensive fuzzy quality score for each actual grid based on all information quality indicators corresponding to each actual grid. The comprehensive fuzzy quality score and grid density of each actual grid are concatenated to construct the clustering input feature vector for each actual grid. A clustering algorithm based on the clustering input feature vector is performed on all actual grids to divide all actual grids into several clusters.

[0106] In one possible implementation, the second processing module is further configured to: For any observation station in each actual grid, each information quality index in its corresponding observation quality feature vector is input into the corresponding fuzzy membership function to obtain the fuzzy membership degree corresponding to each information quality index. The initial fuzzy quality score of the observation station is obtained by weighting all the fuzzy membership degrees corresponding to the observation station according to the preset fuzzy weight vector. The average of the initial fuzzy quality scores of all observation stations in each actual grid is used to obtain the comprehensive fuzzy quality score of each actual grid.

[0107] In one possible implementation, the second processing module is further configured to: For any cluster, the comprehensive fuzzy quality score of each actual grid is normalized to obtain the first normalized score of each actual grid, and the grid density of each actual grid is normalized to obtain the second normalized score of each actual grid. The first and second normalized scores of each actual grid are weighted and summed to obtain the comprehensive score of each actual grid.

[0108] In one possible implementation, the system further includes a model training module for: For any optimized grid, obtain its grid feature information at each target time and its actual observed concentration at each target time; the actual observed concentration is calculated based on the observation data of all observation stations in the optimized grid at the corresponding target time. The grid feature information of each target time is input into the assimilation concentration model of the previous time to obtain the assimilation concentration sample of each target time output by the assimilation concentration model of the previous time; among them, the assimilation concentration model of the initial time is an artificial intelligence model. Based on the error between the assimilated concentration sample and the corresponding actual observed concentration at each target time, calculate the first loss function value at each target time. Based on the first loss function values ​​of all target times in each optimization grid, the second loss function values ​​of each optimization grid are obtained; Calculate the combined loss function value based on the second loss function values ​​of all optimized grids; The assimilation concentration model of the previous time step is trained with the minimum comprehensive loss function as the training objective to obtain the assimilation concentration model of the current time step.

[0109] In one possible implementation, the model training module is further used for: Assign influence weights to the first loss function value at each target time, where the influence weight of historical times that are closer to the current time is greater; The second loss function value is obtained by weighting and summing each first loss function value with its corresponding influence weight.

[0110] In one possible implementation, the acquisition module is further configured to: extract simulation information of all initial grids at the current moment from the model simulation results of the variable grid, and acquire environmental information of the unit area where all initial grids are located at the current moment; the simulation information includes one or more of grid area, wind speed, temperature, humidity, air pressure and emissions; the environmental information includes one or more of terrain, population density, nighttime light intensity and road network density.

[0111] It should be noted that the deep learning-based variable grid atmospheric chemical data assimilation system provided by the present invention can execute the deep learning-based variable grid atmospheric chemical data assimilation method of any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0112] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a deep learning-based variable grid atmospheric chemical data assimilation method. This method includes: acquiring the grid feature information of all initial grids in the variable grid at the current time; the number of observation stations in each initial grid is continuously updated as time changes; inputting all the grid feature information at the current time into the assimilation concentration model at the current time to obtain the actual assimilation concentration at the current time output by the assimilation concentration model at the current time; wherein the assimilation concentration model at the current time is trained on the assimilation concentration model at the previous time based on multiple training samples; the training samples include the grid feature information of at least one optimized grid at the target time; the target time includes the current time and historical times; the optimized grid is selected from multiple actual grids, and the actual grid is the initial grid that includes at least one observation station at the current time.

[0113] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is able to execute the deep learning-based variable grid atmospheric chemical data assimilation method provided in the above embodiments.

[0115] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the deep learning-based variable grid atmospheric chemical data assimilation method provided in the above embodiments.

[0116] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the method of each embodiment or some parts of the embodiments.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for assimilating variable grid atmospheric chemical data based on deep learning, characterized in that, include: Obtain the current time-based mesh feature information of all initial meshes in the variable mesh; The number of observation stations within each initial grid is continuously updated over time; Input all the grid feature information at the current time into the assimilation concentration model at the current time to obtain the actual assimilation concentration at the current time output by the assimilation concentration model at the current time; The assimilation concentration model at the current time is obtained by training the assimilation concentration model at the previous time based on multiple training samples; the training samples include grid feature information of at least one optimized grid at the target time. The target time includes the current time and historical time; the optimized grid is obtained by screening multiple actual grids, and the actual grid is the initial grid that includes at least one of the observation stations at the current time.

2. The method for assimilation of variable grid atmospheric chemical data based on deep learning according to claim 1, characterized in that, Also includes: For any of the actual grids, an observation quality feature vector is constructed for each of the observation stations therein, the observation quality feature vector including at least two information quality indicators for characterizing the integrity and reliability of the observation data; Obtain the grid density corresponding to each of the actual grids; Based on all the information quality indicators and the grid density corresponding to each actual grid, all the actual grids are divided into several clusters; Calculate the overall score for each actual grid in each cluster; The actual grid with the highest comprehensive score is selected from each cluster and used as the optimized grid.

3. The method for assimilating variable grid atmospheric chemical data based on deep learning according to claim 2, characterized in that, The step of dividing all the actual grids into several clusters based on all the information quality indicators and the grid density corresponding to each actual grid includes: Calculate the comprehensive fuzzy quality score for each actual grid based on all the information quality indicators corresponding to each actual grid. The comprehensive fuzzy quality score and the grid density of each actual grid are concatenated to construct the clustering input feature vector of each actual grid; A clustering algorithm based on the clustering input feature vector is performed on all the actual grids to divide all the actual grids into several clusters.

4. The method for assimilating variable grid atmospheric chemical data based on deep learning according to claim 3, characterized in that, The step of calculating the comprehensive fuzzy quality score of each actual grid based on all the information quality indicators corresponding to each actual grid includes: For any observation station in each of the actual grids, each information quality index in the corresponding observation quality feature vector is input into the corresponding fuzzy membership function to obtain the fuzzy membership degree corresponding to each information quality index. The initial fuzzy quality score of the observation station is obtained by weighting all the fuzzy membership degrees corresponding to the observation station according to the preset fuzzy weight vector. The average of the initial fuzzy quality scores of all observation stations in each actual grid is used to obtain the comprehensive fuzzy quality score of each actual grid.

5. The deep learning-based variable grid atmospheric chemical data assimilation method according to claim 3, characterized in that, The calculation of the comprehensive score for each actual grid in each cluster includes: For any of the clusters, the comprehensive fuzzy quality score of each of the actual grids is normalized to obtain a first normalized score for each of the actual grids, and the grid density of each of the actual grids is normalized to obtain a second normalized score for each of the actual grids. The first normalized score and the second normalized score of each actual grid are weighted and summed to obtain the comprehensive score of each actual grid.

6. The method for assimilating atmospheric chemical data based on a variable grid according to claim 1, characterized in that, Also includes: For any of the optimized grids, obtain its grid feature information at each of the target times, and obtain its actual observed concentration at each of the target times; The actual observed concentration is calculated based on the observation data of all observation stations in the optimized grid at the corresponding target time. The grid feature information of each target time is input into the assimilation concentration model of the previous time to obtain the assimilation concentration sample of each target time output by the assimilation concentration model of the previous time; wherein, the assimilation concentration model of the starting time is an artificial intelligence model; Based on the error between the assimilated concentration sample and the corresponding actual observed concentration at each target time, calculate the first loss function value at each target time. Based on the first loss function values ​​of all target times in each optimization grid, the second loss function values ​​of each optimization grid are obtained; Calculate the comprehensive loss function value based on the second loss function values ​​of all the optimized grids; The assimilation concentration model at the previous time step is trained with the minimum comprehensive loss function value as the training objective to obtain the assimilation concentration model at the current time step.

7. The deep learning-based variable grid atmospheric chemical data assimilation method according to claim 6, characterized in that, The step of obtaining the second loss function value of each optimization grid based on the first loss function values ​​of all target times in each optimization grid includes: Assign influence weights to the first loss function value at each target time, wherein the influence weight of historical times that are closer to the current time is greater; The second loss function value is obtained by weighting and summing each of the first loss function values ​​with its corresponding influence weight.

8. The method for assimilating variable grid atmospheric chemical data based on deep learning according to claim 1, characterized in that, The process of obtaining the current-time grid feature information of all initial grids in the variable grid includes: Extract the simulation information of all initial grids at the current time from the model simulation results of the variable grid, and obtain the environmental information of the unit region where all the initial grids are located at the current time; The simulation information includes one or more of the following: grid area, wind speed, temperature, humidity, air pressure, and emissions. The environmental information includes one or more of the following: topography, population density, nighttime light intensity, and road network density.

9. A deep learning-based variable grid atmospheric chemical data assimilation system, characterized in that, include: The acquisition module is used to acquire the current mesh feature information of all initial meshes in the variable mesh; The number of observation stations within each initial grid is continuously updated over time; The first processing module is used to input all the grid feature information at the current time into the assimilation concentration model at the current time to obtain the actual assimilation concentration at the current time output by the assimilation concentration model at the current time; wherein, the assimilation concentration model at the current time is trained on the assimilation concentration model at the previous time based on multiple training samples; the training samples include grid feature information at the target time of at least one optimized grid. The target time includes the current time and historical time; the optimized grid is obtained by screening multiple actual grids, and the actual grid is the initial grid that includes at least one of the observation stations at the current time.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the deep learning-based variable grid atmospheric chemical data assimilation method as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based variable grid atmospheric chemical data assimilation method as described in any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based variable grid atmospheric chemical data assimilation method as described in any one of claims 1 to 8.