Background analysis method and system based on atmospheric greenhouse gas monitoring data
By generating intelligent agents and constructing complex states, and using diffusion models and multi-objective optimization techniques to separate interference signals, the problem of insufficient accuracy and robustness of atmospheric greenhouse gas monitoring data in existing technologies has been solved. This has enabled high-precision reconstruction of the background distribution, supporting climate change assessment and emission source apportionment.
Patent Information
- Application Number
- CN202511095393.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Existing baseline analysis of atmospheric greenhouse gas monitoring data suffers from insufficient accuracy and robustness in areas such as multi-source interference separation, boundary effect compensation, and high-precision baseline reconstruction.
By collecting data from monitoring equipment to generate an intelligent agent, a composite state of background data and interference data is constructed. Data reconstruction is performed using diffusion models and multi-objective optimization techniques. Combined with clustering and optimization strategies, the background signal is separated and reconstructed.
It achieves accurate and stable reconstruction of background distribution under complex interference and multi-source heterogeneous scenarios, improving scientific rigor and accuracy, and supporting greenhouse gas emission source apportionment and climate change assessment.
Smart Images

Figure CN120995134A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, specifically to a background analysis method and system based on atmospheric greenhouse gas monitoring data. Background Technology
[0002] With the increasing prominence of global climate change, atmospheric greenhouse gas monitoring and assessment has become one of the core technologies in environmental science and climate policy. Precise monitoring of atmospheric greenhouse gas concentrations (such as carbon dioxide) not only allows for tracking regional emission dynamics and assessing the effectiveness of emission reductions, but also provides data support for the formulation of carbon trading and environmental management policies. Currently, greenhouse gas monitoring systems primarily rely on various methods, including ground-based observation stations, mobile monitoring equipment, and satellite remote sensing. The data collected by these devices possess high spatiotemporal resolution and diverse attributes, reflecting the changing characteristics of atmospheric greenhouse gases in different regions and time periods.
[0003] However, due to variable meteorological conditions, complex geographical environments, and diverse data acquisition methods, actual monitoring data often contains multi-source interference and complex noise, such as anthropogenic emissions, sudden meteorological events, topographical influences, and sensor errors. These factors can lead to significant deviations between the raw observation data and the atmospheric background, thus affecting the accurate determination of greenhouse gas concentration background levels. Accurately distinguishing between background signals and interference signals, and restoring the natural state of the atmospheric environment, is one of the fundamental steps in long-term greenhouse gas change research, emission source apportionment, and policy evaluation.
[0004] Currently, methods for determining greenhouse gas background levels both domestically and internationally mainly rely on traditional approaches such as statistical analysis, outlier removal, spatiotemporal regression, and physical diffusion modeling. These methods have certain limitations in practical applications, particularly in multi-source data fusion, highly disruptive and complex scenarios, and high-precision background estimation, where they still face technical challenges such as insufficient algorithm robustness and limited handling of boundary effects. Summary of the Invention
[0005] In view of the above-mentioned problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by this invention is that the existing background analysis of atmospheric greenhouse gas monitoring data has insufficient accuracy and robustness in terms of multi-source interference separation, boundary effect compensation, and high-precision background reconstruction.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a background analysis method based on atmospheric greenhouse gas monitoring data, comprising:
[0008] Greenhouse gas monitoring data is collected using monitoring equipment; based on the distribution of the monitoring data, an intelligent agent is generated regarding the monitoring space.
[0009] The monitoring data of the intelligent agent is constructed as a composite state consisting of background data and interference data superimposed.
[0010] By performing evolutionary analysis of complex states among different agents, the baseline data of the monitoring space is reconstructed to obtain the baseline analysis results of atmospheric greenhouse gas monitoring data.
[0011] As a preferred embodiment of the background analysis method based on atmospheric greenhouse gas monitoring data described in this invention, the monitoring data includes greenhouse gas concentration data, collection time, geographic spatial location, and meteorological parameters.
[0012] The distribution includes dividing the monitoring space into monitoring units of size U×U×U, and using the monitoring data to cluster the coordinate points in each monitoring unit to obtain a cluster family in each monitoring unit; where U represents the side length of the monitoring unit.
[0013] Each cluster in the monitoring unit is treated as an intelligent agent;
[0014] When dividing the monitoring units, the division scheme with the most monitoring units is selected in the monitoring space; after the division is completed, the part outside the monitoring units is simplified and is not used as a reference for the intelligent agent to perform data analysis.
[0015] As a preferred embodiment of the background analysis method based on atmospheric greenhouse gas monitoring data described in this invention, the monitoring data of the intelligent agent includes taking the geographic spatial location of the cluster center of each intelligent agent as the geographic spatial location of the intelligent agent, and obtaining other monitoring data of the intelligent agent by calculating the mean value of the corresponding monitoring data at all coordinate points in the cluster.
[0016] As a preferred embodiment of the background analysis method based on atmospheric greenhouse gas monitoring data described in this invention, the following steps are taken: The complex state of the intelligent agent is modeled by superimposing the original values of the background data and interference data, performing data diffusion on the superposition result, and using the diffused result as the result of the complex state; the monitoring data of the intelligent agent is then made equal to the result of the complex state, thus completing the modeling of the complex state; expressed as: Z=A+B+∑ j∈Ei F(x i,j );
[0017] Where Z represents the monitoring data of greenhouse gas concentration of agent i, A represents the background data of agent i, B represents the interference data of agent i, j represents the index of agent i; Ei represents all agents in the monitoring space except i, representing the interference of other agents.
[0018] The data diffusion includes using a pre-trained diffusion model F(x) i,j ), quantify the data diffusion behavior; use the monitoring data of agent i and any neighboring agent j as F(x) i,j The input is F(x), which takes the monitoring data of agents i and j as a steady state and the change in the monitoring data of agent i as the input. i,j ) output.
[0019] As a preferred embodiment of the background analysis method based on atmospheric greenhouse gas monitoring data described in this invention, the evolutionary analysis includes: making a diffusion hypothesis for agents less than U away from the edge of the monitoring space based on the spatial characteristics of agents in the monitoring space; and, based on the diffusion hypothesis, performing background data evolution for each agent in the reconstructed monitoring space through an optimization strategy.
[0020] The diffusion hypothesis includes setting up n×m uniform virtual agents at a spatial location U outside the monitoring space; inputting the monitoring data of agents no more than 2U away from the edge of the monitoring space into a pre-trained neural network to generate virtual monitoring data at each of the virtual agents;
[0021] Where n represents the number of virtual intelligent agents preset per unit area, and m represents the area at the edge of the monitoring space.
[0022] As a preferred embodiment of the background analysis method based on atmospheric greenhouse gas monitoring data described in this invention, the optimization strategy includes taking the background data and interference data of all agents as a solution of the monitoring space, and generating the optimal solution through optimization of a multi-objective function;
[0023] The multi-objective function includes: each agent randomly combines background data and interference data, and after all agents have completed the combination, a solution is obtained; the optimal solution that satisfies three objectives is found in the solution space: the sum of the differences between the greenhouse gas concentration simulated by each agent and the actual value is minimized, the distribution characteristics of the interference data are most similar to the distribution characteristics of the anthropogenic emission rate, and the distribution characteristics of the background data are most similar to the distribution characteristics of the negative correlation between the greenhouse gas purification rate and the background data.
[0024] Represented as min(∑ E ΔZ); max(W(P) B ,Q));max(W(P A ,-H));
[0025] Where E represents all nodes in the reconstructed monitoring space; ΔZ represents the difference between the simulated monitoring data and the actual values for each agent; W(...) represents the Wasserstein distance; Q represents the distribution of greenhouse gas emission rates in the monitoring space; P B This indicates the distribution of interference data in the monitoring space; P A H represents the distribution of background data in the monitoring space; H represents the distribution of greenhouse gas purification rates in the monitoring space, W(P) A (-H) is P A The Wasserstein distance, which is negatively correlated with H, represents the distribution of numerical values relative to P after H takes negative values. A Wasserstein distance;
[0026] In the solution space, we obtain the set of all Pareto optimal solutions.
[0027] As a preferred embodiment of the background analysis method based on atmospheric greenhouse gas monitoring data described in this invention, the reconstruction of the background of the monitoring space includes, in the set of Pareto optimal solutions, treating each solution as an individual, the distribution of the background data of each individual as the feature of the individual, and clustering the features of all individuals.
[0028] If there exists a cluster whose number of individuals is greater than the number of individuals in all other clusters, then the cluster is output; if there is no cluster whose number of individuals is greater than the number of individuals in all other clusters, then the fitness function is used to select the cluster, the fitness of each cluster center is calculated, and the cluster corresponding to the maximum fitness value is output.
[0029] The baseline analysis results include taking the maximum and minimum values of the baseline data of each agent in the output cluster as the interval extreme values of the baseline data; and smoothing the interval of the baseline data of each agent in the monitoring space to obtain the baseline data of each data point in the monitoring space.
[0030] Fitness function: max(W(P) B ,Q))+max(W(P A ,-H))-min(∑ E ΔZ).
[0031] A background analysis system based on atmospheric greenhouse gas monitoring data using the method described in this invention, wherein:
[0032] The data acquisition unit collects greenhouse gas monitoring data through monitoring equipment; based on the distribution of the monitoring data, it generates an intelligent agent about the monitoring space.
[0033] The construction unit constructs the monitoring data of the intelligent agent into a composite state consisting of background data and interference data superimposed.
[0034] The analysis unit performs evolutionary analysis of complex states among different intelligent agents, reconstructs the background data of the monitoring space, and obtains the background analysis results of atmospheric greenhouse gas monitoring data.
[0035] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, it implements the steps of the method described in any one of the present invention.
[0036] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the present invention.
[0037] The beneficial effects of this invention are as follows: The background analysis method based on atmospheric greenhouse gas monitoring data provided by this invention can fully utilize multi-source monitoring data and achieve effective separation of background signals and interference signals through spatial partitioning and clustering, intelligent agent complex state modeling, data diffusion mechanisms, and multi-objective optimization. Compared with traditional statistical and physical modeling methods, this invention can more accurately and stably reconstruct the background distribution of atmospheric greenhouse gases in scenarios with complex interference, multi-source heterogeneity, and significant boundary effects, thereby improving the scientific rigor and accuracy of background determination. Furthermore, through multi-objective Pareto optimization and cluster fitness screening, this invention ensures that the output results not only closely match actual monitoring data but also possess strong physical rationality and spatial continuity. This helps provide more reliable data support for applications such as greenhouse gas emission source apportionment and climate change trend assessment, promoting the improvement of atmospheric environmental monitoring and management. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 The first embodiment of the present invention provides an overall flowchart of a background analysis method based on atmospheric greenhouse gas monitoring data. Detailed Implementation
[0040] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0041] Example 1, referring to Figure 1 As an embodiment of the present invention, a background analysis method based on atmospheric greenhouse gas monitoring data is provided, comprising:
[0042] S1: Collect greenhouse gas monitoring data through monitoring equipment; generate an intelligent agent about the monitoring space based on the distribution of the monitoring data.
[0043] Furthermore, the monitoring data includes greenhouse gas concentration data, collection time, geographic location, and meteorological parameters (including but not limited to temperature, humidity, wind speed, and air pressure, which are used to construct the diffusion function for the interaction between different intelligent agents in the following text).
[0044] The distribution involves dividing the monitoring space into monitoring units of size U×U×U. Using the monitoring data, the coordinate points within each monitoring unit are clustered to obtain a cluster family for each monitoring unit. This effectively uncovers spatial correlations and local features within data from different regions, improving the accuracy of data modeling. Here, U represents the side length of the monitoring unit. Each cluster family within the monitoring unit is treated as an intelligent agent. This significantly reduces the overall complexity of data processing and facilitates subsequent separation and dynamic modeling of background and interference signals in different spatial regions.
[0045] When dividing the monitoring space into monitoring units, the division scheme with the largest number of monitoring units is selected. After the division is completed, the areas outside the monitoring units are simplified and not used as a reference for data analysis by the intelligent agent. Prioritizing the scheme with the largest number of monitoring units during spatial division helps improve spatial resolution and preserves the details and diversity of the monitoring space to the greatest extent. Simplifying the areas outside the monitoring space avoids interference from low-density or invalid data at the boundaries, improving the scientific rigor, rationality, and computational efficiency of the background analysis, thus laying a solid foundation for the accurate reconstruction of greenhouse gas distribution and the extraction of background data.
[0046] The monitoring data for each agent includes using the geospatial location of its cluster center as its geospatial location, and calculating the mean value of the monitoring data at all coordinate points within the cluster for other monitoring data. This simplifies the data processing.
[0047] S2: Construct the monitoring data of the intelligent agent into a composite state consisting of background data and interference data superimposed;
[0048] The complex state of the agent is modeled by superimposing the original values of the background data and the interference data, then performing data diffusion on the superposition result, and using the diffused result as the result of the complex state; the monitoring data of the agent is then made equal to the result of the complex state, thus completing the modeling of the complex state; expressed as: Z=A+B+∑ j∈Ei F(x i,j By modeling the monitoring data of each agent as a superposition of background and interference data, the true composition of the observations can be accurately expressed, facilitating the dynamic separation of background and interference components in the future.
[0049] Z represents the monitoring data of greenhouse gas concentration of agent i, A represents the background data of agent i, B represents the interference data of agent i, j represents the index of agent i; Ei represents all agents in the monitoring space except i, representing the interference of other agents.
[0050] It should be noted that the data diffusion includes, through a pre-trained diffusion model F(x) i,j ), quantify the data diffusion behavior; use the monitoring data of agent i and any neighboring agent j as F(x) i,j The input is F(x), which takes the monitoring data of agents i and j as a steady state and the change in the monitoring data of agent i as the input. i,j The output of the model incorporates a data diffusion process to simulate the interaction and mass transfer of greenhouse gases in different spatial units within the atmosphere. By quantifying the diffusion of the superimposed results through a pre-trained diffusion model, spatial correlation and neighborhood influences can be captured. This allows the model to move beyond the linear superposition of single-point data and realistically reflect the dynamic changes in monitoring data caused by environmental factors such as gas flow and diffusion among different agents. This design not only improves the physical fit of the complex state model but also provides a theoretical and data foundation for subsequent background data reconstruction and interference source tracing, making the analysis results more interpretable, robust, and scientifically valuable.
[0051] In this embodiment, the diffusion model is a pre-trained fitting function model. Specifically, it is obtained by fitting and training a large amount of historical monitoring data to obtain an empirical function of the diffusion relationship between the monitoring data of the intelligent agents. This fitting function can be a linear regression model, a multinomial regression model, or other simple regression algorithms. The input in the diffusion process is the original monitoring data of the target intelligent agent and its neighboring intelligent agents, and the output is the monitoring data after the diffusion of the target intelligent agent. In other optional embodiments, the diffusion model can also be implemented using methods such as neural network models, physical diffusion equations, convolutional neural networks, random walks, or Gaussian process regression.
[0052] It's also important to understand that real-world greenhouse gas diffusion is essentially a process where gas concentrations in different regions gradually reach equilibrium through molecular motion and other means in space and time. In physical diffusion theory, the system eventually reaches a steady-state distribution, at which point the diffusion flow between points tends to be static or changes minimal. Therefore, when simulating data diffusion between agents, the true diffusion result of the gas under that spatial structure can only be represented when the monitoring data reaches a "steady state." If a data change at a certain step is directly taken as the output before diffusion stabilizes and the system is still in the adjustment phase, this change may be affected by initial value sensitivity or may not fully reflect the cooperative diffusion effect of all neighborhoods and the overall environment. Only in a steady state does the change in monitoring data have physical unique determinism and interpretability, truly reflecting the net result of the neighborhood diffusion process. After reaching a steady state, the data of agent i and neighboring agent j undergo multiple rounds of mutual influence and adjustment, and the output result is no longer sensitive to initial values and accidental perturbations, thereby improving the robustness and stability of the final diffusion output.
[0053] S3: By performing evolutionary analysis of complex states among different agents, the baseline data of the monitoring space is reconstructed to obtain the baseline analysis results of atmospheric greenhouse gas monitoring data.
[0054] By monitoring the spatial characteristics of agents in the monitoring space, a diffusion hypothesis is made for agents whose distance from the edge of the monitoring space is less than U; based on the diffusion hypothesis, the monitoring space is reconstructed, and the background data of each agent is evolved through an optimization strategy.
[0055] The diffusion hypothesis involves setting up n×m uniform virtual agents at a spatial location U outside the monitoring space. Monitoring data from agents located no more than 2U from the edge of the monitoring space (ensuring at least two agents exist perpendicular to the boundary, facilitating analysis of data trends) is input into a pre-trained neural network to generate virtual monitoring data at each agent's location. This data (actually, a simplified network is trained to analyze the greenhouse gas concentration distribution "from the outer layer inward" in current meteorological data, determining whether external emissions are increasing or decreasing within the monitoring space—essentially determining whether each analytical unit at the edge of the monitoring space is inhaling or releasing greenhouse gases) is doing so. Here, n represents the number of virtual agents per unit area, and m represents the area at the edge of the monitoring space. It's important to note that the agents at the edge of the monitoring space and their surrounding areas are often directly affected by non-background factors such as human activities and meteorological changes. Without modeling these external disturbances, the monitoring space's interference data can easily be underestimated or misjudged. By establishing n×m virtual agents outside the spatial boundary and using pre-trained neural networks to extrapolate the "interference" data of the virtual agents based on inputs such as edge agent data and meteorological characteristics, it is helpful to quantify and compensate for the input and output effects of external interference components on the internal agents. At observation points near the boundary, the collected gas concentration changes are often the result of a "mixture" of local background and external interference sources. Through virtual agents and diffusion hypotheses, the inflow (emission input) or outflow (purification output) of external interference into the monitoring space can be better decomposed and quantified, avoiding misjudging external disturbances as background changes and improving the accuracy of separating background from interference. With the assistance of outer-layer virtual agents in modeling, a "semi-open boundary" in the real physical environment of the atmosphere can be simulated, making the spatial distribution of interference data closer to actual emissions and environmental flow states, especially helpful for detecting and inverting the distribution of interference sources at the monitoring spatial boundary or downwind area.
[0056] In this embodiment, the neural network is a shallow feedforward neural network (FNN) or a simple multi-layer perceptron (MLP). This network structure is relatively simple, typically including one or two hidden layers with a moderate number of nodes per layer. It can achieve a non-linear fitting relationship between the input (such as monitoring data from boundary agents and virtual agents, meteorological characteristics, etc.) and the output (estimated values of interference data at the virtual agent). This neural network is easy to train quickly and is suitable for scenarios with small amounts of data or real-time inference, ensuring the model's efficiency and stability. In other optional embodiments, the neural network can also be a convolutional neural network, recurrent neural network, graph neural network, hybrid neural network, or other deep learning models, flexibly selected according to the monitoring spatial structure and data characteristics to adapt to the needs of different application scenarios for spatial inference and spatiotemporal modeling of interference data.
[0057] Furthermore, the optimization strategy includes treating the baseline data and interference data of all agents as a solution in the monitoring space, and generating the optimal solution through multi-objective function optimization. The multi-objective function involves each agent randomly combining the baseline data and interference data (ensuring that the values of both data satisfy preset constraints during combination), and completing the combination for all agents to obtain a solution. The optimal solution in the solution space is found to satisfy three objectives: minimizing the sum of the differences between the simulated greenhouse gas concentrations and actual values for each agent, maximizing the similarity between the distribution characteristics of interference data and the distribution characteristics of anthropogenic emissions, and maximizing the similarity between the distribution characteristics of the baseline data and the distribution characteristics of a negative correlation with greenhouse gas purification rates. By minimizing the total difference between the simulated greenhouse gas concentrations and actual monitored values for each agent, the optimization results are ensured to best reflect actual observations, guaranteeing the basic reliability of the model output. By maximizing the similarity between the distribution characteristics of interference data and the distribution characteristics of anthropogenic emissions, the impact of emission sources on the monitoring space can be better captured and reflected, helping to accurately identify and quantify the contribution of external interference to the overall environment. Since a higher background concentration of atmospheric greenhouse gases generally indicates a weaker purification or natural removal capacity, and vice versa, maximizing the similarity between the distribution characteristics of background data and the negative correlation distribution of greenhouse gas purification rates helps to strengthen the coupling between background and purification mechanisms at the distribution level, preventing model outputs from violating natural physical laws, thereby improving the scientific validity and practical explanatory power of the background reconstruction results.
[0058] Represented as min(∑ E ΔZ); max(W(P) B ,Q));max(W(P A ,-H)).
[0059] Meanwhile, the constraint that needs to be satisfied during the calculation is: the greenhouse gas absorption rate + emission rate within the monitoring space + diffusion result of emissions outside the monitoring space (the sum of the diffusion amounts of all intelligent agents in the original monitoring space by the "outer virtual intelligent agent") = 0
[0060] Where E represents all nodes in the reconstructed monitoring space; ΔZ represents the difference between the simulated monitoring data and the actual value for each agent (the simulated greenhouse gas concentration is Z). ’ =A ’ +B ’ +ΔF(x i,j In fact, B ’ +ΔF(x i,j ) is the real interference data, but B ’ +ΔF(x i,j B does not meet the emission distribution characteristics, therefore B ’ +ΔF(x i,j It can be split into two parts; A ’ B represents the background data analog quantity; ’ The analog quantity representing the interference data, i.e., the interference data before it spreads; ΔF(x) i,j ) represents the diffusion result during the simulation process, after all agents A have completed the process. ’ +B ’ After the arrangement, diffusion results are generated; via ZZ ’ That is, we obtain ΔZ for each agent; W(...) represents the Wasserstein distance; the Wasserstein distance is used to measure the similarity between the distribution of interference data and the distribution of emission rate. By maximizing the consistency of their distributions, the separated interference components are made to better match the real emission characteristics, which is convenient for interference source tracing and emission monitoring.
[0061] Q represents the distribution of greenhouse gas emission rates within the monitored space; P B This indicates the distribution of interference data in the monitoring space; P A H represents the distribution of background data in the monitoring space; H represents the distribution of greenhouse gas purification rates in the monitoring space, W(P) A (-H) is P A The Wasserstein distance, which is negatively correlated with H, represents the distribution of numerical values relative to P after H takes negative values. A The Wasserstein distance is then calculated. In the solution space, the set of all Pareto optimal solutions is obtained.
[0062] Furthermore, the reconstruction of the baseline of the monitoring space includes, in the set of Pareto optimal solutions, treating each solution as an individual, the distribution of the baseline data of each individual as the feature of the individual, and clustering the features of all individuals.
[0063] Prioritize selecting the baseline analysis results with the highest consensus and representativeness in the solution space to improve the robustness and objectivity of the final output. If the number of individuals in a cluster is greater than the number of individuals in all other clusters, then that cluster is output. If the number of individuals in a particular cluster is significantly greater than other clusters, it indicates that this cluster represents the common characteristics of most optimal solutions, reflecting the natural consistency reached between objectives during the multi-objective optimization process. Outputting this "main cluster" as the final result can effectively reduce the uncertainty caused by individual outliers, noisy solutions, or a few divergent solutions, improve the statistical credibility and environmental interpretability of the baseline reconstruction, and thus ensure the scientific rigor, stability, and application reliability of the greenhouse gas baseline analysis results. If no cluster has a number of individuals greater than the number of individuals in all other clusters, then a fitness function is used to select the cluster. The fitness of each cluster center is calculated, and the cluster corresponding to the maximum fitness value is output. When the optimal solutions in the solution space consist of multiple clusters of similar size, i.e., there is no clearly dominant master cluster, a fitness function is introduced for scientific selection to ensure that the final baseline analysis results have the highest overall goodness. By calculating the fitness of each cluster center separately, the performance of each cluster in terms of reconstruction accuracy, physical consistency, and distribution rationality can be systematically evaluated, avoiding the influence of randomness or local optima on the results. Finally, the cluster with the highest fitness value is selected for output, which not only improves the scientificity and objectivity of the analysis results, but also maximizes the balance between the multi-objective needs of monitoring data and environmental physical constraints, providing a more reliable and robust baseline reconstruction scheme for practical applications.
[0064] Fitness function: max(W(P) B ,Q))+max(W(P A ,-H))-min(∑ E ΔZ).
[0065] The baseline analysis results include: taking the maximum and minimum values of the baseline data for each agent in the output cluster as the interval extreme values of the baseline data; and smoothing the intervals of the baseline data for each agent in the monitoring space to obtain the baseline data for each data point in the monitoring space. Specifically:
[0066] First, the maximum and minimum baseline data values for each agent in the output cluster are used as the baseline data interval for the agent's representative location (coordinate point). This step effectively assigns interval labels to a finite number of known points in the monitoring space.
[0067] For each coordinate point in the monitoring space (regardless of whether it is the center of the agent), the baseline data of the coordinate point is generated and assigned according to the baseline data range of the surrounding agents by using spatial interpolation algorithms (such as nearest neighbor interpolation, inverse distance weighted interpolation, Kriging interpolation, etc.) or rules based on physical diffusion / correlation.
[0068] The interpolation rules can be set as follows: the nearest agent's interval extreme value; if it is within the coverage of multiple agent intervals, then use a weighted average or spatial weighting; or directly use spatial interpolation / spatial simulation methods to generate the background interval of new points based on surrounding known points.
[0069] In this way, after smoothing, each spatial coordinate point within the monitoring space is assigned a background data range, achieving a continuous representation of the background distribution across the entire space, rather than data only existing at the agent's coordinates. This "smoothing" step refers to using interpolation and mapping to smoothly extend the spatial data from a discrete distribution to a continuous distribution. This yields the range of greenhouse gas concentrations at each coordinate point throughout the entire space.
[0070] On the other hand, this embodiment also provides a background analysis system based on atmospheric greenhouse gas monitoring data, which includes:
[0071] The data acquisition unit collects greenhouse gas monitoring data through monitoring equipment; based on the distribution of the monitoring data, it generates an intelligent agent about the monitoring space.
[0072] The construction unit constructs the monitoring data of the intelligent agent into a composite state consisting of background data and interference data superimposed on each other.
[0073] The analysis unit performs evolutionary analysis of complex states among different intelligent agents, reconstructs the background data of the monitoring space, and obtains the background analysis results of atmospheric greenhouse gas monitoring data.
[0074] If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0075] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0076] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0077] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0078] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A background analysis method based on atmospheric greenhouse gas monitoring data, characterized in that, include: Greenhouse gas monitoring data is collected through monitoring equipment; Based on the distribution of the monitoring data, an intelligent agent is generated regarding the monitoring space; The monitoring data of the intelligent agent is constructed as a composite state consisting of background data and interference data superimposed. By performing evolutionary analysis of complex states among different agents, the baseline data of the monitoring space is reconstructed to obtain the baseline analysis results of atmospheric greenhouse gas monitoring data.
2. The background analysis method based on atmospheric greenhouse gas monitoring data as described in claim 1, characterized in that: The monitoring data includes greenhouse gas concentration data, collection time, geographic location, and meteorological parameters; The distribution includes dividing the monitoring space into monitoring units of size U×U×U, and using the monitoring data to cluster the coordinate points in each monitoring unit to obtain a cluster family in each monitoring unit; where U represents the side length of the monitoring unit. Each cluster in the monitoring unit is treated as an intelligent agent; When dividing the monitoring units, the division scheme with the most monitoring units is selected in the monitoring space; after the division is completed, the part outside the monitoring units is simplified and is not used as a reference for the intelligent agent to perform data analysis.
3. The background analysis method based on atmospheric greenhouse gas monitoring data as described in claim 2, characterized in that: The monitoring data of the intelligent agent includes taking the geospatial location of the cluster center of each intelligent agent as the geospatial location of the intelligent agent, and obtaining other monitoring data of the intelligent agent by calculating the mean value of the corresponding monitoring data at all coordinate points in the cluster.
4. The background analysis method based on atmospheric greenhouse gas monitoring data as described in claim 3, characterized in that: The complex state of the agent is modeled by superimposing the original values of the background data and the interference data, then performing data diffusion on the superposition result, and using the diffused result as the result of the complex state; the monitoring data of the agent is then made equal to the result of the complex state, thus completing the modeling of the complex state; expressed as: Z=A+B+∑ j∈Ei F(x i,j ); Where Z represents the monitoring data of greenhouse gas concentration of agent i, A represents the background data of agent i, B represents the interference data of agent i, j represents the index of agent i; Ei represents all agents in the monitoring space except i, representing the interference of other agents. The data diffusion includes using a pre-trained diffusion model F(x) i,j ), quantify the data diffusion behavior; use the monitoring data of agent i and any neighboring agent j as F(x) i,j The input is F(x), which takes the monitoring data of agents i and j as a steady state and the change in the monitoring data of agent i as the input. i,j The output of ).
5. The background analysis method based on atmospheric greenhouse gas monitoring data as described in claim 4, characterized in that: The evolutionary analysis includes making a diffusion hypothesis for agents whose distance from the edge of the monitored space is less than U by monitoring the spatial characteristics of agents in the space; Based on the aforementioned diffusion hypothesis, the reconstructed monitoring space uses an optimization strategy to evolve the background data for each agent. The diffusion hypothesis includes setting up n×m uniform virtual agents at a spatial location U outside the monitoring space; inputting the monitoring data of agents no more than 2U away from the edge of the monitoring space into a pre-trained neural network to generate virtual monitoring data at each of the virtual agents; Where n represents the number of virtual intelligent agents preset per unit area, and m represents the area at the edge of the monitoring space.
6. The background analysis method based on atmospheric greenhouse gas monitoring data as described in claim 5, characterized in that: The optimization strategy includes taking the baseline data and interference data of all agents as a solution in the monitoring space, and generating the optimal solution through optimization of a multi-objective function; The multi-objective function includes: each agent randomly combines background data and interference data, and after all agents have completed the combination, a solution is obtained; by finding the optimal solution in the solution space that satisfies three objectives: the sum of the differences between the greenhouse gas concentration simulated by each agent and the actual value is minimized, the distribution characteristics of the interference data are most similar to the distribution characteristics of the anthropogenic emission rate, and the distribution characteristics of the background data are most similar to the distribution characteristics of the negative correlation between the greenhouse gas purification rate and the background data. Represented as min(∑ E ΔZ); max(W(P) B ,Q)); max(W(P A ,-H)); Where E represents all nodes in the reconstructed monitoring space; ΔZ represents the difference between the simulated monitoring data and the actual values for each agent; W(...) represents the Wasserstein distance; Q represents the distribution of greenhouse gas emission rates in the monitoring space; P B This indicates the distribution of interference data in the monitoring space; P A H represents the distribution of background data in the monitoring space; H represents the distribution of greenhouse gas purification rates in the monitoring space, W(P) A (-H) is P A The Wasserstein distance, which is negatively correlated with H, represents the distribution of numerical values relative to P after H takes negative values. A Wasserstein distance; In the solution space, we obtain the set of all Pareto optimal solutions.
7. The background analysis method based on atmospheric greenhouse gas monitoring data as described in claim 6, characterized in that: The reconstruction of the baseline of the monitoring space includes, in the set of Pareto optimal solutions, treating each solution as an individual, the distribution of the baseline data of each individual as the feature of the individual, and clustering the features of all individuals; If there exists a cluster whose number of individuals is greater than the number of individuals in all other clusters, then the cluster is output; if there is no cluster whose number of individuals is greater than the number of individuals in all other clusters, then the fitness function is used to select the cluster, the fitness of each cluster center is calculated, and the cluster corresponding to the maximum fitness value is output. The baseline analysis results include taking the maximum and minimum values of the baseline data of each agent in the output cluster as the interval extreme values of the baseline data; and smoothing the interval of the baseline data of each agent in the monitoring space to obtain the baseline data of each data point in the monitoring space. Fitness function: max(W(P) B ,Q))+max(W(P A ,-H))-min(∑ E ΔZ).
8. A background analysis system based on atmospheric greenhouse gas monitoring data using the method described in any one of claims 1-7, characterized in that: The data acquisition unit collects greenhouse gas monitoring data through monitoring equipment; based on the distribution of the monitoring data, it generates an intelligent agent about the monitoring space. The construction unit constructs the monitoring data of the intelligent agent into a composite state consisting of background data and interference data superimposed. The analysis unit performs evolutionary analysis of complex states among different intelligent agents, reconstructs the background data of the monitoring space, and obtains the background analysis results of atmospheric greenhouse gas monitoring data.
9. A computer device, comprising: A memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.