Atmospheric water vapor content prediction method, system and equipment based on AI large model
By using a hybrid WT-LSTM model, combining wavelet transform and LSTM for multivariate nonlinear modeling, the nonlinear bottleneck of traditional methods in atmospheric water vapor content prediction is solved, achieving high-precision spatiotemporal prediction and supporting extreme weather early warning and agricultural water use optimization.
Patent Information
- Application Number
- CN202511165726.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional methods struggle to effectively handle the nonlinearity and multi-scale variability of atmospheric water vapor content, resulting in insufficient prediction accuracy. This is particularly true in complex climate zones, where existing AI models fall short in multi-factor integration and wavelet time-frequency decomposition, limiting the assessment of global multi-scale variability and regional contributions.
A hybrid WT-LSTM model is adopted, which combines wavelet transform multi-scale decomposition and LSTM multivariate nonlinear modeling. The spatiotemporal prediction of atmospheric water vapor content is carried out through an end-to-end framework. The model integrates total column water vapor, surface temperature, precipitation and evaporation data, uses wavelet transform to extract multi-scale feature subsequences, and performs multi-factor nonlinear training through LSTM.
It significantly improves the prediction accuracy of atmospheric water vapor content, with an interannual prediction R² increase of 36% and an RMSE decrease of 56.52%, supporting extreme weather early warning and agricultural water use optimization, and providing high-precision spatiotemporal prediction results.
Smart Images

Figure CN121578415A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of atmospheric remote sensing prediction technology, and in particular relates to methods, systems and equipment for predicting atmospheric water vapor content based on AI large models. Background Technology
[0002] Total Column Water Vapor (TCWV), a crucial component of the climate system, directly influences the water cycle, energy balance, and extreme weather events, playing a vital role in ecological monitoring, agricultural production, and weather forecasting. Against the backdrop of global warming, the spatiotemporal variability of TCWV has intensified, leading to more frequent extreme precipitation and drought events, severely threatening agricultural sustainability and water resource management. According to the World Meteorological Organization (WMO) 2025 report, events such as the Southeast Asian floods and the African drought highlight the central role of TCWV in driving climate dynamics; its changes directly affect crop yields and ecological stability by regulating precipitation and humidity fluctuations. Studies show that for every 1 kg / m³ increase in TCWV… 2 This could lead to an increase in precipitation in tropical regions by about 5-10%, thereby improving agricultural irrigation efficiency, but also amplifying the risk of flooding. Traditional methods rely on statistical models such as ARIMA or global climate models (GCMs), which can capture linear trends, but they are difficult to handle non-stationarity and multi-scale nonlinear coupling, resulting in insufficient prediction accuracy, especially in complex climate zones such as the tropical monsoon zone and the polar amplification effect region.
[0003] In recent years, the rise of reanalysis datasets such as ERA5 and NCEP / NCAR has provided high-resolution (0.25° × 0.25°) spatiotemporal data support for TCWV research. These datasets integrate satellite observations and ground station data, revealing the global trend of TCWV: an average increase of 0.78 kg / m² between 1959 and 2023. 2 Ten years, the highest in the tropics (1.29 kg / m³). 2 / ten years), the lowest in the frigid zone (0.16kg / m 2 ( / decade). Regional studies show that the TCWV in the South Asian monsoon region exhibits a 2-8 year cycle, influenced by the El Niño-Southern Oscillation (ENSO), with an amplitude of up to 10%; the frequency of droughts in western North America has increased by 15%, synchronized with the interdecadal variation of TCWV; seasonal moisture changes in Europe significantly regulate annual precipitation patterns; tropical moisture growth in the Amazon basin promotes rainforest precipitation; the dry-wet transition in Australia is driven by moisture input from the Indian Ocean; and the strengthening of TCWV in the East Asian typhoon region has led to a 25% increase in typhoon rainfall. These findings lay the foundation for regional moisture dynamics, but are mostly based on short-term (<30 years) and low-resolution data, limiting the assessment of global multi-scale variability and regional contributions.
[0004] Research on the driving mechanisms of TCWV emphasizes its nonlinear coupling with surface temperature, precipitation, and evaporation. In sub-Saharan Africa, moisture and temperature coupling exacerbates drought frequency by 25%; wind speed coupling in the East Asian monsoon region enhances typhoon precipitation intensity; positive feedback in the Amazon promotes precipitation growth; coupling in Australia drives wet-dry transitions; sea surface temperature coupling in the Indian Ocean increases extreme rainfall by 20%; and Arctic moisture input accelerates ice sheet melting, affecting mid-latitude weather. However, the differences in coupling between the tropics and mid-latitudes are not yet fully revealed, and single data sources limit system comparisons. Traditional forecasting models such as ARIMA struggle to capture nonlinear dynamics; GCMs simulate greenhouse gas and land-use change, but are computationally intensive and have high uncertainty. Deep learning, such as LSTM, shows great potential in time series forecasting; CNN and TCN improve regional accuracy; and random forests and support vector machines handle variability, but single-factor inputs ignore coupling, limiting accuracy under complex climatic conditions.
[0005] By 2025, AI models have made significant progress in atmospheric water vapor forecasting. Global AI weather models such as FourCastNet have achieved high-precision forecasts, while Fengwu and Pangu have been used for real-time ZTD (tropospheric delay, related to water vapor) estimation. Sensitivity analysis has validated AI's learning of atmospheric physics, such as spatiotemporal links. AI is applied to atmospheric river forecasting, combining with deep learning to reduce uncertainty and support water management; AI-Air optimizes urban forecasts in air quality forecasting systems; and deep learning architectures such as ConvLSTM are used for precipitation nowcasting. Machine learning reviews emphasize the evolution of ML / DL in weather and climate, including nonlinear modeling and data-driven Earth system science. Challenges include data scarcity, black-box problems, and computational resources, but AI shows advantages compared to NWP, such as in extreme weather forecasting. Although existing technologies have made progress, insufficient integration of multiple factors and wavelet time-frequency decomposition lead to poor handling of multi-scale variability, resulting in significant forecasting bottlenecks. Summary of the Invention
[0006] To address the aforementioned technical challenges, this invention proposes an atmospheric water vapor content prediction method, system, and device based on an AI large-scale model. It proposes a WT-LSTM hybrid large-scale model, integrating wavelet transform multi-scale decomposition with LSTM multivariate nonlinear modeling, significantly improving accuracy and overcoming the bottlenecks of traditional nonlinear prediction. This provides a scientific basis for extreme weather warnings, agricultural water use optimization, and climate adaptation. Through an end-to-end framework, it achieves global grid-level spatiotemporal prediction, supporting agricultural decision-making, such as tropical irrigation optimization and temperate drought warning.
[0007] To achieve the above objectives, this invention provides a method for predicting atmospheric water vapor content based on an AI large model, including: acquiring atmospheric water vapor content data and related meteorological variable data;
[0008] Wavelet transform was used to perform time-frequency decomposition on atmospheric water vapor content data and related meteorological variable data to extract multi-scale feature subsequences.
[0009] A large AI model is constructed based on a long short-term memory network. The extracted multi-scale feature subsequences are input into the large model for multi-factor nonlinear training to obtain the trained large model.
[0010] The trained large model is used to make spatiotemporal predictions of future atmospheric water vapor content, and the prediction results are obtained.
[0011] Optionally, time-frequency decomposition of atmospheric water vapor content data and related meteorological variable data using wavelet transform includes:
[0012] Discrete wavelet transform is used as the decomposition method;
[0013] Set the number of decomposition levels to 3-4;
[0014] Use the Daubechies mother wavelet as the basis function;
[0015] High-frequency and low-frequency subsequences are generated to capture short-term fluctuations and long-term trends of atmospheric water vapor, respectively.
[0016] Optionally, large-scale AI models built on long short-term memory networks include:
[0017] Construct a network structure containing two hidden layers;
[0018] Configure the Adam optimizer and mean squared error loss function;
[0019] An early stop mechanism is used to control the training process.
[0020] Optional variables for multi-factor nonlinear training integration include: total column water vapor data; surface temperature data; precipitation data; and evaporation data.
[0021] Optionally, the method further includes: using run theory to extract high water vapor events and low water vapor events; and quantifying the duration, frequency, intensity, and severity of the events.
[0022] Optionally, the output prediction results include: generating a spatiotemporal distribution map of atmospheric water vapor content; and displaying interannual trend curves and seasonal scale variations.
[0023] To achieve the above objectives, this invention provides an atmospheric water vapor content prediction system based on an AI large model, comprising:
[0024] The data acquisition module is used to acquire atmospheric water vapor content data and related meteorological variable data;
[0025] The feature extraction module is used to perform time-frequency decomposition on atmospheric water vapor content data and related meteorological variable data using wavelet transform, and extract multi-scale feature subsequences.
[0026] The model building and training module is used to build a large AI model based on the long short-term memory network. The extracted multi-scale feature subsequences are input into the large model for multi-factor nonlinear training to obtain the trained large model.
[0027] The prediction module is used to make spatiotemporal predictions of future atmospheric water vapor content using a trained large model, and obtain prediction results.
[0028] To achieve the above objectives, the present invention provides an atmospheric water vapor content prediction device based on an AI large model, comprising: a processor and a memory storing computer program instructions;
[0029] When the processor executes the computer program instructions, it implements the atmospheric water vapor content prediction method based on the AI large model.
[0030] Technical advantages of this invention: This invention discloses a method, system, and device for predicting atmospheric water vapor content based on an AI large-scale model. Through a WT-LSTM hybrid large-scale model, it can predict interannual R... 2 This invention achieves a 36% improvement and a 56.52% reduction in RMSE, overcoming the nonlinear bottleneck of traditional models and providing high-precision spatiotemporal predictions. The technology is applicable to global extreme weather early warning and supports agricultural water optimization, such as a 10% improvement in tropical irrigation efficiency. The modular design of the system and equipment facilitates expansion and deployment, enhancing climate adaptability. This invention is highly innovative, filling a gap in multi-scale nonlinear prediction and providing a scientific basis for ecological and agricultural management. Attached Figure Description
[0031] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0032] Figure 1 The technology roadmap includes (a) multi-source data integration, (b) driving factors and trend analysis, and (c) TCWV prediction and validation.
[0033] Figure 2 Flowchart of the WT-LSTM model;
[0034] Figure 3 This is a schematic diagram of continuous wavelet transform analysis of surface factors, where a is the total water column vapor, b is the surface temperature, c is the total precipitation, d is the evapotranspiration, e is the 10m U component, and f is the 10mV wind component.
[0035] Figure 4This is a schematic diagram of cross-wavelet transform analysis of TCWV and surface factors, where a is TCWV, b is surface temperature, c is TCWV, d is total precipitation, e is large TCWV, f is evapotranspiration, g is TCWV, (h) is the 10m U-Wind component, i is TCWV, and j is the 10m V wind component.
[0036] Figure 5 The seasonal trend frequency variations of four factors are: (a) total atmospheric water vapor, (b) temperature, (c) precipitation, and (d) evaporation.
[0037] Figure 6 The decomposition and prediction of 780 grid points are shown, where a is the A1 high-frequency sequence, b is the D1 low-frequency sequence, c is the D2 low-frequency sequence, d is the D3 low-frequency sequence, e is the final prediction result of the WT-LSTM hybrid model, and f is the final prediction result of the LSTM single model.
[0038] Figure 7 The diagram shows the detailed metrics and comparative analysis of the predictions for all grid points, where a represents the prediction results of the WT-LSTM hybrid model and the LSTM single model; b represents the percentage improvement in prediction performance of the WT-LSTM hybrid model compared to the LSTM single model.
[0039] Figure 8 The following is a scatter plot of site verification results for different regions: a) China; b) USA; c) Australia; d) Argentina; e) France; f) Kenya.
[0040] Figure 9 Verify the Taylor chart for the site, where a represents China, b represents the United States, c represents Australia, d represents Argentina, e represents France, and f represents Kenya;
[0041] Figure 10 Compare the prediction results for all grid points using Taylor plots, where a is the WT-LSTM hybrid model and b is the LSTM single model;
[0042] Figure 11 This is a prediction of atmospheric water vapor changes over the next 36 months, where a is a comparison of historical data and predicted atmospheric water vapor, b is the interannual trend of atmospheric water vapor, and c is the kernel density distribution of atmospheric water vapor changes. Detailed Implementation
[0043] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0044] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0045] like Figure 1 As shown, this embodiment provides a method for predicting atmospheric water vapor content based on an AI large model, including:
[0046] Step 1: Data Preparation and Preprocessing. First, acquire atmospheric water vapor content data and related meteorological variable data. Specifically, Figure 1 (a) presents multi-source data integration, including the Generation 5 Reanalysis Dataset (ERA5) from the European Centre for Medium-Range Weather Forecasts (ECMWF), which provides hourly global hourly estimates with a spatial resolution of 0.25° × 0.25° and covers the period from 1959 to 2023; the NCEP / NCAR Reanalysis Dataset from the National Oceanic and Atmospheric Administration (NOAA) and the National Center for Atmospheric Research (NCAR), as a supplementary perspective with a resolution of 2.5° × 2.5°; and vertical profile observations from the NOAA Integrated Global Radiosonde Archive (IGRA), including temperature, humidity, and pressure, used for independent validation and calibration of the reanalysis data. IGRA stations are distributed across major global climate zones, such as the tropics (Amazon Basin), subtropics (Indian subcontinent), temperate zones (North American Great Plains), and polar zones (Arctic Greenland), with more than six stations in total, ensuring data representativeness. Key variables extracted include total column water vapor (TCWV, kg / m³). 2 The variables include surface temperature (Temp, in °C), precipitation (TP, in m), potential evapotranspiration (Evap, in m), and the 10-meter wind speed component (U10 / V10, in m / s). These variables characterize the interaction between the surface and the atmosphere. Precipitation and evapotranspiration reflect the dynamics of the water cycle, temperature drives evaporation, and wind speed affects water vapor transport.
[0047] The preprocessing steps include data cleaning, missing value handling, and standardization. First, outliers (such as cloud-polluted pixels in ERA5) are filtered using quality control markers. Missing values are filled using linear interpolation or nearest neighbor methods, with a filling rate of <5%. Then, spatiotemporal alignment is performed by upsampling the NCEP / NCAR data to 0.25° resolution using bilinear interpolation to align with ERA5. Standardization uses the Z-score method, with the formula z = (x - μ) / σ, where μ is the mean and σ is the standard deviation, ensuring consistent variable scales and avoiding gradient explosion. Simultaneously, a global grid dataset is constructed with a total sample size exceeding 10^6 monthly records, covering all seven continents and four oceans, divided into latitudinal zones (tropical 23.5°S-23.5°N, subtropical 23.5°-35°N / S, temperate 35°-66.5°N / S, and polar >66.5°N / S) to reflect climatic heterogeneity. This step also includes auxiliary data fusion, such as weighted average fusion of ERA5 and NCEP / NCAR, with weights based on the observation error covariance, using the formula \hat{x}=W_1x_1+W_2x_2, W_i=\sigma_j^2 / (\sigma_1^2+\sigma_2^2), which improves data accuracy by approximately 10%, especially in high-latitude ice sheet regions and tropical convection active regions. After data preparation, the total dataset size is approximately TB, supporting subsequent multi-scale analysis.
[0048] Step 2: Time-frequency feature extraction and decomposition. Figure 1 (b) involves the analysis of driving factors and changing trends. Wavelet transform is used to perform time-frequency decomposition on the preprocessed data to extract multi-scale feature subsequences. This step employs Discrete Wavelet Transform (DWT), selecting the Daubechies mother wavelet (db3 or db4) as the basis function, and setting the decomposition level to 3-4 levels to balance resolution and computational efficiency. The mathematical basis of DWT is W_{j,k}=∫f(t)ψ_{j,k}(t)dt, where ψ_{j,k}(t)=2^{-j / 2}ψ(2^{-j}tk), j is the scale, and k is the displacement. A high-pass filter is used to extract detail signals (high-frequency components, such as A1 capturing seasonal fluctuations over 6-12 months) and a low-pass filter is used to extract approximate signals (low-frequency components, such as D1-D3 capturing interannual and interdecadal trends). For TCWV sequences, the focal period is 8-16 months (related to ENSO and monsoon rains) and 128 months (long-term climate change), which are decomposed into A1 (high-frequency local variation), D1 (short-term trend), D2 (medium-term cycle) and D3 (long-term low frequency).
[0049] To enhance the analysis, this step also incorporates Continuous Wavelet Transform (CWT) as an auxiliary method, with the formula W_f(a,b)=(1 / √a)∫f(t)ψ^((tb) / a)dt, used for time-spectrum visualization to reveal the energy distribution of TCWV in the 12-month dominant cycle (seasonal cycle) and the 6-month secondary cycle (semi-annual oscillation). High-energy regions are concentrated in the 24–36 months of 1970–1980, the 48–60 months and 120–144 months of 1980–1990, the 48–60 months of 2000–2010, and the 6–12 months of 2015–2023, consistent with monsoon rains and ENSO driving forces. Cross-wavelet transform (XWT) further quantifies the coherence between TCWV and driving factors, with the formula W_{xy}(a,b)=W_x(a,b)W_y^(a,b), revealing the phase relationship of TCWV lag temperature by 1-2 months, supporting the evaporation feedback mechanism. Wavelet reconstruction verifies the decomposition accuracy, with the reconstructed signal formula ∇f(t)=\sum W_{j,k}ψ_{j,k}(t), and the error <5×10^{-4}. Feature vectors with tens of dimensions are extracted, including peak energy, phase lag, and periodic intensity, supporting AI input. This step is implemented in a Python environment using the PyWavelets library, with a processing time of approximately minutes, significantly reducing data noise and improving the multi-scale representation of non-stationary signals.
[0050] Step 3: AI Large-Scale Model Construction and Training. A large-scale AI model is constructed based on a Long Short-Term Memory (LSTM) network, using the extracted time-frequency feature subsequences as input for multi-factor nonlinear training. The large-scale model architecture is designed as an end-to-end framework, including an input layer (multivariate feature vectors), two hidden layers (128 units in the first layer and 64 units in the second layer, using the ReLU activation function), a Dropout layer (with a dropout rate of 0.2 to prevent overfitting), and an output layer (linear regression outputting TCWV predicted values). The core mechanism of LSTM captures long-term dependencies through gating units. The forget gate formula is f_t = σ(W_f·[h_{t-1},x_t] + b_f), the input gate is i_t = σ(W_i·[h_{t-1},x_t] + b_i), the output gate is o_t = σ(W_o·[h_{t-1},x_t] + b_o), the cell state is C_t = f_t*C_{t-1} + i_t*tanh(W_c·[h_{t-1},x_t] + b_c), and the hidden state is h_t = o_t*tanh(C_t). The input integrates multiple factors such as TCWV, Temp, TP, and Evap, and is transformed into a low-dimensional vector through the embedding layer to support nonlinear modeling.
[0051] The training process uses the Adam optimizer with a learning rate of 0.001 (dynamic decay), a batch size of 32, and the loss function is the mean squared error MSE = (1 / n)∑(y - ŷ)^2. L2 regularization (λ = 0.001) is added to prevent overfitting. Dataset division: 80% for the training set (from 1959 to 2011), 20% for the test set (from 2012 to 2023), and a 20% validation subset is drawn from the training set. Train for 100 epochs and set an early stopping mechanism (patience of 10 epochs, monitoring the validation loss). The innovation lies in multi-task learning. Each wavelet subsequence independently trains an LSTM branch and is then integrated through weighted coefficients. The weights w_i = 1 / var(y_i), and the low-frequency D3 has a higher weight. This step embeds physical constraints, such as the Clausius-Clapeyron relationship, as an auxiliary loss L_phy = λ|dTCWV / dTemp - 7%*TCWV / Temp| to improve interpretability. The training is executed on a GPU cluster (NVIDIA A100). The total number of parameters is in the millions. After convergence, the model supports parallel prediction for global grid points, and the time is < hour level. The accuracy of this large model is annual R 2 is 0.941, and in summer it is 0.949, which is better than the single-factor LSTM.
[0052] Step 4: Spatiotemporal prediction and extreme event analysis. Use the trained large model to perform spatiotemporal prediction on the future atmospheric water vapor content and output the prediction results. Input future scenario data (such as temperature and precipitation projections under CMIP6 RCP4.5), generate a 0.25° grid-level TCWV map with a time span of 36 months. The prediction formula is ŷ = ∑ w_i·LSTM(D_i), and the inverse DWT reconstruction is f̂(t) = ∑Ŵ{j,k}ψ{j,k}(t). The results show that the future TCWV will increase slowly, with a significant increase in summer, consistent with the historical trend (0.78 kg / m 2 / decade). The spatiotemporal distribution is visualized as a heat map, supporting analysis of latitude bands, such as an increase of 1.29 kg / m 2 / decade in the tropics.
[0053] Extreme event analysis uses the run theory to extract high / low water vapor events: The standardized sequence z = (x - μ) / σ, with thresholds at the 90% and 10% quantiles. A run is identified as a subsequence that continuously satisfies z > thresh_high or z < thresh_low. Quantify the duration dur = end - start, the frequency freq = count / year, the intensity int = ∑|x_t - thresh| / dur, and the severity sev = int * freq. The global high water vapor event lasts for 2 months with an intensity of 7 kg / m 2 , and the severity in the tropical monsoon region is 72 kg / m2 • Per event per year, supporting a 5-10% increase in rice yield; low moisture event intensity 5 kg / m 2 In temperate arid regions, there is a 10% risk of yield reduction. Seasonal analysis shows that high moisture levels and flooding in the Amazon during summer reduce yields by 10%, while low moisture levels in the Arctic during winter inhibit ecological recovery. This analysis is integrated into forecasts, providing early warning reports to support agricultural applications such as irrigation optimization.
[0054] Step 5: Model validation and optimization. Figure 1 (c) is for TCWV prediction and validation, and the model performance is evaluated through multiple indicators: MAE = (1 / n)\sum|y-\hat{y}|, RMSE = sqrt(MSE), MAPE = (1 / n)\sum|(y-\hat{y}) / y|, MASE = MAE / MAE_naive, R 2 =1-SSE / SST. Taylor plots visualize standard deviation, correlation, and error; WT-LSTM points are closer to the reference circle. Site validation selects 6 sites (e.g., East Asia site in China, North America site in the United States), cross-validation K=5 fold, R... 2 >0.92. Optimizations included grid search hyperparameters (learning rate 0.0001-0.01) and sensitivity analysis to quantify variable contributions (Temp 30%). Compared to LSTM, WT-LSTM achieves a 46% reduction in MAE, making it suitable for global early warning systems. The overall implementation is in Python / TensorFlow and supports scaling to higher resolution data.
[0055] The WT-LSTM model modeling process is as follows: Figure 2 As shown, the process includes data preprocessing, DWT decomposition to generate subsequences A1 and D1-D3, LSTM training, and inverse DWT prediction and reconstruction. The dataset was divided into a training set (January 1959 to December 2011, 80%) and a test set (January 2012 to December 2023, 20%), with the 20% of the training set reserved as a validation set for hyperparameter optimization. The LSTM model contained two hidden layers (128 and 64 units respectively), with a learning rate of 0.001, a batch size of 32, and a dropout rate of 0.2 to enhance generalization ability. The model used the Adam optimizer, trained for 100 epochs based on the mean squared error (MSE) loss function, and implemented an early stopping mechanism after 10 epochs to prevent overfitting. Model performance was evaluated using various metrics, including mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), mean absolute percentage error (MAPE), correlation coefficient (R), and coefficient of determination (R²). 2 ), to quantify the prediction accuracy of the test set.
[0056] like Figure 3denoted as Continuous wavelet transform analysis of surface factors, where a is the total water column vapor, b is the surface temperature, c is the total precipitation, d is the evapotranspiration, e is the 10m U component, and f is the 10mV wind component.
[0057] like Figure 4 As shown, cross wavelet transform (XWT) and wavelet coherence (WTC) analysis further revealed the interaction between TCWV and surface variables in the time-frequency domain. A color gradient represents the resonance intensity, a thick black line indicates the 95% confidence interval, and arrows indicate the direction of the phase difference. The association between TCWV and key meteorological variables has been quantified, providing fundamental support for time-frequency dynamics. Figure 4 The TCWV and surface temperature exhibit strong resonance within 8-16 month and 128 month cycles, with significantly enhanced short-cycle fluctuations after 1970. TCWV lags behind temperature by approximately 1-2 months, reflecting that temperature changes indirectly drive water vapor accumulation through seasonal evaporation. Figure 4 b) indicates that WTC analysis confirms a high coherence between the two within an 8-16 month period. The lag effect may stem from increased evaporation rates due to surface heating, consistent with the strong positive correlation between TCWV and temperature in tropical regions (such as the western mountainous areas of the North American Great Plains) (r>0.6, p<0.05). Increased evaporation caused by rising temperatures supports an increase of approximately 5-8% in rice and maize yields. Figure 4 The c-value indicates that TCWV and precipitation show significant resonance within an 8-16 month cycle. Between 2008 and 2018, TCWV lagged behind precipitation by approximately one month, suggesting that precipitation replenishes soil moisture and then enhances water vapor content through evaporation feedback. Figure 4 The analysis of d for WTC reveals a high synchronicity between the two variables within this cycle, echoing the significant correlation between TCWV and precipitation in subtropical monsoon regions (such as the Indian subcontinent) (r = 0.65, p < 0.05). Precipitation-driven water vapor accumulation effectively ensures irrigation stability and reduces the drought risk in northwestern India. Figure 4 The value of e, TGWV, and evaporation show a high degree of synchronization within an 8-16 month cycle, such as... Figure 4 The f-value for WTC highlights the driving role of seasonal evaporation peaks on TCWV, consistent with the strong correlation between TCWV and evaporation in temperate oceanic regions (such as the North Atlantic and Northwest Pacific) (r>0.6, p<0.05). This process provides crucial water vapor supply for agricultural irrigation in humid Southeast Asian regions (such as the Mekong Delta in Vietnam), increasing water vapor transport efficiency by approximately 10-15% during the peak summer evaporation season. Figure 4 The g and i values represent the TCWV and the 10-meter wind speed components U10 and V10, respectively. These components show some resonance within an 8-16 month period, but do not reach the significance threshold within a longer period of 104-256 months. For example... Figure 4The h and j values, respectively, from WTC analysis reveal high synchronicity within short periods, indicating that wind speed promotes water vapor transport in the short term, possibly modulated by changes in airflow patterns induced by global warming. However, the correlation between TCWV and wind speed is weak in cold regions (such as the Arctic Circle and Antarctic coast) (r < 0.6, p > 0.05). XWT and WTC results further confirm that the contribution of wind-driven short-period water vapor transport is limited, and the inhibitory effects of ice and snow cover and polar vortex are particularly evident in the arid zone around the Sahel.
[0058] like Figure 5 The seasonal trend frequency variations of four factors are further illustrated: (a) total atmospheric water vapor, (b) temperature, (c) precipitation, and (d) evaporation. All variables in the four seasons (spring, summer, autumn, and winter) conform to a normal distribution, with TCWV and surface temperature showing particularly pronounced distribution characteristics and a more obvious central tendency. The range of trend values is consistent with the slope, with TCWV mainly distributed within -2 × 10⁻⁶. -4 Up to 2×10 -4 kg / m 2 / year, with peak values concentrated between 0 and 1×10 -4 kg / m 2 / year reflects the dynamic changes in water vapor driven by seasonal climate. Surface temperature and precipitation trend values are mostly concentrated between 0 and 2 × 10⁻⁶. -4 ℃ / year and 2×10 - 8 mm / year, with higher values in summer and spring likely related to global warming and increased rainfall, while values are lower in winter. Evaporation trend values are mainly distributed in the range of -5×10. -8 The negative values, down to 0 mm / year, predominate, suggesting a weakening of the seasonality of water loss, with strong symmetry in the seasonal distribution. The Shapiro-Wilk test further confirms the normal distribution of all variables (p>0.05), indicating that these trend values are statistically robust.
[0059] like Figure 6 The decomposition and prediction of 780 grid points are presented, where a represents the A1 high-frequency sequence, b represents the D1 low-frequency sequence, c represents the D2 low-frequency sequence, d represents the D3 low-frequency sequence, e represents the final prediction result of the WT-LSTM hybrid model, and f represents the final prediction result of the LSTM single model. The predicted values of each subsequence and the final result are compared with the actual observed values. LSTM exhibits excellent performance for different subsequences, with the D3 subsequence (MAE 0.166, R...) showing the highest performance. 2The 0.968 value was the best performing LSTM model, particularly in low-frequency long-term trend modeling, highlighting its advantage in capturing long-term water vapor variations. This superior performance is attributed to the high efficiency and low computational complexity of LSTM in feature extraction, demonstrating significant advantages in multi-scale time series modeling. Finally, through wavelet reconstruction, the multi-factor WT-LSTM hybrid model outperformed the single LSTM model in all metrics, significantly improving prediction accuracy. The reconstructed curves showed that the multi-factor WT-LSTM predictions closely matched the actual observations at peaks and troughs, while the single-factor LSTM predictions performed poorly at these points, failing to accurately capture extreme water vapor variations. Based on the analysis results of the 780 grid points, the model performance was comprehensively validated on a larger scale by extending the analysis to all grid points.
[0060] like Figure 7 Figure 'a' compares the performance of the single-factor LSTM and WT-LSTM models in predicting changes across all grid points of TCWV using a bar chart, evaluating their predictive capabilities at interannual and seasonal scales. The single-factor LSTM model shows better accuracy at the interannual scale (MAE 0.5 kg / m²). 2 R 2 The value of 0.688 can capture the overall trend of TCWV, but it is insufficient in predicting seasonal variations, especially in spring (RMSE 9.262 kg / m³). 2 R 2 0.643) and winter (RMSE 8.498kg / m 2 R 2 The error (0.693) is relatively high, and the correlation is weak. Table 1 shows that its comprehensive index is RMSE 8.726 kg / m³. 2 MAE 4.434kg / m 2 MAPE 0.318, MASE 10.707. In contrast, WT-LSTM, by combining wavelet transform (WT) decomposition of periodic components with LSTM nonlinear modeling, significantly improves prediction accuracy, reducing the interannual RMSE to 3.793 kg / m². 2 MAE up to 2.394 kg / m 2 MAPE to 0.204, MASE to 5.781, R 2 Reaching 0.941. In terms of seasonal performance, WT-LSTM performed better in summer (RMSE 3.479 kg / m³). 2 R 2 0.949) and autumn (RMSE 2.917kg / m 2 R 2 The RMSE (0.961) is particularly excellent, especially in spring (RMSE 4.349 kg / m³). 2 R 20.921) and winter (RMSE 4.245kg / m 2 R 2 The 0.925 value is also significantly improved, highlighting its advantage in capturing the seasonal cycle and nonlinear characteristics of TCWV.
[0061] like Figure 7 Table b further illustrates the percentage performance improvement of WT-LSTM compared to single-factor LSTM in TCWV prediction, with Table 2 listing quantitative indicators for interannual and seasonal scales. At the interannual scale, WT-LSTM shows a significant improvement in accuracy, with MAE increasing from 0.5 kg / m³. 2 Reduced to 0.3 kg / m 2 (Decreased by 46.01%), RMSE from 8.726 kg / m 2 Reduced to 3.793 kg / m 2 (Decreased by 56.52%), MAPE decreased from 0.318% to 0.204% (decreased by 35.84%), MASE decreased from 10.707 to 5.781 (decreased by 45.99%), R 2 The MAE improved from 0.688 to 0.941 (an improvement of 36.77%). On a seasonal scale, autumn showed the best performance, with MAE decreasing to 1.966 kg / m³. 2 (Reduced by 52.10%), RMSE decreased to 2.917 kg / m³. 2 (reduced by 63.35%), MAPE decreased by 39.60%, R 2 The MAE reached 0.961 (an improvement of 55.84%); summer was the second best season, with MAE decreasing to 2.273 kg / m³. 2 (Reduced by 50.63%), RMSE decreased to 3.479 kg / m³. 2 (reduced by 60.12%), MAPE decreased by 43.30%, R 2 The MAE reached 0.949 (an improvement of 38.54%). Improvements were similar in spring and winter, with the MAE decreasing to 2.727 kg / m³ in spring. 2 (Reduced by 44.26%), RMSE decreased to 4.349 kg / m³. 2 (reduced by 53.04%), MAPE decreased by 33.90%, R 2 The MAE reached 0.921 (an improvement of 43.23%); the MAE decreased to 2.610 kg / m³ in winter. 2 (Reduced by 37.11%), RMSE decreased to 4.245 kg / m². 2 (reduced by 50.04%), MAPE decreased by 25.09%, R 2 The value reached 0.925 (an improvement of 33.47%).
[0062] Table 1
[0063]
[0064] Table 2
[0065]
[0066] To verify the performance of the WT-LSTM mixture model in predicting TCWV changes from 1959 to 2023, meteorological station observation data were used for evaluation based on six representative stations listed in Table 3 (located in China, the United States, Australia, Argentina, France, and Kenya, covering tropical, temperate, and arid climate zones). Figure 8 The scatter plot visually demonstrates the fit between the WT-LSTM predictions and the observed values. The predicted points are closely clustered around the diagonal, indicating high consistency of the model across all sites. Figure 9 Taylor plots further quantified the model performance, showing the superior performance of predicted and observed values in terms of standard deviation, correlation coefficient, and root mean square error, particularly at the French and Argentine sites. In terms of correlation coefficient (R²),... 2 Regarding WT-LSTM, it has applications in China, the United States, Australia, Argentina, France, and Kenya. 2 The values were 0.936, 0.948, 0.929, 0.976, 0.981, and 0.915, respectively, all higher than 0.9, reflecting the high correlation and robustness of the model across different climate zones. The mean absolute error (MAE) assessment showed that the MAEs for each station were 0.413, 0.189, 0.531, 0.231, 0.138, and 0.368 kg / m³, respectively. 2 The average MAE is less than 0.4 kg / m 2 (range 0.138-0.531 kg / m) 2 Overall, the WT-LSTM predictions are in high agreement with the observed values, especially in France (R). 2 0.981, MAE 0.138kg / m 2 ) and Argentina (R 2 0.976, MAE 0.231kg / m 2 Its performance was particularly outstanding, validating its superior ability to predict TCWV under diverse global climatic conditions.
[0067] Table 3
[0068]
[0069] like Figure 10Taylor plots comparing the prediction results for all grid points are presented, where a represents the WT-LSTM hybrid model and b represents the LSTM single model. The WT-LSTM predictions are closer to the reference circle, the standard deviation is highly consistent with the observed values, the correlation coefficient is better, and the root mean square error is significantly reduced. These characteristics indicate that combining multi-factor data with wavelet transform significantly enhances the model's accuracy and stability. WT-LSTM, with its superior information extraction and fitting capabilities, especially in capturing complex climate trends, has established itself as the preferred model for atmospheric water vapor prediction.
[0070] Figure 11 Historical data was used to predict atmospheric water vapor changes over the next 36 months, and the results were explained through visualization analysis. Figure 11 Figure 'a' compares the training data (black), test data (yellow), and predicted data (blue), showing that the future water vapor trend is consistent with historical patterns, but generally shows a gradual upward trend, with the summer increase being particularly significant. Figure 11 Figure b depicts the interannual variation from 1959 to 2026, with the magenta shading representing historical data and the cyan shading representing projected data. The fitted curves indicate that despite local fluctuations over the past 60 years and the next two years, the overall trend of atmospheric water vapor continues to rise slowly. Figure 11 The value of 'c' indicates the nuclear density, and the range of atmospheric water vapor variation is still mainly concentrated around 22 kg / m³. 2 and 27kg / m 2 This prediction aligns with the anticipated increase in water vapor due to global warming, validating the model's reliability and providing a multi-scale forecasting basis for long-term planning in agricultural water resources management.
[0071] This embodiment provides an atmospheric water vapor content prediction system based on an AI large-scale model, including: a data acquisition module, a feature extraction module, a prediction module, a validation module, and an output module, supporting distributed deployment and real-time processing. The data acquisition module acquires data from ERA5, NCEP / NCAR, and IGRA via API interfaces, supporting multi-threaded downloading and fusion, and processing terabyte-level datasets. The feature extraction module implements DWT / CWT, using the Python PyWavelets library for parallel decomposition of grid point sequences. The prediction module runs WT-LSTM, trained using the TensorFlow framework, and supports GPU acceleration. The validation module calculates indicators and generates Taylor plots. The output module uses Matplotlib / GIS tools to visualize maps and reports. Deployed on a cloud platform, the system has a response time of less than minutes and is suitable for meteorological bureaus and agricultural departments.
[0072] This embodiment provides an atmospheric water vapor content prediction device based on an AI large-scale model, including: a processor (multi-core GPU), memory (>32GB RAM), a network interface, and a display (4K resolution). The processor executes the program to implement the method and supports parallel prediction; the memory stores the dataset and model parameters; the network interface transmits data in real time; and the display interactively shows the results. This device is portable, consumes less than 500W of power, is suitable for field stations, and supports mobile APP integration.
[0073] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An AI large model-based atmospheric water vapor content prediction method, characterized by, The method comprises: obtaining atmospheric water vapor content data and related meteorological variable data; using wavelet transform to perform time-frequency decomposition on the atmospheric water vapor content data and the related meteorological variable data, and extracting a multi-scale feature subsequence; constructing an AI large model based on a long short-term memory network, inputting the extracted multi-scale feature subsequence into the large model for multi-factor nonlinear training, and obtaining a trained large model; using the trained large model to perform spatio-temporal prediction on future atmospheric water vapor content, and obtaining a prediction result. 2.The AI big model-based atmospheric water vapor content prediction method of claim 1, wherein, The time-frequency decomposition of the atmospheric water vapor content data and the related meteorological variable data using wavelet transform comprises: using discrete wavelet transform as the decomposition method; setting the decomposition layer number to 3-4 layers; using Daubechies mother wavelet as the base function; generating high-frequency subsequences and low-frequency subsequences to capture short-term fluctuations and long-term trends of atmospheric water vapor, respectively. 3.The AI big model-based atmospheric water vapor content prediction method of claim 1, wherein, The construction of the AI large model based on the long short-term memory network comprises: constructing a network structure containing two hidden layers; setting the Adam optimizer and the mean square error loss function; adopting an early stopping mechanism to control the training process. 4.The AI big model-based atmospheric water vapor content prediction method of claim 1, wherein, The variables integrated by the multi-factor nonlinear training include: total column water vapor data; surface temperature data; precipitation data; and evaporation data. 5.The AI big model-based atmospheric water vapor content prediction method of claim 1, wherein, The method further comprises: extracting high water vapor events and low water vapor events using run theory; and quantifying the duration, frequency, intensity, and severity of the events. 6.The AI big model-based atmospheric water vapor content prediction method of claim 1, wherein, The output prediction result comprises: generating a spatio-temporal distribution map of atmospheric water vapor content; and displaying an interannual trend curve and a seasonal scale change.
7. An atmospheric water vapor content prediction system based on an AI large model, characterized by, A system for implementing the AI large model-based atmospheric water vapor content prediction method according to any one of claims 1-6, the system comprising: a data acquisition module for obtaining atmospheric water vapor content data and related meteorological variable data; a feature extraction module for using wavelet transform to perform time-frequency decomposition on the atmospheric water vapor content data and the related meteorological variable data, and extracting a multi-scale feature subsequence; a model construction and training module for constructing an AI large model based on a long short-term memory network, inputting the extracted multi-scale feature subsequence into the large model for multi-factor nonlinear training, and obtaining a trained large model; a prediction module for using the trained large model to perform spatio-temporal prediction on future atmospheric water vapor content, and obtaining a prediction result. 8.An atmospheric water vapor content prediction device based on an AI large model, characterized by, The system comprises: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement the AI large model-based atmospheric water vapor content prediction method according to any one of claims 1-6.