Conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclone

By using a conditional diffusion generative deep learning method, multi-source meteorological data of tropical cyclones can be reconstructed from infrared brightness temperature, which solves the problem of insufficient spatiotemporal resolution of low-orbit satellites and achieves high-quality reconstruction of multi-source meteorological data.

CN121009764APending Publication Date: 2025-11-25NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510843038.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing technologies, the spatiotemporal resolution of microwave observations by low Earth orbit satellites is insufficient, resulting in a scarcity of multi-source meteorological data for tropical cyclone reconstruction and making it difficult to achieve high spatiotemporal resolution observations.

Method used

A conditional diffusion generative deep learning method is used to reconstruct multi-source meteorological data from infrared radiometer data of geostationary satellites. By constructing a conditional diffusion model, data such as infrared brightness temperature, polarization-corrected temperature, and sea surface wind speed are generated, achieving high spatiotemporal resolution reconstruction.

Benefits of technology

It improved the peak signal-to-noise ratio, structural similarity, and learning-aware image patch similarity of the reconstruction results, thereby enhancing the reconstruction quality of multi-source meteorological data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009764A_ABST
    Figure CN121009764A_ABST
Patent Text Reader

Abstract

The invention discloses a conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclones. The conditional diffusion method comprises the following steps: acquiring an infrared radiation brightness set, a polarization correction temperature set and a sea surface wind speed of different tropical cyclones in a specified historical period; constructing a polarization correction temperature reconstruction model based on the conditional diffusion model; the input of the polarization correction temperature reconstruction model is an infrared radiation brightness temperature set, and the output target is a polarization correction temperature set; carrying out transfer learning on the trained polarization correction temperature reconstruction model to form a sea surface wind speed reconstruction model; the input of the sea surface wind speed reconstruction model is an infrared radiation brightness temperature set, and the output target is the sea surface wind speed. Therefore, the mapping relation from the infrared radiation brightness temperature to the multi-source meteorological data is established through generative deep learning of conditional diffusion, so that the multi-source meteorological data with the temporal-spatial resolution the same as that of geostationary orbit observation is reconstructed, and good performance is shown on various meteorological indexes and computer vision indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones, belonging to the field of meteorological data analysis and processing technology. Background Technology

[0002] Typical spectral bands for satellite observations can be divided into visible light (0.38 to 0.75 μm or 400 to 790 THz), infrared (IR, 0.75 μm to 1 mm or 300 GHz to 400 THz), and microwave (1 mm to 1 m or 0.3 to 300 GHz). Among them, IR can be used to observe cloud top information of tropical cyclones (TC) in all weather conditions (Dvorak, 1975; Velden et al., 2006; Olander and Velden, 2007), while microwaves can penetrate cloud tops to reveal the convective structure of TCs (Hawkins et al., 2001; Lee et al., 2002).

[0003] Microwaves can be used to estimate and predict turbulent current (TC) intensity (Jiang et al., 2019; Wimmers et al., 2019), and their advantage of observing convection through cloud tops can, along with IR, enhance the estimation capability of TC intensity (Wimmers and Velden, 2010, 2016; Olander and Velden, 2019). The three-dimensional structure of TCs can also be obtained from some temperature and humidity-sensitive channels of microwave radiometers (Boukabara et al., 2011). Furthermore, microwave data plays an important role in data assimilation, reducing initial analysis errors (Bauer et al., 2015) and providing effective guidance before TC landfall (Mcnally et al., 2014).

[0004] The lack of continuous microwave observation data has been a major challenge for TC (Temporal Tunneling) research, as TCs operate almost entirely over the ocean and are constantly moving throughout their lifespan. The development of microwave radiometers is limited by the size of the actual aperture antenna and the difficulty of obtaining high-sensitivity receivers, currently restricting their operation to low Earth orbit (LEO). Sensors operating in LEO are further constrained by their orbital altitude, allowing only one or two scans of the same location within 24 hours (Olander and Velden, 2019), a temporal resolution far lower than that of geostationary orbit (GEO) observations. To improve the temporal resolution of microwave observations, Staelin and Rosenkranz (1978) proposed the concept of a GEO microwave radiometer, which has been tested in various countries and organizations (Lambrigtsen et al., 2004; Christensen et al., 2007; Liu et al., 2011). In recent years, microsatellites have offered new avenues for improving the temporal resolution of microwave observations by increasing the number of LEO sensors (Blackwell et al., 2018; Roy et al., 2023). In addition, near real-time microwave data can be provided by algorithms that combine GEO IR data (Turk et al., 2000) and image morphology techniques (Wimmers and Velden, 2007).

[0005] Artificial intelligence (AI) models excel at extracting information and features from complex data and identifying patterns and regularities hidden within it, significantly improving weather and climate prediction capabilities (Ham et al., 2019; Chen et al., 2023; Bi et al., 2023), and demonstrating advantages in monitoring and forecasting total solar eclipse (TC) (Zhuo and Tan, 2021; Liu et al., 2024). AI models have also improved the accuracy and efficiency of synthetic microwave data (Haynes et al., 2024; Li et al., 2024). Meng et al. (2022) first attempted an image transformation task in TC research, using a generative adversarial network (GAN) to generate microwave-estimated precipitation images from IR images. However, the nature of the image transformation task limits the model's ability to obtain accurate quantitative precipitation, and the limited number of satellite channels and low horizontal resolution in the dataset also restricts the model's performance. Ortiz et al. (2024) used Bayesian deep learning to generate microwave brightness temperature and its uncertainty from GEO IR data, which helps to understand the source of uncertainty and guide improvements to the model.

[0006] Given the limited variety of real-time data and the potential for inaccurate reconstruction of understory convection structures using IR data (Hilburn et al., 2021), generative deep learning methods, including GANs and diffusion models, are increasingly being applied in areas such as atmospheric downscaling (Stengel et al., 2020; Ling et al., 2024; Mardani et al., 2025), transformations between meteorological variables (Xiao et al., 2024), nowcasting (Gong et al., 2024), and medium-term forecasting (Price et al., 2024). They are also suitable for reconstructing multi-source meteorological data from IR data. Diffusion models, in particular, can generate multiple ensemble members based on probability distributions and are more stable than GANs during training, making them suitable as the architecture for this invention. They generate target samples from noise, where the diffusion process is inspired by nonequilibrium thermodynamics (Sohl-Dickstein et al., 2015): a bidirectional Markov chain is defined, dividing the diffusion process into forward and reverse diffusion processes. The forward diffusion process gradually introduces Gaussian noise into the original input x0, generating x1…x T Then, the reverse diffusion process from p(x) t-1 |x t Sampling in the image x, iteratively removing noise. t Noise in the data. To generate new data samples, samples are drawn from a normal Gaussian distribution N(0,σ). 2 max Extracting Gaussian noise image x from I) T As model input, model ∈ θ Based on the current state x of each diffusion step t T Gradually predict and remove noise ∈ θ (x T The conditional diffusion model introduces conditional variables into the reverse diffusion process, allowing each step of the denoising process to utilize conditional information to generate more qualified samples.

[0007] The observation data from the synthetic aperture radar (SAR) operating in LEO is transformed through a geophysical model function to obtain the cross-polarized normalized radar cross-section backscattering, which can yield the TC sea surface wind speed (Li et al., 2013). The geophysical model function is obtained by matching SAR observations to effective wind speeds from different sources and then fitting the parameters of the SAR cross-polarized geophysical model function. GPROF (Goddard Profiling Algorithm) is an algorithm that combines Bayesian statistical methods to estimate surface precipitation and vertical precipitation structure from LEO passive microwave radiometer data, and it has important applications in atmospheric science research (Randel et al., 2020).

[0008] Transfer learning is a machine learning technique whose core idea is to use the parameters of a model already trained on one task as the initial parameters of a model for another related task, thereby accelerating the learning process of improving the new task, especially when data for the target task is scarce. In atmospheric science, transfer learning has been shown to play an important role in multiple fields, including TC (Li et al., 2019; Wang et al., 2023). Summary of the Invention

[0009] This invention addresses the scarcity of observational data caused by the insufficient spatiotemporal resolution of LEO observations by providing a conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclones (TCs). Using tropical cyclones, a common mesoscale hazardous weather system, as the experimental subject, this invention employs an artificial intelligence method to reconstruct multi-source meteorological data, including microwave polarization-corrected temperature and sea surface wind speed, from brightness temperature observations by geostationary satellite infrared radiometers. Compared to existing methods, this invention establishes a mapping relationship from infrared brightness temperature to the aforementioned multi-source meteorological data through generative deep learning in conditional diffusion, thereby reconstructing multi-source meteorological data with the same spatiotemporal resolution as geostationary orbit observations. Furthermore, it demonstrates excellent performance across various meteorological and computer vision metrics.

[0010] To achieve the above-mentioned technical objectives, the present invention will adopt the following technical solution:

[0011] A conditional diffusion method for reconstructing multi-source meteorological data from tropical cyclones includes the following steps:

[0012] Obtain the infrared brightness temperature set, polarization-corrected temperature set, and sea surface wind speed of different tropical cyclones within a specified historical period;

[0013] The obtained infrared radiation brightness temperature set and polarization correction temperature set are used to form the first dataset, and are allocated according to the preset division ratio to form the first training set, the first validation set and the first test set. At the same time, it is ensured that the infrared radiation brightness temperature set and polarization correction temperature set corresponding to the same tropical cyclone are in the same dataset.

[0014] The obtained infrared radiation brightness temperature set and sea surface wind speed are combined to form a second dataset, and then allocated according to a preset division ratio to form a second training set, a second validation set and a second test set. At the same time, it is ensured that the infrared radiation brightness temperature set and sea surface wind speed corresponding to the same tropical cyclone are in the same dataset during the division process.

[0015] A polarization-corrected temperature reconstruction model is constructed based on the conditional diffusion model, and is trained using the first training set, tested using the first test set, and validated using the first validation set. The input of the polarization-corrected temperature reconstruction model is the set of infrared radiation brightness temperatures, and the output target is the set of polarization-corrected temperatures.

[0016] The trained polarization correction temperature reconstruction model is transferred to a second training set, tested on a second test set, and validated on a second validation set to form a sea surface wind speed reconstruction model. The input of the sea surface wind speed reconstruction model is the infrared radiation brightness temperature set, and the output target is the sea surface wind speed.

[0017] Preferably, the obtained infrared radiation brightness temperature set, polarization correction temperature set, and sea surface wind speed are allocated according to a preset division ratio to form a training set, a validation set, and a test set, which is specifically achieved through the following steps:

[0018] Step 1.1, Data Filtering:

[0019] The data that meets the condition of the proportion of non-empty values ​​in the model input variables and model output variables is denoted as dataset df filter;

[0020] For the first dataset, the model input variable is the set of infrared radiation brightness temperatures corresponding to each tropical cyclone, and the model output variable is the set of polarization correction temperatures corresponding to each tropical cyclone.

[0021] For the second dataset, the model input variable is the set of infrared radiation brightness temperatures corresponding to each tropical cyclone; the model output variable is the sea surface wind speed corresponding to each tropical cyclone.

[0022] Step 1.2: Obtain a unique identifier:

[0023] Based on the dataset df_filter, a unique identifier is generated for each tropical cyclone according to the basic information of each tropical cyclone.

[0024] Assign the identifiers corresponding to each tropical cyclone to the dataset df_filter to generate a dataset containing the corresponding identifiers, denoted as the ID list;

[0025] Step 1.3: Shuffle the ID list:

[0026] The ID list is shuffled using a specified random number seed;

[0027] Step 1.4: Extract the identifiers that must be included in the test set:

[0028] Extract the identifiers corresponding to typical tropical cyclones from the shuffled ID list obtained in step 1.3 into the test set, and remove these identifiers from the shuffled ID list. The ID list after removal is denoted as the n-ID list.

[0029] Step 1.5: Calculate the number of samples contained in each identifier:

[0030] Based on the dataset df_filter and the n-ID list, count the number of samples contained in each identifier and output the data, which is denoted as the ID_counts table;

[0031] Step 1.6: Calculate the size of each dataset:

[0032] Based on the dataset df_filter, the number of samples in each dataset is calculated according to the preset partitioning ratio;

[0033] Step 1.7: Assign identifiers to each dataset:

[0034] Based on the n-ID list and ID_counts table, according to the calculated number of samples in each dataset and the distribution of hurricane wind force level of each sample in the n-ID list, the identifiers of each tropical cyclone are assigned to the corresponding datasets, and it is ensured that the total proportion of the three datasets of training set, validation set and test set and the distribution proportion of samples included in each hurricane wind force level in each dataset basically meet the preset division proportion.

[0035] Step 1.8: Construct the datasets:

[0036] Based on the identifiers assigned to each dataset obtained in step 1.7, samples with the corresponding identifiers are extracted from the dataset df_filter to construct the training set, validation set, and test set.

[0037] Preferably, step 1.7, when assigning identifiers to each dataset, specifically includes the following steps:

[0038] Step 1.7.1: Initialize the ID lists contained in the training and test sets;

[0039] Step 1.7.2: According to the hurricane wind force distribution, obtain the intensity of all samples from the dataset df_filter;

[0040] Step 1.7.3: Initialize counters for all hurricane wind force levels in each dataset;

[0041] Step 1.7.4: Traverse the n-ID list:

[0042] a) Obtain the cumulative number of samples for the current identifier and the corresponding hurricane wind force level;

[0043] b) Assign identifiers to each dataset based on the cumulative number of samples included in each identifier and the corresponding hurricane wind force level:

[0044] If the cumulative number of samples for the current identifier is less than the number of samples in the training set, or if there are no samples of the hurricane wind force level for the current identifier in the training set, then the current identifier will be assigned to the training set.

[0045] If the cumulative number of samples for the current identifier is less than the sum of the number of samples in the training set and the validation set, or if there are no samples of the hurricane wind force level for the current identifier in the validation set, then the current identifier will be assigned to the validation set.

[0046] Conversely, the current identifier is assigned to the test set;

[0047] c) Update the cumulative number of samples and hurricane wind force level counts for the training, test, and validation datasets.

[0048] Preferably, the hurricane wind force level is SSHWS level.

[0049] Preferably, in step 1.1, when filtering data, the proportion of non-null values ​​of the model input variables must be greater than or equal to 0.95; the proportion of non-null values ​​of the model output variables must be greater than or equal to 0.4.

[0050] Preferably, in step 1.2, the identifier generated for each tropical cyclone is constructed by extracting the ocean basin, number, and season of the corresponding tropical cyclone from the dataset df_filter.

[0051] Preferably, the obtained infrared radiation brightness temperature set, polarization correction temperature set, and sea surface wind speed are allocated according to a preset division ratio to form a training set, a validation set, and a test set, which is accomplished through a Python function.

[0052] Preferably, the polarization correction temperature set includes 37 channels of microwave polarization correction temperature (PCT). 37 and 89-channel microwave polarization correction temperature (PCT) 89 Specifically, it is calculated using the following formula:

[0053] PCT 37 =2.15T 37V -1.15T 37H

[0054] PCT 89 =1.7T 89V -0.7T 89H

[0055] Where: PCT 37 Indicates the microwave polarization correction temperature for the 37GHz channel; T37V T represents the microwave polarization vertical temperature of the 37GHz channel; 37H PCT represents the microwave polarization level temperature of the 37 GHz channel. 89 Indicates the microwave polarization correction temperature of the 89GHz channel; T 89V T represents the microwave polarization vertical temperature of the 89GHz channel; 89H This indicates the microwave polarization level temperature of the 89GHz channel.

[0056] Preferably, the infrared radiation brightness temperature set includes the brightness temperature BT of multiple channels. IR Elevation and mask; brightness temperature BT for multiple channels IR This includes the brightness temperatures of the first channel from the 6.15μm spectral band, the second channel from the 7.00μm spectral band, the third channel from the 7.40μm spectral band, the fourth channel from the 8.50μm spectral band, the fifth channel from the 9.70μm spectral band, the sixth channel from the 10.30μm spectral band, the seventh channel from the 11.20μm spectral band, the eighth channel from the 12.30μm spectral band, and the ninth channel from the 13.30μm spectral band.

[0057] Another technical objective of the present invention is to provide an electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, the computer program executing the above-described conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclones.

[0058] Based on the above-mentioned technical objectives, the present invention has the following advantages compared with the prior art:

[0059] The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones described in this invention, compared with the existing BT method of low earth orbit (LEO) radiometers, offers advantages over conventional methods. MW The observation method in this invention is provided by a geostationary orbit (GEO) satellite. IR This allows for the acquisition of information, achieving the same spatiotemporal resolution as GEO observations. This is in contrast to BT reconstructed using a Generative Adversarial Network (GAN) with a traditional convolutional neural network as the generator. MWIn comparison, the reconstruction results of the microwave brightness temperature reconstruction method using the conditional diffusion model have a higher peak signal-to-noise ratio (PSNR), higher structural similarity (SSIM), and lower learned perceptual image patch similarity (LPIPS). Attached Figure Description

[0060] Figure 1 This is a flowchart of the conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones as described in this invention.

[0061] Figure 2 This is a schematic diagram of the conditional diffusion model used in this invention.

[0062] Figure 3 This is an example of reconstructing the PCT using the present invention. In the figure: (a) BT of channel 13. IR (b)PCT 89 Objective, (c) Reconstruct PCT 89 (d) Ensemble average, and (e) Continuously ranked probability score (CRPS) calculated from ensemble members and targets. The performance diagram is also included. The reconstructed PCT is shown in Figure 4. 89 Set members. The set average and the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and image similarity index (LPIPS) of each set member relative to the target value are displayed in the lower left corner of the subplot.

[0063] Figure 4 This is an example of reconstructing PCT and its attention map according to the present invention. (a) BT IR (b)PCT 89 Objective, (c)PCT 89 The set average (d) is calculated using the set members and the target, (eh) is patched according to the model settings (ad), and (ip) is the attention map for each point, with points marked with red squares and their indices marked with red numbers. The brighter the color of the attention map, the greater the weight.

[0064] Figure 5 To test BT samples with and without binocular walls IR Differences. IR B08 IR B09 IR B10 IR B11 IR B12 IRB13 IR B14 IR B15 and IR B16 Corresponding to (a) to (i).

[0065] Figure 6 This is an example of using the present invention to reconstruct SAR sea surface wind speed. (a) BT IR (b) SAR sea surface wind speed target value, (c) reconstructed SAR sea surface wind speed ensemble mean, (d) reanalysis data 10-meter wind speed and instantaneous wind speed, (e) CRPS calculated from ensemble members, (fo) reconstructed SAR sea surface wind speed ensemble members. Detailed Implementation

[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Unless otherwise specifically stated, the relative arrangement, expressions, and values ​​of components and steps set forth in these embodiments do not limit the scope of the present invention. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0067] like Figure 1 , Figure 2As shown, the conditional diffusion method for reconstructing multi-source meteorological data according to this invention includes forward diffusion and backward diffusion, with a Transformer backbone network. Forward and backward diffusion are used during the training phase, while only backward diffusion is used after training is complete. First, the data required for training and inference needs to be prepared. Then, forward and backward diffusion are performed sequentially during training. Conditions are added to the backbone network at each time step of backward diffusion to guide the model's prediction of noise. After each complete backward diffusion process, the current model is evaluated to obtain the optimal model parameters for subsequent inference. The above process is used to train both the PCT reconstruction model (i.e., the polarization-corrected temperature reconstruction model) and the SAR sea surface wind reconstruction model (i.e., the sea surface wind speed reconstruction model). The difference is that the SAR sea surface wind reconstruction model is trained using transfer learning based on the PCT reconstruction model. Simultaneously, attention maps are used to perform interpretability analysis on the results after the PCT reconstruction model training is completed. The specific steps are as follows:

[0068] Step 1: Data Preparation

[0069] Obtain the infrared radiation brightness temperature set, polarization correction temperature set, and sea surface wind speed of different tropical cyclones within a specified historical period to prepare the dataset for training the PCT reconstruction model, SAR sea surface wind reconstruction model, and inference. The obtained data needs to be divided into training set, test set, and validation set according to a preset division ratio.

[0070] To train, test, and validate the PCT reconstruction model, this invention requires constructing a first dataset from the acquired infrared radiation brightness temperature set and polarization correction temperature set, and then allocating them according to a preset ratio to form a first training set, a first validation set, and a first test set. Simultaneously, it ensures that the infrared radiation brightness temperature set and polarization correction temperature set corresponding to the same tropical cyclone are all within the same dataset. To train, test, and validate the SAR sea surface wind reconstruction model, this invention requires constructing a second dataset from the acquired infrared radiation brightness temperature set and sea surface wind speed, and then allocating them according to a preset ratio to form a second training set, a second validation set, and a second test set. Simultaneously, it ensures that during the partitioning process, the infrared radiation brightness temperature set and sea surface wind speed corresponding to the same tropical cyclone are all within the same dataset.

[0071] Specifically, the data categories include BT data from multiple channels listed in Table 1. IR Vertical and horizontal polarized microwave (BT) data, elevation and masking, and SAR sea surface wind speed. Among these, BT... MW PCT is formed by linearly combining Equations (1) and (2) to reduce the impact of surface emissivity differences and better identify precipitation-related ice scattering signals (Cecil and Chronis, 2018).

[0072] PCT 37 =2.15T 37V -1.15T 37H (1)

[0073] PCT 89 =1.7T 89V -0.7T 89H (2)

[0074] Where: PCT 37 Indicates the microwave polarization correction temperature for the 37GHz channel; T 37V T represents the microwave polarization vertical temperature of the 37GHz channel; 37H PCT represents the microwave polarization level temperature of the 37 GHz channel. 89 Indicates the microwave polarization correction temperature of the 89GHz channel; T 89V T represents the microwave polarization vertical temperature of the 89GHz channel; 89H This indicates the microwave polarization level temperature of the 89GHz channel.

[0075] Table 1 Data Preparation Instructions

[0076]

[0077] When dividing the acquired data into training, testing, and validation sets, it is necessary to ensure that all samples from the same tropical cyclone appear in only one dataset to avoid data leakage. Traditional partitioning methods often divide by year, which is only suitable for situations where the number and intensity of TC samples are evenly distributed each year. The microwave data (polarization-corrected temperature set) and SAR data (infrared radiation brightness temperature set) used in this invention are polar-orbiting satellite observation data, which have the characteristics of discontinuous temporal and spatial coverage. Therefore, traditional methods cannot be used for dataset partitioning.

[0078] Therefore, this invention provides a novel dataset construction method, applied to the sample allocation process of the first and second datasets respectively, to ensure that the infrared radiation brightness temperature set and sea surface wind speed corresponding to the same tropical cyclone are all in the same dataset. Furthermore, since the method of allocating the first dataset according to a preset partitioning ratio to form the first training set, first validation set, and first test set is the same as the method of allocating the second dataset according to a preset partitioning ratio to form the second training set, second validation set, and second test set, for ease of explanation, the dataset construction method described in this invention does not use "first" or "second" before each dataset for distinction. It is specifically implemented through the following steps:

[0079] Step 1.1, Data Filtering:

[0080] The dataset df_filter is a selection of data from the model input and output variables whose proportion of non-empty values ​​meets certain criteria; where:

[0081] For the first dataset, the model input variable is the set of infrared radiation brightness temperatures corresponding to each tropical cyclone, and the model output variable is the set of polarization correction temperatures corresponding to each tropical cyclone. For the second dataset, the model input variable is the set of infrared radiation brightness temperatures corresponding to each tropical cyclone, and the model output variable is the sea surface wind speed corresponding to each tropical cyclone.

[0082] When filtering data, the proportion of non-null values ​​in the model input variables must be greater than or equal to 0.95; the proportion of non-null values ​​in the model output variables must be greater than or equal to 0.4.

[0083] Step 1.2: Obtain a unique identifier:

[0084] Based on the dataset df_filter, a unique identifier is generated for each tropical cyclone according to its basic information, so that the dataset can be divided according to the identifier in the future. Then, the identifier corresponding to each tropical cyclone is assigned to the dataset df_filter to generate a dataset containing the corresponding identifier, which is called the ID list.

[0085] Specifically, the identifier generated for each tropical cyclone is constructed by extracting the ocean basin, number, and season of the corresponding tropical cyclone from the dataset df_filter. For example, the identifier "WP012020" represents the 1 (01)TC (tropical cyclone) in the Northwest Pacific (WP) in 2020.

[0086] Step 1.3: Shuffle the ID list:

[0087] Shuffle the ID list using a specified random number seed.

[0088] Specifically, given a random number seed, the list of IDs is randomly shuffled using the built-in Python method numpy.random.shuffle to ensure the reproducibility and randomness of the subsequent allocation process.

[0089] Step 1.4: Extract the identifiers that must be included in the test set:

[0090] The identifiers corresponding to typical tropical cyclones are extracted from the shuffled ID list obtained in step 1.3 and added to the test set. These identifiers are then removed from the shuffled ID list, and the resulting ID list is denoted as the n-ID list.

[0091] Therefore, this invention improves the flexibility of the model by specifying that certain TCs must be included in the test set, which facilitates the evaluation of TCs of particular interest.

[0092] Step 1.5: Calculate the number of samples contained in each identifier:

[0093] Based on the dataset df_filter and the n-ID list, count the number of samples contained in each identifier and output the data, which is denoted as the ID_counts table;

[0094] Step 1.6: Calculate the size of each dataset:

[0095] Based on the dataset df_filter, the number of samples in each dataset is calculated according to the preset partitioning ratio.

[0096] Step 1.7: Assign identifiers to each dataset:

[0097] Based on the n-ID list and ID_counts table, according to the calculated sample count of each dataset and the hurricane wind force distribution of each sample in the n-ID list, the identifiers of each tropical cyclone are assigned to the corresponding datasets (training set, validation set, and test set), and it is ensured that the total proportion of the three datasets (training set, validation set, and test set) and the distribution proportion of samples included in each hurricane wind force in each dataset (training set, validation set, and test set) basically conform to the preset division proportion.

[0098] Therefore, it can be seen that when constructing each dataset, this invention takes into account both the overall partitioning ratio and the TC intensity level partitioning ratio (i.e., the distribution ratio of samples included in each hurricane wind force level in each dataset). It not only follows the partitioning ratio of the overall dataset, but also tries its best to maintain these ratios at different hurricane wind force levels (represented by the Saffir-Simpson Hurricane Wind Force Level, i.e., SSHWS intensity level), thus ensuring the balance between datasets to the greatest extent.

[0099] Specifically, the process of assigning identifiers to each dataset includes the following steps:

[0100] Step 1.7.1: Initialize the ID lists contained in the training and test sets;

[0101] Step 1.7.2: According to the hurricane wind force distribution, obtain the hurricane wind force intensity of all samples from the dataset df_filter;

[0102] Step 1.7.3: Initialize the counts for all hurricane wind force levels in each dataset;

[0103] Step 1.7.4: Traverse the n-ID list:

[0104] a) Obtain the cumulative number of samples for the current identifier and the corresponding hurricane wind force level;

[0105] b) Assign identifiers to each dataset based on the cumulative number of samples included in each identifier and the corresponding hurricane wind force level:

[0106] If the cumulative number of samples for the current identifier is less than the number of samples in the training set, or if there are no samples of the hurricane wind force level for the current identifier in the training set, then the current identifier will be assigned to the training set.

[0107] If the cumulative number of samples for the current identifier is less than the sum of the number of samples in the training set and the validation set, or if there are no samples of the hurricane wind force level for the current identifier in the validation set, then the current identifier will be assigned to the validation set.

[0108] Conversely, the current identifier is assigned to the test set;

[0109] c) Update the cumulative number of samples and hurricane wind force level counts for the training, test, and validation datasets.

[0110] After the identifiers are assigned, a statistical summary is generated to show the distribution of samples across the training set, validation set, test set, and all SSHWS intensity levels (see Table 2) to determine whether the total proportion of the three datasets and the proportion of samples at each intensity level conform to the partitioning ratio. The specific steps include the following:

[0111] Based on the identifiers assigned to each dataset obtained in step 1.7.4, the corresponding samples are filtered out from df_filter and stored in the corresponding datasets, thereby obtaining the training set, validation set and test set;

[0112] Each dataset is grouped according to the SSHWS intensity level, and the number of samples in each dataset that fall under the corresponding SSHWS intensity level is calculated.

[0113] The results were combined into a single table (see Table 2), with row headings for SSHWS strength levels (including “TD”, “TS”, “C1”, “C2”, “C3”, “C4”, “C5” and “Total”) and column headings for datasets (including “Training Set”, “Validation Set”, “Test Set” and “Total”).

[0114] Table 2. Partitioning of the First Dataset

[0115]

[0116] In Table 2, the numbers in parentheses represent the ratio (integers, unit: %) of the number of samples in each dataset to the total number of samples at the corresponding SSHWS intensity level. For example, when the SSHWS intensity level is TD, the total number of samples is 4659, and 3180 samples are allocated to the training set. Therefore, the ratio of samples with the SSHWS intensity level of TD allocated to the training set is 3180 / 4659*100%≈68%, which matches the ratio of the number of samples in the training set (9253) to the total number of samples (13213).

[0117] Step 1.8: Construct the datasets:

[0118] Based on the identifiers assigned to each dataset obtained in step 1.7, samples with the corresponding identifiers are extracted from the dataset df_filter to construct the training set, validation set, and test set.

[0119] Step 1.9: Calculate the range of variables: Calculate the minimum and maximum values ​​of the input and output variables based on the training set.

[0120] Calculate the minimum and maximum values ​​of each variable in the model input variables, model output variables, and additional variables (such as elevation);

[0121] The minimum and maximum values ​​are stored in the dictionaries dict_min and dict_max, respectively.

[0122] Step 2: Train the PCT reconstruction model:

[0123] A PCT reconstruction model was constructed based on the conditional diffusion model, and was trained using the first training set, tested using the first test set, and validated using the first validation set.

[0124] The input to the PCT reconstruction model is the brightness temperature (BT) of multiple channels observed by an infrared radiometer mounted on a geostationary satellite. IR Elevation and mask; the output of the PCT reconstruction model is a 37-channel microwave polarization-corrected temperature (PCT) model. 37 and 89-channel microwave polarization correction temperature (PCT) 89 ( Figure 2 ).

[0125] For a general diffusion model, firstly, a bidirectional Markov chain of length T is defined, consisting of forward diffusion and backward diffusion processes. The forward diffusion process gradually adds Gaussian noise to the original input x0, generating x1…x T The reverse diffusion process starts from probability p(x) t-1 |x t Sampling, iteratively removing noise from image x t Noise in the model. To generate new data samples, the model is derived from a normal Gaussian distribution. Extract Gaussian noise map x t As input, noise is progressively predicted and removed at each diffusion step t to complete the reverse diffusion process. Compared to general diffusion models, the conditional diffusion model used in this invention introduces the brightness temperature BT of multiple channels in the reverse diffusion. IR Condition variables Gaussian noise map x t Both are used as inputs to the backbone network DiT, enabling each step of the denoising process to utilize conditional information from the joint probability. Sampling removes noise, guiding the model to generate a brightness temperature (BT) that more closely approximates the real microwave brightness temperature. MW The sample.

[0126] Referring to Table 1, the brightness temperature (BT) of multiple input channels is used when training the PCT reconstruction model. IR This includes the brightness temperatures of the first channel from the 6.15μm spectral band, the second channel from the 7.00μm spectral band, the third channel from the 7.40μm spectral band, the fourth channel from the 8.50μm spectral band, the fifth channel from the 9.70μm spectral band, the sixth channel from the 10.30μm spectral band, the seventh channel from the 11.20μm spectral band, the eighth channel from the 12.30μm spectral band, and the ninth channel from the 13.30μm spectral band.

[0127] After constructing the PCT reconstruction model, this invention uses attention maps to perform interpretability analysis on the output results of the PCT reconstruction model.

[0128] Figure 3 The case shown has a clear TC double-wall structure (the central primary eyewall and the outer secondary eyewalls are separated by a circular moat area with little or no precipitation, Houze, (2010)). We selected eight locations and visualized the attention maps for these locations [see...]. Figure 4 [From (i) to (p)]. Three locations in the moat region [ Figure 4 The attention map weights from (i) to (k) are all concentrated in BT. IR The lower-lying areas indicate that information for reconstructing the moat region cannot be directly obtained from these points; that is, the model lacks the ability to penetrate the infrared cloud top, which aligns with physical facts. Therefore, the information for reconstructing the moat region is learned from other areas. Locations 4 to 6 are distributed along the strong convection, weak convection, and outer rainband of the secondary eyewall. Figure 4 [(l) to (n)]. The attention map weights at these locations are also concentrated in BT. IRLower-weighted areas, but they also assign lower weights to the eye and moat areas (visually appearing darker). This feature is more pronounced in the attention maps of thin clouds and clear skies. Figure 4 [o and p]. Location 7, located in the thin cloud region, assigns weights to cloudy areas while further emphasizing the lower weights of the eye, moat, and clear sky. The attention map at location 8 is almost the opposite of the attention map at location 4, located in the secondary eyewall. The attention map at location 8 assigns higher weights to clear sky and thin cloud regions similar to those at location 8, which is consistent with the self-attention mechanism (Zhang et al., 2019). Therefore, the model learns the mapping relationship between thin clouds and clear sky from IR to PCT, as well as the mapping relationship from IR cloud tops to PCT deep convection. The information reconstructed from the moat region is indirectly learned through these two relationships.

[0129] To further investigate whether the model truly learned the features of the double eyelid wall, we compared the BT (Bit-Based Transformation) results of samples where the model successfully reconstructed the double eyelid wall with those where it successfully reconstructed samples without the double eyelid wall. IR Differences ( Figure 5 The results showed that, compared with samples without binocular walls, all BT samples with binocular walls... IR The channels are all relatively low. This difference is particularly pronounced in the outer core, indicating that samples with double eye walls exhibit a wider outer rainband. The large outer core size or organized outer rainband has been shown to favor the formation and maintenance of double eye walls in ideal simulations (Wang and Tan, 2020). Therefore, the PCT reconstruction model trained in this invention has learned the key dynamic characteristics of double eye walls.

[0130] Step 3: Construct and train a SAR sea surface wind speed reconstruction model.

[0131] Based on the trained PCT reconstruction model, transfer learning was performed, and a second training set was used for training, a second test set for testing, and a second validation set for validation to construct a SAR sea surface wind speed reconstruction model. The SAR sea surface wind speed reconstruction model uses brightness temperature (BT) as its benchmark. IR Using the mask as input, the target is reconstructed as the SAR sea surface wind field. Figure 6 ).

[0132] The SAR sea surface wind speed reconstruction model described in this invention is constructed using a PCT reconstruction model trained through transfer learning because SAR-retrieved sea surface wind speed data is extremely scarce (Table 3). Furthermore, given the limited sea surface wind speed data prepared when constructing the SAR sea surface wind speed reconstruction model, this invention allocates the second dataset only to the second training set and the second test set when dividing the second dataset.

[0133] Table 3 shows the partitioning of the second dataset.

[0134]

[0135] The IR data used in this invention comes from GEO radiometers, while microwave and sea surface wind speed data come from LEO radiometers. Furthermore, sea surface wind speed can be obtained by inversion from microwave data (Stogryn et al., 1994; Wang et al., 2017), indicating a correlation between the two. Therefore, this invention is well-suited for use with transfer learning, i.e., first training a model to convert IR data to microwave data, and then fine-tuning this model to obtain a model to convert IR data to sea surface wind speed data.

[0136] Example 2

[0137] The present invention also provides a storage medium, wherein the program stored in the storage medium executes the above-described conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclones when running.

[0138] Example 3

[0139] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclones through the computer program.

[0140] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0141] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0144] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0145] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0146] This invention provides a conditional diffusion method for reconstructing multi-source meteorological data from tropical cyclones. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the principles of this invention, and these modifications and improvements should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A conditional diffusion method for reconstructing multi-source meteorological data from tropical cyclones, characterized in that, Includes the following steps: Obtain the infrared brightness temperature set, polarization-corrected temperature set, and sea surface wind speed of different tropical cyclones within a specified historical period; The obtained infrared radiation brightness temperature set and polarization correction temperature set are used to form the first dataset, and are allocated according to the preset division ratio to form the first training set, the first validation set and the first test set. At the same time, it is ensured that the infrared radiation brightness temperature set and polarization correction temperature set corresponding to the same tropical cyclone are in the same dataset. The obtained infrared radiation brightness temperature set and sea surface wind speed are combined to form a second dataset, and then allocated according to a preset division ratio to form a second training set, a second validation set and a second test set. At the same time, it is ensured that the infrared radiation brightness temperature set and sea surface wind speed corresponding to the same tropical cyclone are in the same dataset during the division process. A polarization-corrected temperature reconstruction model is constructed based on the conditional diffusion model, and is trained using the first training set, tested using the first test set, and validated using the first validation set. The input of the polarization-corrected temperature reconstruction model is the set of infrared radiation brightness temperatures, and the output target is the set of polarization-corrected temperatures. The trained polarization correction temperature reconstruction model is transferred to a second training set, tested on a second test set, and validated on a second validation set to form a sea surface wind speed reconstruction model. The input of the sea surface wind speed reconstruction model is the infrared radiation brightness temperature set, and the output target is the sea surface wind speed.

2. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 1, characterized in that, The method of dividing the first dataset into a first training set, a first validation set, and a first test set according to a preset partitioning ratio is the same as the method of dividing the second dataset into a second training set, a second validation set, and a second test set according to a preset partitioning ratio. This can be achieved through the following steps: Step 1.1, Data Filtering: The dataset df_filter is the set of data from the model input variables and model output variables that meets the condition of having a certain proportion of non-empty values. For the first dataset, the model input variable is the set of infrared radiation brightness temperatures corresponding to each tropical cyclone, and the model output variable is the set of polarization correction temperatures corresponding to each tropical cyclone. For the second dataset, the model input variable is the set of infrared radiation brightness temperatures corresponding to each tropical cyclone; The model output variable is the sea surface wind speed corresponding to each tropical cyclone; Step 1.2: Obtain a unique identifier: Based on the dataset df_filter, a unique identifier is generated for each tropical cyclone according to its basic information. The identifiers corresponding to each tropical cyclone are then assigned to the dataset df_filter to generate a dataset containing the corresponding identifiers, denoted as the ID list. Step 1.3: Shuffle the ID list: The ID list is shuffled using a specified random number seed; Step 1.4: Extract the identifiers that must be included in the test set: Extract the identifiers corresponding to typical tropical cyclones from the shuffled ID list obtained in step 1.3 into the test set, and remove these identifiers from the shuffled ID list. The ID list after removal is denoted as the n-ID list. Step 1.5: Calculate the number of samples contained in each identifier: Based on the dataset df_filter and the n-ID list, count the number of samples contained in each identifier and output the data, which is denoted as the ID_counts table; Step 1.6: Calculate the size of each dataset: Based on the dataset df_filter, the number of samples in each dataset is calculated according to the preset partitioning ratio; Step 1.7: Assign identifiers to each dataset: Based on the n-ID list and ID_counts table, according to the calculated number of samples in each dataset and the distribution of hurricane wind force level of each sample in the n-ID list, the identifiers of each tropical cyclone are assigned to the corresponding datasets, and it is ensured that the total proportion of the three datasets of training set, validation set and test set and the distribution proportion of samples included in each hurricane wind force level in each dataset basically meet the preset division proportion. Step 1.8: Construct the datasets: Based on the identifiers assigned to each dataset obtained in step 1.7, samples with the corresponding identifiers are extracted from the dataset df_filter to construct the training set, validation set, and test set.

3. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 2, characterized in that, Step 1.7, assigning identifiers to each dataset, specifically includes the following steps: Step 1.7.1: Initialize the ID lists contained in the training and test sets; Step 1.7.2: According to the hurricane wind force distribution, obtain the hurricane wind force intensity of all samples from the dataset df_filter; Step 1.7.3: Initialize counters for all hurricane wind force levels in each dataset; Step 1.7.4: Traverse the n-ID list: a) Obtain the cumulative number of samples for the current identifier and the corresponding hurricane wind force level; b) Assign identifiers to each dataset based on the cumulative number of samples included in each identifier and the corresponding hurricane wind force level: If the cumulative number of samples for the current identifier is less than the number of samples in the training set, or if there are no samples of the hurricane wind force level for the current identifier in the training set, then the current identifier will be assigned to the training set. If the cumulative number of samples for the current identifier is less than the sum of the number of samples in the training set and the validation set, or if there are no samples of the hurricane wind force level for the current identifier in the validation set, then the current identifier will be assigned to the validation set. Conversely, the current identifier is assigned to the test set; c) Update the cumulative number of samples and hurricane wind force level counts for the training, test, and validation datasets.

4. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 2, characterized in that, The hurricane wind force level is SSHWS.

5. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 2, characterized in that, In step 1.1, when filtering data, the proportion of non-null values ​​for the model input variables must be greater than or equal to 0.95; the proportion of non-null values ​​for the model output variables must be greater than or equal to 0.

4.

6. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 2, characterized in that, In step 1.2, the identifier generated for each tropical cyclone is constructed by extracting the ocean basin, number, and season of the corresponding tropical cyclone from the dataset df_filter.

7. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 2, characterized in that, The obtained infrared radiation brightness temperature set, polarization correction temperature set, and sea surface wind speed are allocated according to a preset division ratio to form a training set, a validation set, and a test set, which is accomplished through Python functions.

8. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 1, characterized in that, The polarization correction temperature set includes 37 channels of microwave polarization correction temperature (PCT). 37 and 89-channel microwave polarization correction temperature PCT 89 Specifically, it is calculated using the following formula: PCT 37 =2.15T 37V -1.15T 37H PCT 89 =1.7T 89V -0.7T 89H Where: PCT 37 Indicates the microwave polarization correction temperature for the 37GHz channel; T 37V T represents the microwave polarization vertical temperature of the 37GHz channel; 37H PCT represents the microwave polarization level temperature of the 37 GHz channel. 89 Indicates the microwave polarization correction temperature of the 89GHz channel; T 89V T represents the vertical temperature of microwave polarization in the 89GHz channel. 89H This indicates the microwave polarization level temperature of the 89GHz channel.

9. The conditional diffusion method for reconstructing multi-source meteorological data of tropical cyclones according to claim 1, characterized in that, Infrared radiation brightness temperature collection includes brightness temperature (BT) values ​​for multiple channels. IR Elevation and mask; brightness temperature BT for multiple channels IR This includes the brightness temperatures of the first channel from the 6.15μm spectral band, the second channel from the 7.00μm spectral band, the third channel from the 7.40μm spectral band, the fourth channel from the 8.50μm spectral band, the fifth channel from the 9.70μm spectral band, the sixth channel from the 10.30μm spectral band, the seventh channel from the 11.20μm spectral band, the eighth channel from the 12.30μm spectral band, and the ninth channel from the 13.30μm spectral band.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The computer program executes the conditional diffusion method for reconstructing multi-source meteorological data for tropical cyclones as described in any one of claims 1 to 9.