Spatial Interpolation Method, Device, Electronic Device and Storage Medium
By predicting the climate data of unknown sites in the target observation area and determining the weight based on spatial autocorrelation, the problem of low accuracy in sparse sites is solved, and higher interpolation accuracy of climate data is achieved.
Patent Information
- Application Number
- CN202410082549.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-01-19
AI Technical Summary
The traditional spatial interpolation method has low accuracy and large errors in sparse sites, and depends on the original remote sensing products to cause large errors between the climate observation product data and the site observation data.
By obtaining the original climate data of known and unknown sites in the target observation area, the crude predicted climate data of unknown sites are predicted, and the prediction weight is determined based on the spatial autocorrelation between unknown sites and known sites, and the accuracy of climate data is finally improved through interpolation method.
The accuracy of traditional spatial interpolation methods in sparse sites is improved, the error of interpolation results is reduced, and the accuracy of climate observation products is improved.
Smart Images

Figure CN118096509B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer processing, and in particular, to a spatial interpolation method, device, electronic device and storage medium. Background Art
[0002] At present, the methods for generating spatially gridded climate observation products usually include interpolation generation, remote sensing products and reanalysis data.
[0003] Among them, for interpolation, there are mainly machine learning methods and traditional interpolation methods. Machine learning methods rely to a large extent on surrounding variables. If spatial autocorrelation is not considered and the correlation with surrounding stations is not large, the accuracy of the interpolation results will be low. Traditional interpolation methods have greater uncertainty for interpolation in areas with sparse stations. Remote sensing products will have large errors due to environmental factors.
[0004] In addition, there is also the reanalysis and data fusion of remote sensing data based on deep learning. Since it mainly depends on the original remote sensing products, there is still a large error between the data of the climate observation products and the actual data obtained from station observations. Summary of the Invention
[0005] In view of this, the purpose of the embodiments of the present application is to provide a spatial interpolation method, device, electronic device and storage medium, which can improve the problems of low accuracy and large error of the interpolation results for geographical locations with sparse stations in traditional spatial interpolation methods.
[0006] To achieve the above technical purpose, the technical solution adopted by the present application is as follows:
[0007] In a first aspect, an embodiment of the present application provides a spatial interpolation method, and the method includes:
[0008] Obtain a target observation area, where the target observation area includes a plurality of known stations, a plurality of unknown stations, and the original climate data corresponding to each known station;
[0009] Predict the rough predicted climate data of each unknown station according to the original climate data;
[0010] Determine the prediction weight of each unknown station according to the spatial autocorrelation between each unknown station and the known stations that meet the preset conditions in space;
[0011] Determine the predicted climate data of each unknown station according to the rough predicted climate data and the prediction weight corresponding to each unknown station;
[0012] Perform climate data interpolation on the climate data to be predicted according to the predicted climate data to obtain an interpolation result.
[0013] In combination with the first aspect, in some alternative embodiments, determining the predicted climate data of each unknown site is achieved through a spatial interpolation model, wherein the training process of the spatial interpolation model is as follows:
[0014] Obtain a target training set, where the target training set includes a plurality of training samples, each training sample including known sites and unknown sites in a target observation area, and initial data corresponding to the known sites;
[0015] Perform gated convolution operations on each training sample in the training set through a gated convolution layer in a preset initial model to predict the rough prediction results of the unknown sites in each training sample;
[0016] Obtain weight factors corresponding to each known site and each unknown site;
[0017] According to the initial data and the weight factors corresponding to each known site, determine the spatial weights corresponding to each unknown site through an attention module in the preset initial model, where the spatial weights represent the spatial autocorrelation between the current unknown site and each known site;
[0018] Take the product of each rough prediction result and the spatial weight corresponding to each rough prediction result as the prediction result corresponding to each unknown site, and determine the model loss through a loss function according to each prediction result and the corresponding preset real data;
[0019] Update the parameters of the preset initial model according to the model loss until the model loss meets the preset model convergence condition to obtain a spatial interpolation model.
[0020] In combination with the first aspect, in some alternative embodiments, obtaining a target training set includes:
[0021] Obtain climate samples of the target observation area at multiple different times, where the climate samples include known sites, unknown sites in the target observation area, and precipitation values corresponding to the known sites;
[0022] Perform masking on the climate samples to obtain masked data corresponding to the climate samples;
[0023] Generate multiple first feature maps with different scales corresponding to the masked data according to the masked data;
[0024] Perform a weighted pooling operation on each first feature map to extract features from the precipitation values of the known sites in each first feature map to obtain a second feature map corresponding to each first feature map as the training sample, and use the precipitation value after feature extraction as the initial data.
[0025] Combined with the first aspect, in some alternative embodiments, the formula for the gated convolution operation is as follows:
[0026]
[0027]
[0028]
[0029] In the formula, represents the j-th rough prediction result obtained after the gated convolution operation, represents the initial data in the convolution region of the i-th convolution kernel of the (l - 1)-th layer; represents the weight of the i-th convolution kernel of the l-th layer; represents the conventional convolution operation; b i l represents the bias of the i-th convolution kernel of the l-th layer; M represents the convolution region, and f(·) represents the activation function.
[0030] Combined with the first aspect, in some alternative embodiments, the initial data includes precipitation values representing the precipitation of the known sites within a unit time, and the weight factors include site coordinates, elevation, and terrain undulation;
[0031] Then, according to the initial data and the weight factors corresponding to each known site, the spatial weight corresponding to each unknown site is determined through the attention module in the preset initial model, including:
[0032] Determine the Euclidean distance between the unknown site and each target known site through the attention module in the preset initial model:
[0033]
[0034] In the formula, distance i represents the Euclidean distance, represents the coordinate difference between the current unknown site and the target known site, represents the elevation difference between the current unknown site and the target known site, represents the difference in terrain undulation between the current unknown site and the target known site, represents the difference between the initial predicted value corresponding to the current unknown site and the precipitation value corresponding to the target known site, and the target known site represents a known site within a preset distance range from the current unknown site;
[0035] Normalize the Euclidean distance to determine the prediction weight:
[0036] weight i = softmin(distance i , {distance0, distance1, …, distance n )
[0037] where weight i represents the prediction weight, and {distance0, distance1, …, distance n} represents the set composed of each Euclidean distance.
[0038] Combined with the first aspect, in some alternative embodiments, the loss function is as follows:
[0039]
[0040] where n represents the number of samples, h(x k ) represents the prediction result of the kth training sample, y k represents the preset true data corresponding to the kth training sample, and I i,j represents the prediction result corresponding to the unknown site at (i, j).
[0041] In a second aspect, an embodiment of the present application further provides a spatial interpolation device, and the device includes:
[0042] A first acquisition unit, configured to acquire a target observation area, where the target observation area includes a plurality of known sites, a plurality of unknown sites, and the original climate data corresponding to each known site;
[0043] A prediction unit, configured to predict the rough prediction climate data of each unknown site according to the original climate data;
[0044] A first determination unit, configured to determine the prediction weight of each unknown site according to the spatial autocorrelation in space between each unknown site and the known sites that meet the preset conditions;
[0045] A second determination unit, configured to determine the prediction climate data of each unknown site according to the rough prediction climate data and the prediction weight corresponding to each unknown site;
[0046] An interpolation unit, configured to perform climate data interpolation on the to-be-predicted climate data according to the prediction climate data to obtain an interpolation result.
[0047] Combined with the second aspect, in some alternative embodiments, the device further includes:
[0048] A second acquisition unit, configured to acquire a target training set, where the target training set includes a plurality of training samples, and each training sample includes known sites and unknown sites in a target observation area, and initial data corresponding to the known sites;
[0049] A convolution unit, configured to perform gated convolution operations on each training sample in the training set through a gated convolution layer in a preset initial model, so as to predict a rough prediction result of the unknown site in each training sample;
[0050] A third acquisition unit, configured to acquire weight factors corresponding to each known site and each unknown site;
[0051] A third determination unit, configured to determine a spatial weight corresponding to each unknown site through an attention module in the preset initial model according to the initial data and the weight factors corresponding to each known site, where the spatial weight represents the spatial autocorrelation between the current unknown site and each known site;
[0052] A fourth determination unit, configured to use the product of each rough prediction result and the spatial weight corresponding to each rough prediction result as the prediction result corresponding to each unknown site, and determine a model loss through a loss function according to each prediction result and corresponding preset true data;
[0053] A training unit, configured to update parameters of the preset initial model according to the model loss until the model loss meets a preset model convergence condition, so as to obtain a spatial interpolation model.
[0054] In a third aspect, an embodiment of the present application further provides an electronic device, where the electronic device includes a processor and a memory that are coupled to each other, and a computer program is stored in the memory. When the computer program is executed by the processor, the electronic device is enabled to execute the above method.
[0055] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is enabled to execute the above method.
[0056] The invention adopting the above technical solution has the following advantages:
[0057] In the technical solution provided by this application, first, a target observation area is obtained. The target observation area includes multiple known stations, multiple unknown stations, and the original climate data corresponding to each known station. Then, the rough predicted climate data of each unknown station is predicted based on the original climate data, and the prediction weight of each unknown station is determined according to the spatial autocorrelation between each unknown station and all known stations in space. Then, the predicted climate data of each unknown station is determined according to the rough predicted climate data and the prediction weight corresponding to each unknown station. Finally, climate data interpolation is performed on the climate data to be predicted based on the predicted climate data to obtain an interpolation result. In this way, the problem that the interpolation result of the traditional spatial interpolation method has low accuracy and large error for geographical locations with sparse stations can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] This application can be further illustrated by the non-limiting embodiments given in the drawings. It should be understood that the following drawings only show some embodiments of this application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.
[0059] Figure 1 It is a block diagram of an electronic device provided by an embodiment of this application.
[0060] Figure 2 It is a schematic flowchart of a spatial interpolation method provided by an embodiment of this application.
[0061] Figure 3 It is a schematic diagram of the training process of a spatial interpolation model provided by an embodiment of this application.
[0062] Figure 4 It is an example diagram of weight pooling provided by an embodiment of this application.
[0063] Figure 5 It is an example diagram of the process of gated convolution operation provided by an embodiment of this application.
[0064] Figure 6 It is a schematic diagram of the determination process of spatial weights provided by an embodiment of this application.
[0065] Figure 7 It is a schematic flowchart of performing spatial interpolation on training samples provided by an embodiment of this application.
[0066] Figure 8 It is a block diagram of a spatial interpolation device provided by an embodiment of this application.
[0067] Icons: 100 - Electronic device; 101 - Processor; 102 - Memory; 300 - Spatial interpolation device; 310 - First acquisition unit; 320 - Prediction unit; 330 - First determination unit; 340 - Second determination unit; 350 - Interpolation unit. Detailed implementation manners
[0068] The present application will be described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that in the description of the drawings or the specification, similar or identical parts are all denoted by the same reference numerals, and the implementation manners not shown or described in the drawings are forms known to those of ordinary skill in the art. In the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0069] Please refer to Figure 1 , an electronic device 100 provided in an embodiment of the present application may include a processor 101 and a memory 102. A computer program is stored in the memory 102. When the computer program is executed by the processor 101, the electronic device 100 can execute the corresponding steps in the following spatial interpolation method.
[0070] In this embodiment, the electronic device 100 may be a personal computer, a laptop computer, a cloud server, etc., which is used to obtain a target observation area, then predict the rough predicted climate data of each unknown site according to the original climate data carried by the target observation area, and determine the prediction weight of each unknown site according to the spatial autocorrelation of each unknown site and all known sites in space, and then determine the predicted climate data of each unknown site according to the rough predicted climate data and the prediction weight corresponding to each unknown site, and finally perform climate data interpolation on the climate data to be predicted according to the predicted climate data to obtain an interpolation result.
[0071] Please refer to Figure 2 , the present application also provides a spatial interpolation method, which can be applied to the above-mentioned electronic device 100, and each step in the method is executed or implemented by the electronic device 100. Among them, the spatial interpolation method may include the following steps:
[0072] Step 110, obtain a target observation area, where the target observation area includes a plurality of known sites, a plurality of unknown sites, and the original climate data corresponding to each known site;
[0073] Step 120, predict the rough predicted climate data of each unknown site according to the original climate data;
[0074] Step 130, determine the prediction weight of each unknown site according to the spatial autocorrelation of each unknown site and the known sites that meet the preset conditions in space;
[0075] Step 140: Determine the predicted climate data of each unknown site according to the rough predicted climate data and the prediction weights corresponding to each unknown site.
[0076] Step 150: Perform climate data interpolation on the to-be-predicted climate data according to the predicted climate data to obtain an interpolation result.
[0077] In the above implementation, first, a target observation area is obtained. The target observation area includes multiple known sites, multiple unknown sites, and the original climate data corresponding to each known site. Then, the rough predicted climate data of each unknown site is predicted based on the original climate data, and the prediction weights of each unknown site are determined according to the spatial autocorrelation between each unknown site and all known sites in space. Then, the predicted climate data of each unknown site is determined according to the rough predicted climate data and the prediction weights corresponding to each unknown site. Finally, climate data interpolation is performed on the to-be-predicted climate data according to the predicted climate data to obtain an interpolation result. In this way, the problem that the interpolation result has low accuracy and large error in the traditional spatial interpolation method for geographical locations with sparse sites can be improved.
[0078] The following will elaborate on each step of the spatial interpolation method in detail as follows:
[0079] In step 110, in order to solve the problem of inaccurate climate observation in areas with sparse sites in this embodiment, the target observation area can be any geographical range with relatively few climate observation sites. In this embodiment, the Qinghai-Tibet Plateau region is taken as an example.
[0080] It can be understood that the technical solution takes climate observation sites as the interpolation objects to supplement and improve climate data in areas with sparse site distributions. Therefore, for the convenience of implementing the technical solution, the target observation area is rasterized. Among them, a grid with a climate observation site and whose climate data at the location of the climate observation site can be directly observed through the climate observation site is used as a known site, and each known site carries the original climate data obtained by observing the location of the site; the grid without a climate observation site is used as an unknown site.
[0081] In this embodiment, the acquisition of the target observation area can be data obtained in real time based on a preset rule. Taking the Qinghai-Tibet Plateau region as an example, the location distributions of known and unknown stations in this region can be obtained at 23:59 on the last day of each month, and the monthly climate data (the climate data can be temperature, humidity, precipitation value, etc.) observed at each known station in this month in this region can be used as the original climate data; alternatively, the target observation area can be storing the above monthly climate data in the memory 102 of the electronic device 100 as historical data, and based on the user's operation instruction, calling the data that meets the user's needs from the historical data, such as the location distributions of known and unknown stations in the Qinghai-Tibet Plateau region in the previous month, and the monthly average temperature observed at each known station. Here, the acquisition method of the target observation area is not specifically limited.
[0082] In step 120, through gated convolution operation, data prediction is performed on each unknown station according to the original climate data of each known station, and the rough predicted climate data corresponding to each unknown station is obtained.
[0083] Specifically, for the operation process of gated convolution, refer to the description of gated convolution operation in the training of the following spatial interpolation model, and details are not elaborated here.
[0084] In step 130, the spatial autocorrelation (i.e., coordinate difference, elevation difference, terrain undulation difference, data difference, etc.) between each unknown station and the known stations that meet the preset conditions (i.e., the known stations within a preset range of distance from the unknown station, which usually refers to all known stations in practical applications) is used as the prediction weight for correcting the rough predicted climate data.
[0085] In this way, the data prediction of the unknown stations can be made more accurate. Specifically, for the determination method of the prediction weight, refer to the description of the spatial weight in the training of the following spatial interpolation model, and details are not elaborated here.
[0086] In step 140, the product of the rough predicted climate data and the prediction weight is used as the predicted climate data of each unknown station.
[0087] In step 150, the predicted climate data predicted for each unknown station is used as the interpolation result of this unknown station, and the climate data to be predicted including known and unknown stations is interpolated to obtain the interpolation result.
[0088] Please refer to Figure 3 , in this embodiment, steps 120 to 140 are implemented through a spatial interpolation model, and the training process of the spatial interpolation model is as follows:
[0089] Step 210: Obtain a target training set, where the target training set includes multiple training samples, and each training sample includes known sites and unknown sites in a target observation area, as well as initial data corresponding to the known sites;
[0090] Step 220: Perform gated convolution operations on each training sample in the training set through a gated convolution layer in a preset initial model to predict a rough prediction result of the unknown sites in each training sample;
[0091] Step 230: Obtain weight factors corresponding to each known site and each unknown site;
[0092] Step 240: According to the initial data and the weight factors corresponding to each known site, determine the spatial weight corresponding to each unknown site through an attention module in the preset initial model, where the spatial weight represents the spatial autocorrelation between the current unknown site and each known site;
[0093] Step 250: Take the product of each rough prediction result and the spatial weight corresponding to each rough prediction result as the prediction result corresponding to each unknown site, and determine the model loss through a loss function according to each prediction result and the corresponding preset real data;
[0094] Step 260: Update the parameters of the preset initial model according to the model loss until the model loss meets the preset model convergence condition to obtain a spatial interpolation model.
[0095] In step 210, for the convenience of implementing the technical solution of the present application, in this embodiment, the daily precipitation data of the target observation area (in this embodiment, to solve the problem of inaccurate climate observation in areas with sparse stations, so the target observation area can be any area with relatively few observation stations, and this embodiment takes the Qinghai-Tibet Plateau area as an example) from 1980 to 2020 is used as the data source. For the convenience of data processing and optimizing the presentation effect of precipitation data, the daily precipitation data of the target observation area is merged into monthly precipitation data. Among them, the monthly precipitation data includes the precipitation values (i.e., the above initial data) of all observation stations (or simply referred to as stations, that is, stations carrying precipitation values) in the target observation area. In practical applications, for the convenience of calculating the spatial autocorrelation between different stations subsequently, the monthly precipitation data may also include the elevation, terrain undulation, and longitude and latitude data of each station.
[0096] In this embodiment, obtaining the target training set may include:
[0097] Obtain climate samples of the target observation area at multiple different times, where the climate samples include known sites, unknown sites, and precipitation values corresponding to the known sites in the target observation area;
[0098] Mask the climate sample to obtain the mask data corresponding to the climate sample;
[0099] Generate multiple first feature maps with different scales corresponding to the mask data according to the mask data;
[0100] Perform a weighted pooling operation on each of the first feature maps to extract features from the precipitation values of the known stations in each of the first feature maps, obtain a second feature map corresponding to each of the first feature maps as the training sample, and use the precipitation values after feature extraction as the initial data.
[0101] It can be understood that for the convenience of data processing, the aforementioned monthly precipitation data also needs to be preprocessed in practical applications. Specifically, since this technical solution performs feature extraction and interpolation based on stations, in order to reduce the interference of non-effective data (i.e., the precipitation values carried by unknown stations, usually noise) in the feature extraction process, first, according to the longitude and latitude of each station in the target observation area, rasterize the initial image (with a size of 1km×1km) and generate mask data (that is, the raster with a station is valid data, and the mask data is assigned 1, and the raster without a station is invalid data, and the mask data is assigned 0) to identify the validity of the raster data.
[0102] In this embodiment, according to the mask data obtained after masking, with months as the time unit, generate multiple first feature maps with different scales corresponding to the monthly mask data of the target observation area between 1980 and 2020. In this way, it is easier to obtain the spatial variation trend of the variable to be interpolated in a larger area range, and the feature decomposition from the trend to the spatial details can be completed.
[0103] In this embodiment, all the rasters in the first feature map are regarded as a station, and all stations are divided into known stations and unknown stations according to the validity of the data in the raster, that is, the raster with the aforementioned mask data of 1 is defined as a known station, and the raster with the aforementioned mask data of 0 is defined as an unknown station.
[0104] In this embodiment, perform a weighted pooling operation on the known stations, so as to extract features from the first feature maps with different scales, and obtain a second feature map corresponding to each first feature map. In this way, by using the multi-scale feature maps (i.e., the second feature maps) as the subsequent convolution input, the problem of directly restoring high-resolution data in climate interpolation prediction is converted into the problem of gradually restoring high-resolution data from low-resolution data, reducing the model training difficulty and increasing the stability of the model.
[0105] Specifically, refer to Figure 4 , Figure 4Among them, Weight pooling represents weight pooling, Image Representation is the image representation characterizing the first feature map, and Mask Representation is the mask representation characterizing the mask data corresponding to the first feature map. In Mask Representation, 1 represents the valid value of the current grid as a known site, and 0 represents the invalid value of the current grid as an unknown site. In this technical solution, pooling is performed on the precipitation values in the first feature map with a kernel (pooling kernel) of a 2×2 grid size and a stride of 2 grids. During the pooling process, when each feature extraction is performed by the kernel, if all the grids in the kernel are invalid values, the pooling weight corresponding to the kernel is 0; if the grids in the kernel contain valid values, the pooling weight corresponding to the kernel is the reciprocal of the number of valid values.
[0106] Exemplarily, taking Figure 4 the first feature extraction in as an example, during this feature extraction process, the kernel covers four grids (0, 1, 0, 0), and its corresponding Mask Representation is (0, 1, 1, 0). That is, during the first feature extraction process, the first grid is an invalid value, the second and third grids are valid values, and the fourth grid is an invalid value. The total number of valid values in the kernel is 1, and the number of valid values in the corresponding Mask Representation is 1 + 1 = 2. Then Weight pooling(0, 1, 0, 0) / (0, 1, 1, 0) = 1 / 2 = 0.5. Similarly, during the second feature extraction process, the kernel covers four grids (2, 0, 0, 0), the total number of valid values is 2, and its corresponding Mask Representation is (1, 0, 0, 0), that is, the number of valid values is 1. Then Weight pooling(2, 0, 0, 0) / (1, 0, 0, 0) = 2 / 1 = 2. And so on, feature extraction is performed on the precipitation values in the first feature map with four grids as a kernel and a stride of 2 until all the precipitation values in the first feature map are covered, that is, the weight pooling operation is completed to obtain the second feature map. Multiple second feature maps of different scales corresponding to the mask data of each month are jointly used as a training sample.
[0107] In step 220, the formula for the gated convolution operation is as follows:
[0108]
[0109]
[0110]
[0111] In the formula, represents the j-th rough prediction result obtained after the gated convolution operation, represents the initial data in the convolution region of the i-th convolution kernel in the (l - 1)-th layer; represents the weight of the i-th convolution kernel in the l-th layer; represents the conventional convolution operation; b i l represents the bias of the i-th convolution kernel in the l-th layer; M represents the convolution region, and f(·) represents the activation function.
[0112] It can be understood that the conventional convolution mainly extracts high-level semantic features from the data. The convolution is a matrix operation between a partial region of the input data and the convolution kernel. After adding the corresponding bias to the convolution output value, a non-linear transformation is performed through the activation function. In a convolutional neural network, the filter performs convolution calculations on local input data. After calculating the local data within each data window, the data window continuously slides until all data is calculated. The conventional convolution calculation process is shown in Equation (3).
[0113] In this embodiment, however, gated convolution is used to fill in the data for the grids without stations (i.e., unknown stations) in the second feature maps of different scales (i.e., training samples). By filling the missing regions of the compact latent features into the low-level features (with higher resolution and richer details), the coding efficiency is further improved. Among them, the gated convolution is an improved version of the conventional convolution. The conventional convolution considers all surrounding features, but when interpolating data of climate types, not all surrounding features are effective, and the influence of invalid values on the convolution needs to be eliminated. Using the gated convolution enables the network model to learn a dynamic feature selection mechanism for each channel and each spatial position to assign precipitation values to the unknown station areas. Among them, the calculation process of the gated convolution is shown in Equation (1) and Equation (2). As shown in Equation (2), the conventional convolution operation is normalized through the sigmoid activation function to serve as the weight of the gated convolution operation. The result of the conventional convolution is updated by multiplying this weight by the result of the conventional convolution, and the gated convolution result corresponding to each second feature map is obtained. The data obtained after the gated convolution at each unknown station is used as the rough prediction result.
[0114] Exemplarily, referring to Figure 5 , as Figure 5 shown in the first horizontal row in the middle part, the second feature map with 3 channels (M×N×3, where M and N represent the length and width of the second feature map, and 3 represents the number of channels) is subjected to conventional convolution, and the bias b1 is added, and then the conventional convolution result is obtained through the activation function ReLU; as Figure 5As shown in the second horizontal row of the middle part, after convolving the second feature map, a bias b2 is added, and the convolution result after adding the bias b2 is normalized through the sigmoid activation function to obtain the weight of the gated convolution operation; then the conventional convolution result of the first horizontal row is multiplied by the convolution weight of the second horizontal row to obtain the single-channel gated convolution result (M×N×1).
[0115] In step 230, the weight factor may include site coordinates, elevation, and terrain undulation.
[0116] In step 240, according to the initial data and the weight factor corresponding to each known site, the spatial weight corresponding to each unknown site can be determined through the attention module in the preset initial model, including:
[0117] Determine the Euclidean distance between the unknown site and each target known site through the attention module in the preset initial model:
[0118]
[0119] where distance i represents the Euclidean distance, represents the coordinate difference between the current unknown site and the target known site, represents the elevation difference between the current unknown site and the target known site, represents the difference in terrain undulation between the current unknown site and the target known site, represents the difference between the initial predicted value corresponding to the current unknown site and the precipitation value corresponding to the target known site, and the target known site represents a known site within a preset distance range from the current unknown site;
[0120] Normalize the Euclidean distance to determine the prediction weight:
[0121] weight i = softmin(distance i ,{distance0,distance1,…,distance n ) (5)
[0122] where weight i represents the prediction weight, {distance0,distance1,…,distance n} represents the set composed of each Euclidean distance.
[0123] In this embodiment, the spatial weight is determined by the attention module. This spatial weight takes into account the location and terrain information (i.e., elevation, terrain undulation, etc.) between the unknown site and the known sites, and introduces the difference in precipitation values between the two as one of the variables, so that the predicted value of the unknown site not only pays attention to the spatial autocorrelation between the unknown site and the known sites, but also pays attention to the characteristic association between the precipitation values. During the process of gradually filling the unknown site, the generation of data at all levels has a more reliable reference, improving the accuracy of the model in predicting the precipitation value interpolation of the unknown site.
[0124] Specifically, referring to Figure 6 , Figure 6 In, Station features represents the site features, Precipitation features represents the precipitation value features, flatten means flattening the raster, for example, transforming a raster of size 2×2 into a raster of size 1×4 for similarity measurement, reshape means restoring the raster, for example, restoring a raster of size 1×4 to a raster of size 2×2, pdist represents the Euclidean distance calculation, Attention map represents the attention map composed of the weight factor and the prediction weight, f(x), g(x), and h(x) respectively represent different linear functions, which can be understood as the fully connected layers in the attention module, where x represents the input composed of the weight factor. In practical applications, the weight factor is linearly transformed by f(x) and g(x) and the Euclidean distance is calculated, and the precipitation features are obtained by linearly transforming the precipitation value by h(x). Multiplying the Euclidean distance by the precipitation features can obtain the aforementioned attention map.
[0125] In step 250, the preset real data can be the real precipitation value obtained by field rainfall observation of the geographical location corresponding to the unknown site. Substitute this real precipitation value and its corresponding prediction result into the preset loss function of the model, and the final model loss is determined by the preset loss function. This model loss characterizes the error magnitude between the prediction result and the real precipitation value.
[0126] Among them, the loss function is as follows:
[0127]
[0128] In the formula, n represents the number of samples, h(x k ) represents the prediction result of the kth training sample, y k represents the preset real data corresponding to the kth training sample, and I i,j represents the prediction result corresponding to the unknown site located at (i,j).
[0129] In step 260, after the model loss is determined by the loss function in step 250, the parameters of the preset initial model are updated according to the model loss until the model loss meets the preset model convergence condition, the model training is completed, and the spatial interpolation model is obtained.
[0130] In this embodiment, the preset model convergence condition may be that the model loss is stable and at a minimum value.
[0131] In summary, refer to Figure 7 , mask the target observation area containing stations and initial data in the training sample (Stations masking) to obtain mask data (Stations value), and perform weight pooling (Weight pooling) on the mask data to obtain multiple second feature maps of different scales, and then perform gated convolution operation (GateConv) on the second feature map to obtain the gated convolution result, and then determine the spatial weight of each unknown station according to the spatial autocorrelation (weight factor) between the unknown station and the known station, and perform feature fusion of the gated convolution result with the attention map (Attention) composed of the weight factor and the spatial weight to obtain the prediction result corresponding to each second feature map, and finally perform feature fusion on the prediction results corresponding to multiple second feature maps through conventional convolution (Conv) to obtain the interpolated precipitation map of the target observation area. Among them, the prediction results corresponding to the second feature maps of different scales are feature fused through convolution to obtain the final interpolation result, which is a technical means known to technicians in the field of convolutional neural networks, and the technical solution proposed in this application will not be described in detail.
[0132] Please refer to Figure 8 The present application also provides a spatial interpolation device 300, which includes at least one software function module that can be stored in the memory 102 in the form of software or firmware or fixed in the operating system (OS) of the electronic device 100. The processor 101 is used to execute the executable modules stored in the memory 102, such as the software function modules and computer programs included in the spatial interpolation device 300.
[0133] The spatial interpolation device 300 includes a first acquisition unit 310, a prediction unit 320, a first determination unit 330, a second determination unit 340 and an interpolation unit 350, and the functions of each unit may be as follows:
[0134] A first acquisition unit 310 is used to acquire a target observation area, where the target observation area includes a plurality of known sites, a plurality of unknown sites, and original climate data corresponding to each known site;
[0135] A prediction unit 320, configured to predict the rough predicted climate data of each unknown site according to the original climate data;
[0136] A first determination unit 330, configured to determine the prediction weight of each unknown site according to the spatial autocorrelation in space between each unknown site and the known sites that meet the preset conditions;
[0137] A second determination unit 340, configured to determine the predicted climate data of each unknown site according to the rough predicted climate data and the prediction weight corresponding to each unknown site;
[0138] An interpolation unit 350, configured to perform climate data interpolation on the climate data to be predicted according to the predicted climate data to obtain an interpolation result.
[0139] Optionally, the spatial interpolation device 300 further includes:
[0140] A second acquisition unit, configured to acquire a target training set, where the target training set includes a plurality of training samples, and each training sample includes known sites and unknown sites in a target observation area, and initial data corresponding to the known sites;
[0141] A convolution unit, configured to perform gated convolution operations on each training sample in the training set through a gated convolution layer in a preset initial model to predict the rough prediction result of the unknown site in each training sample;
[0142] A third acquisition unit, configured to acquire weight factors corresponding to each known site and each unknown site;
[0143] A third determination unit, configured to determine the spatial weight corresponding to each unknown site through an attention module in the preset initial model according to the initial data and the weight factors corresponding to each known site, where the spatial weight represents the spatial autocorrelation between the current unknown site and each known site;
[0144] A fourth determination unit, configured to use the product of each rough prediction result and the spatial weight corresponding to each rough prediction result as the prediction result corresponding to each unknown site, and determine the model loss through a loss function according to each prediction result and the corresponding preset real data;
[0145] A training unit, configured to update the parameters of the preset initial model according to the model loss until the model loss meets a preset model convergence condition to obtain a spatial interpolation model.
[0146] Optionally, the second acquisition unit is further configured to:
[0147] Obtain climate samples of the target observation area at multiple different times, where the climate samples include known stations, unknown stations within the target observation area, and precipitation values corresponding to the known stations;
[0148] Perform masking on the climate samples to obtain masked data corresponding to the climate samples;
[0149] Generate multiple first feature maps with different scales corresponding to the masked data according to the masked data;
[0150] Perform weighted pooling operations on each of the first feature maps to extract features from the precipitation values of the known stations in each of the first feature maps, obtain second feature maps corresponding to each of the first feature maps as the training samples, and use the precipitation values after feature extraction as the initial data.
[0151] Optionally, the formula for the gated convolution operation is as follows:
[0152]
[0153]
[0154]
[0155] In the formula, represents the jth coarse prediction result obtained after the gated convolution operation, represents the initial data in the convolution region of the ith convolution kernel of the (l - 1)th layer; represents the weight of the ith convolution kernel of the lth layer; represents the conventional convolution operation; b i l represents the bias of the ith convolution kernel of the lth layer; M represents the convolution region, and f(·) represents the activation function.
[0156] Optionally, the third determination unit is further configured to:
[0157] Determine the Euclidean distance between the unknown station and each target known station through the attention module in the preset initial model:
[0158]
[0159] In the formula, distance i represents the Euclidean distance, represents the coordinate difference between the current unknown station and the target known station, represents the elevation difference between the current unknown station and the target known station, represents the difference in terrain undulation between the current unknown station and the target known station, It represents the difference between the initial predicted value corresponding to the current unknown site and the precipitation value corresponding to the target known site, where the target known site represents a known site within a preset distance range from the current unknown site;
[0160] Normalize the Euclidean distance to determine the prediction weight:
[0161] weight i = softmin(distance i ,{distance0,distance1,…,distance n )
[0162] In the formula, weight i represents the prediction weight, and {distance0,distance1,…,distance n} represents the set composed of each Euclidean distance.
[0163] Optionally, the loss function is as follows:
[0164]
[0165] In the formula, n represents the number of samples, h(x k ) represents the prediction result of the k-th training sample, y k represents the preset true data corresponding to the k-th training sample, and I i,j represents the prediction result corresponding to the unknown site at (i,j).
[0166] In this embodiment, the processor 101 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 101 can be a general-purpose processor. For example, the processor 101 can be a Central Processing Unit (CPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0167] The memory 102 can be, but is not limited to, a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc. In this embodiment, the memory 102 can be used to store a target observation area, coarse predicted climate data, preset conditions, prediction weights, predicted climate data, interpolation results, a target training set, coarse prediction results, weight factors, spatial weights, prediction results, model losses, preset model convergence conditions, a spatial interpolation model, etc. Of course, the memory 102 can also be used to store programs, and after receiving an execution instruction, the processor 101 executes the program.
[0168] It can be understood that Figure 1 the structure of the electronic device 100 shown in Figure 1 is only a schematic structural diagram, and the electronic device 100 may further include more Figure 1 components than those shown.
[0169] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described electronic device 100 can refer to the corresponding processes of each step in the foregoing method, and will not be elaborated herein.
[0170] This application embodiment also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is enabled to execute the spatial interpolation method as described in the above embodiment.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented through hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of this application.
[0172] In summary, the embodiments of the present application provide a spatial interpolation method, apparatus, electronic device, and storage medium. In the technical solution, first, a target observation area is obtained, where the target observation area includes a plurality of known stations, a plurality of unknown stations, and the original climate data corresponding to each known station. Then, the rough predicted climate data of each unknown station is predicted based on the original climate data, and the prediction weight of each unknown station is determined according to the spatial autocorrelation in space between each unknown station and all known stations. Then, the predicted climate data of each unknown station is determined according to the rough predicted climate data and the prediction weight corresponding to each unknown station. Finally, climate data interpolation is performed on the climate data to be predicted based on the predicted climate data to obtain an interpolation result. In this way, the problem that the interpolation result has low accuracy and large error in the traditional spatial interpolation method for geographical locations with sparse stations can be improved.
[0173] In the embodiments provided in the present application, it should be understood that the disclosed apparatus, system, and method can also be implemented in other ways. The embodiments of the apparatus, system, and method described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, the functional modules in each embodiment of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0174] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A spatial interpolation method, characterized in that: The method comprises: Obtain a target observation area, the target observation area including a plurality of known sites, a plurality of unknown sites and original climate data corresponding to each known site; wherein the target observation area is rasterized, the unknown site represents any raster where no climate observation site is set, the known site represents any raster where a climate observation site is set, and the climate observation site directly observes and obtains the climate data of the location where the climate observation site is located; Predicting the crude predicted climate data of each unknown station based on the original climate data; Determine the prediction weight of each unknown site based on the spatial autocorrelation between each unknown site and the known sites that meet the preset conditions; Determining predicted climate data for each unknown site according to the coarse predicted climate data corresponding to each unknown site and the prediction weight; Performing climate data interpolation on the climate data at the location of the current unknown site according to the predicted climate data corresponding to the current unknown site to obtain an interpolation result corresponding to the current unknown site; Determining the predicted climate data for each unknown station is achieved through a spatial interpolation model, wherein the training process of the spatial interpolation model is as follows: Acquire a target training set, wherein the target training set includes a plurality of training samples, each of the training samples includes known sites and unknown sites in a target observation area, and initial data corresponding to the known sites; Performing a gated convolution operation on each training sample in the training set by using a gated convolution layer in a preset initial model to predict a rough prediction result of the unknown site in each training sample; Obtaining a weight factor corresponding to each of the known sites and each of the unknown sites; According to the initial data and the weight factor corresponding to each known site, a spatial weight corresponding to each unknown site is determined by an attention module in the preset initial model, wherein the spatial weight represents the spatial autocorrelation between the current unknown site and each known site; The product of each of the rough prediction results and the spatial weight corresponding to each of the rough prediction results is used as the prediction result corresponding to each of the unknown sites, and according to each of the prediction results and the corresponding preset real data, the loss of the preset initial model is determined by the loss function; The parameters of the preset initial model are updated according to the loss until the loss meets the preset model convergence condition, thereby obtaining a spatial interpolation model.
2. The method according to claim 1, characterized in that: Get the target training set, including: Acquire climate samples of the target observation area at multiple different times, wherein the climate samples include known stations, unknown stations, and precipitation values corresponding to the known stations within the target observation area; Masking the climate sample to obtain mask data corresponding to the climate sample; Generating a plurality of first feature maps of different scales corresponding to the mask data according to the mask data; A weight pooling operation is performed on each of the first feature maps to extract features of the precipitation values of the known stations in each of the first feature maps, to obtain a second feature map corresponding to each of the first feature maps as the training sample, and the precipitation values after feature extraction are used as the initial data.
3. The method according to claim 1, characterized in that The formula for the gated convolution operation is as follows: In the formula, represents the jth rough prediction result obtained after the gated convolution operation, Represents the initial data in the convolution area of the i-th convolution kernel of the l-1th layer; Represents the weight of the i-th convolution kernel in the l-th layer; represents the conventional convolution operation; b i l represents the bias of the i-th convolution kernel of the l-th layer; M represents the convolution area, and f(·) represents the activation function.
4. The method according to claim 1, characterized in that: The initial data includes a precipitation value representing the precipitation amount of the known site in a unit time, and the weight factor includes the site coordinates, elevation and terrain relief; Then, according to the initial data and the weight factor corresponding to each known site, the spatial weight corresponding to each unknown site is determined by the attention module in the preset initial model, including: The Euclidean distance between the unknown site and each target known site is determined by the attention module in the preset initial model: Where, distance i represents the Euclidean distance, Indicates the coordinate difference between the current unknown site and the target known site. Indicates the elevation difference between the current unknown site and the target known site. Indicates the difference in terrain relief between the current unknown site and the target known site. represents the difference between the initial prediction value corresponding to the current unknown site and the precipitation value corresponding to the target known site, where the target known site represents a known site within a preset distance range from the current unknown site; The Euclidean distance is normalized to determine the spatial weight: weight i =softmin(distance i ,{distance0,distance1,…,distance n }) In the formula, weight i represents the spatial weight, {distance0,distance1,…,distance n } represents the set consisting of various Euclidean distances.
5. The method according to claim 1, characterized in that: The loss function is as follows: In the formula, n represents the number of training samples, h(x k ) represents the prediction result of the kth training sample, y k represents the preset real data corresponding to the kth training sample, I i,j Represents the prediction result corresponding to the unknown site at (i, j).
6. A spatial interpolation device, characterized in that: The device comprises: A first acquisition unit is used to acquire a target observation area, wherein the target observation area includes a plurality of known sites, a plurality of unknown sites, and original climate data corresponding to each known site; wherein the target observation area is rasterized, the unknown site represents any raster where no climate observation site is set, and the known site represents any raster where a climate observation site is set, and the climate observation site directly observes and obtains the climate data of the location where the climate observation site is located; A prediction unit, used for predicting the rough predicted climate data of each unknown station according to the original climate data; A first determination unit is used to determine a prediction weight of each unknown site according to a spatial autocorrelation between each unknown site and a known site that meets a preset condition; A second determining unit, configured to determine the predicted climate data of each unknown site according to the coarse predicted climate data corresponding to each unknown site and the prediction weight; An interpolation unit, configured to interpolate the climate data at the location of the current unknown station according to the predicted climate data corresponding to the current unknown station, to obtain an interpolation result corresponding to the current unknown station; The device also includes: A second acquisition unit is used to acquire a target training set, wherein the target training set includes a plurality of training samples, each of which includes known sites and unknown sites in a target observation area, and initial data corresponding to the known sites; A convolution unit, configured to perform a gated convolution operation on each training sample in the training set through a gated convolution layer in a preset initial model, so as to predict a rough prediction result of the unknown site in each training sample; A third acquisition unit, used to acquire a weight factor corresponding to each of the known sites and each of the unknown sites; A third determination unit is used to determine the spatial weight corresponding to each of the unknown sites through the attention module in the preset initial model according to the initial data corresponding to each of the known sites and the weight factor, wherein the spatial weight represents the spatial autocorrelation between the current unknown site and each of the known sites; A fourth determination unit is used to take the product of each of the rough prediction results and the spatial weight corresponding to each of the rough prediction results as the prediction result corresponding to each of the unknown sites, and determine the loss of the preset initial model through a loss function according to each of the prediction results and the corresponding preset real data; A training unit is used to update the parameters of the preset initial model according to the loss until the loss meets the preset model convergence condition to obtain a spatial interpolation model.
7. An electronic device, characterized in that: The electronic device comprises a processor and a memory coupled to each other, wherein the memory stores a computer program. When the computer program is executed by the processor, the electronic device executes the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Rainfall space interpolation method and system, electronic equipment and readable storage medium
CN113553696A
Rainfall spatial data processing method and device, electronic equipment and storage medium
CN114036446A