A high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account spatiotemporal heterogeneity

By recombining high-dimensional specific wet data into locally similar sub-tensor blocks, and using tensor CP decomposition and mask tensor construction, taking into account the structural characteristics of spatiotemporal heterogeneity, performing feature extraction and reconstruction, the problem of existing methods ignoring spatiotemporal heterogeneity when reconstructing sparse spatiotemporal field data is solved, and the reconstruction accuracy of extremely sparse specific wet data is improved.

CN114004056BActive Publication Date: 2025-05-16YANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110993497.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-27
Publication Date
2025-05-16
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

When reconstructing sparse spatiotemporal field data, the existing methods ignore spatiotemporal heterogeneity, resulting in low high-dimensional interpolation accuracy of extremely sparse and wet data.

Method used

A high-dimensional specific wet data reconstruction method based on extremely sparse sampling is proposed. By reorganizing the original data into locally similar sub-tensor blocks, and constructing it using tensor CP decomposition and mask tensor, it takes into account the spatial and temporal heterogeneity structural characteristics, and performs feature extraction and reconstruction.

Benefits of technology

Effectively take into account space-time heterogeneity, improve the reconstruction accuracy of extremely sparse specific wet data, and achieve good reconstruction of atmospheric multidimensional signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114004056B_ABST
    Figure CN114004056B_ABST
Patent Text Reader

Abstract

The invention discloses a high-dimensional specific humidity data reconstruction method based on extremely sparse sampling that takes into account temporal and spatial heterogeneity, comprising the following steps: reorganizing original multidimensional specific humidity data into locally similar sub-tensor blocks; treating each obtained sub-tensor block as a whole, performing sample screening on each sub-tensor block, selecting representative samples, and constructing a mask tensor; extracting characteristic components of the atmospheric environment multidimensional signal on each sub-tensor block data at different scales based on the atmospheric distribution structure pattern; further performing significance test on the characteristic component sets at different scales, screening the significant components and performing feature reconstruction; and performing weighted summation on the feature reconstruction data of each block data and the original sub-tensor block data to complete the reconstruction of the original multidimensional specific humidity data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of physical geography and meteorological dynamics technology, and specifically relates to a high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account temporal and spatial heterogeneity. Background Art

[0002] Early data collection showed discrete distribution characteristics in space and time. However, with the continuous development of data collection, storage and communication technology, the volume of data has grown exponentially, and the massive generalized spatiotemporal data has shown a global trend in space and a continuous trend in time. Affected by the complex internal physical mechanisms and the coupling correlation between multiple factors, atmospheric signals show significant spatiotemporal heterogeneity. At the same time, the existing monitoring data often have incomplete data problems. Affected by economic and geographical conditions, atmospheric data in some areas are difficult to obtain. Therefore, the lack and sparse distribution of spatiotemporal data have also become a common phenomenon. And with the continuous deepening of data analysis needs and the continuous subdivision of spatiotemporal granularity, effectively taking into account spatiotemporal heterogeneity in the process of reconstructing sparse spatiotemporal field data has an important impact on more in-depth research.

[0003] At present, the interpolation methods for high-dimensional sparse meteorological data can be roughly divided into two categories: statistical methods and machine learning methods. Generally speaking, complex spatiotemporal statistical interpolation methods for the atmosphere usually require solving partial differential equations point by point to calculate the optimal weights of interpolation samples. Such as high-precision surface models, Kriging models, etc. These methods take into account the spatial autocorrelation and spatial heterogeneity of data distribution to a certain extent, and have a certain effect on the reconstruction accuracy, but due to the complexity of model solution, they are usually difficult to deploy. Machine learning methods such as traditional tensor decomposition and semi-supervised learning. It is usually necessary to construct a target function for solution, and use numerical calculation methods such as gradient descent to iteratively train the model to achieve the optimal reconstruction accuracy. However, traditional tensor decomposition usually requires the use of an exhaustive method to complete the selection of the number of features R, and its internal mechanism is not clear. These methods regard the spatiotemporal field as a whole and ignore the spatiotemporal heterogeneity, which will affect the accuracy of high-dimensional interpolation of extremely sparse atmosphere with spatiotemporal heterogeneity. Summary of the invention

[0004] Purpose of the invention: In order to solve the problem that the existing methods ignore the spatiotemporal heterogeneity of the specific humidity data because they regard the specific humidity data as a whole, and then have poor accuracy when performing high-dimensional interpolation on extremely sparse specific humidity data with spatiotemporal heterogeneity, the present invention proposes a high-dimensional specific humidity data reconstruction method based on extremely sparse sampling that takes into account spatiotemporal heterogeneity. This method not only considers the multi-dimensional coupling characteristics of spatiotemporal field data, but also effectively takes into account the spatiotemporal heterogeneous structural characteristics of spatiotemporal field data. It is a major breakthrough in the reconstruction of sparsely sampled high-dimensional specific humidity data.

[0005] Technical solution: A high-dimensional specific humidity data reconstruction method based on extremely sparse sampling that takes into account temporal and spatial heterogeneity, comprising the following steps:

[0006] Step 1: Reorganize the original high-dimensional specific humidity data into multiple locally similar sub-tensor blocks Recorded as the original sub-tensor block data;

[0007] Step 2: Block each sub-tensor Y obtained in step 1 i As a whole, for each sub-tensor block Y i Performing operations S100 to S400 in sequence to complete the reconstruction of high-dimensional specific humidity data;

[0008] S100: Based on the distribution structure of specific humidity data in the atmosphere, data clustering is performed on the sub-tensor block data, and representative samples are selected to construct a mask tensor;

[0009] S200: based on the distribution structure of specific humidity data, extract the characteristic components of each sub-tensor block data from different scales;

[0010] S300: performing a significance test on the feature component sets at different scales extracted in S200, screening the significant components and performing feature reconstruction to obtain feature reconstruction data;

[0011] S400: performing weighted summation on the feature reconstruction data of each sub-tensor block data and the original sub-tensor block data respectively.

[0012] Furthermore, the step 1 specifically includes:

[0013] Step 1.1: Organize the original high-dimensional specific humidity data into a high-order tensor, and divide the high-order tensor into multiple sub-tensor blocks with the same dimension size;

[0014] Step 1.2: Aggregate locally similar sub-tensor blocks through spatiotemporal heterogeneity measures to form locally similar sub-tensor blocks.

[0015] Furthermore, the step 1.1 specifically includes:

[0016] Organize the raw high-dimensional specific humidity data into high-order tensors;

[0017] Partition high-order tensors based on attribute values;

[0018] The high-order tensor is split according to the spatial and temporal dimensions; the temporal dimension is split according to the data update interval, and the spatial dimension is split into regular blocks of the same size.

[0019] Furthermore, the step 1.2 specifically includes:

[0020] Assume that the decomposed sub-tensor blocks For adjacent sub-tensor block data, high-dimensional similarity calculation is performed based on the spatiotemporal heterogeneity measure to obtain the similarity between adjacent sub-tensor block data:

[0021] ρ i,i+1 =SC(X i ,X i+1 ) (2)

[0022] The similarity ρ between adjacent sub-tensor block data i,i+1 Compare with the given similarity threshold δ: If ρ i,i+1 >δ, then Otherwise X i =Y i , X i+1 =Y i+1 .

[0023] Furthermore, S100 specifically includes:

[0024] Combined with the distribution structure of specific humidity data in the atmosphere [min, mid], (mid, max], for the sub-tensor block Y i , the sub-tensor block Y i The data on (Y i ) pjk Different weights are assigned according to their different structures in the entire atmospheric value distribution:

[0025]

[0026] And α+β=1;

[0027] Among them, p, j, k represent the sub-tensor block Y i Index in three dimensions: longitude, latitude and time;

[0028] And use this to construct the sub-tensor block Y i A set of weight coefficients for the data;

[0029] For the weight tensor Weight((Y i ) pjk ), construct the mask tensor

[0030]

[0031] Among them, (Q i ) pjk Denotes the mask tensor Q′ p The value at position (p,j,k).

[0032] Furthermore, S200 specifically includes:

[0033] According to the sub-tensor block And the mask tensor, perform weighted tensor CP decomposition to obtain:

[0034]

[0035] in, for The rth feature in longitude, latitude and time dimensions respectively, For data information that is not captured, is the weight coefficient set of the rth feature; * represents the scalar multiplication of the elements at the corresponding positions of the two tensors, Q i is the mask tensor.

[0036] Furthermore, S300 specifically includes:

[0037] According to formula (6), calculate the sub-tensor block Y i The rth group of eigencomponents on In the set of all feature components The cumulative variance contribution rate of is:

[0038]

[0039] In the formula, is a set of weight coefficients for each sub-tensor block data;

[0040] Based on the calculated cumulative variance contribution rate and the given cumulative variance contribution rate threshold γ, only when CVC (i,r) >γ, the subtensor block Y i The rth group of characteristic components is the significant component;

[0041] Get the sub-tensor block Y i The set of significant components after the above screening is The eigenvalue set is denoted as

[0042] The significant components are combined and the eigenvalue set Perform tensor reconstruction to obtain feature reconstruction data

[0043] Furthermore, S400 specifically includes:

[0044] Based on the following formula, the feature reconstruction data of the sub-tensor block data and the original sub-tensor block data are weighted and summed respectively:

[0045]

[0046] In the formula, Qi is the mask tensor, I represents the size and Q i Same tensor with all positions set to 1, Restructure the data for the feature.

[0047] Beneficial effects: From the perspective of atmospheric field data decomposition and reorganization, the present invention extracts a series of characteristic tensors of the original sparse data, and implements the interpolation of sparsely sampled specific humidity data based on the weighted summation of the characteristic tensors, thereby completing the reconstruction of the multi-dimensional signal of the extremely sparsely sampled atmospheric environment; compared with the prior art, it has the following advantages:

[0048] (1) The method of the present invention organizes the atmospheric multidimensional data into a high-order tensor. In the absence of prior knowledge, the original data is divided into uniform sub-tensors, and then the locally similar sub-tensors are aggregated together through the spatiotemporal heterogeneity measurement to form locally similar sub-tensor blocks. Finally, the tensor CP decomposition is used to complete the feature extraction of atmospheric data with spatiotemporal heterogeneity, which not only considers the retention of the structural characteristics of the original data, but also takes into account the local spatiotemporal heterogeneity. The reconstruction based on extremely sparsely sampled high-order specific humidity data is achieved while taking into account the spatiotemporal heterogeneity.

[0049] (2) The present invention sets different weights for the sub-tensor blocks after the original data is divided into blocks, and then constructs a mask matrix, which objectively reflects the distribution of atmospheric humidity in the real world;

[0050] (3) The method of the present invention sets the cumulative variance contribution rate threshold as the cutoff point for sub-tensor block feature extraction, that is, there is only one input parameter, which minimizes the intervention of human factors and ensures the high efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of the invention;

[0052] Figure 2 It is a schematic diagram of the correlation coefficients of the time principal components of adjacent Shum data;

[0053] Figure 3 It is the graph of the number of characteristic components and the cumulative variance contribution rate;

[0054] Figure 4 This is a schematic diagram of the reconstruction results of extremely sparsely sampled Shum multidimensional signals, where: Figure 4 (a) is the original complete Shum data, Figure 4 (b) is the missing rate of sparse data is 90%, Figure 4 (c) is the extraction result when the number of features is 70. DETAILED DESCRIPTION

[0055] The atmospheric environment multidimensional signal is a discrete high-dimensional signal with a coupling relationship in time and space, and it is naturally affected by natural conditions and has spatiotemporal heterogeneity. At present, when interpolating sparse spatiotemporal fields with structural heterogeneity, existing methods often study the data as a whole, ignoring the characteristics of spatiotemporal heterogeneity. In the present invention, the discrete high-order specific humidity data are organized into a third-order tensor with three dimensions, namely longitude, latitude, and time; in the absence of prior knowledge, the original high-order specific humidity data are reorganized into locally similar sub-tensor blocks; each obtained sub-tensor block is regarded as a whole, and samples are screened for each sub-tensor block, representative samples are selected, and a mask tensor is constructed; based on the atmospheric distribution structure pattern, the characteristic components of the atmospheric environment multidimensional signal on each sub-tensor block data are extracted from different scales; further, the characteristic component sets at different scales are tested for significance, the significant components are screened, and the features are reconstructed; the feature reconstruction data of each block data and the original sub-tensor block data are weighted summed to complete the reconstruction of the original high-order specific humidity data. The method of the present invention reconstructs extremely sparse high-order specific humidity data while taking into account the temporal and spatial heterogeneity. The results show good accuracy and are a major breakthrough in the reconstruction of atmospheric multi-dimensional signals.

[0056] The present invention is further described below in conjunction with the accompanying drawings.

[0057] like Figure 1 As shown, a multi-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account the temporal and spatial heterogeneity of the present invention comprises the following steps:

[0058] Step 1: While taking into account the multidimensional characteristics and spatiotemporal heterogeneity of sparsely sampled specific humidity data, the original high-order specific humidity data is divided into uniform sub-tensor blocks, and locally similar sub-tensor blocks are aggregated together through spatiotemporal heterogeneity measurement to form locally similar sub-tensor blocks. Specifically, it includes the following sub-steps:

[0059] The original high-order specific humidity data is organized into a high-order tensor. Considering the significant spatiotemporal heterogeneity of the spatiotemporal field data, the overall tensor decomposition will lead to feature estimation bias. Therefore, in the absence of prior knowledge, the high-order tensor is split into multiple sub-tensor blocks of the same dimensional size.

[0060] For the original high-order tensor Split the high-order tensor into uniform sub-tensor blocks and define the split operator as follows:

[0061]

[0062] Among them, n represents the number of sub-tensor blocks divided, Represents the decomposed sub-tensor blocks, where each sub-tensor block contains local spatial structure and time information.

[0063] In the implementation, the number of splits for each dimension is customized according to actual needs, and the original high-order tensor is first divided based on the attribute value, because the data range and characteristics between different attributes may vary significantly. Then it is split according to the spatial and temporal dimensions to further reduce the imbalance of dimensions. The division of the time dimension is usually the data update interval, and the spatial dimension is usually divided into regular blocks of the same size.

[0064] Perform high-dimensional similarity calculation (Similarity Calculation (SC)) on adjacent sub-tensor block data:

[0065] ρ i,i+1 =SC(X i ,X i+1 ) (2)

[0066] For a given similarity threshold δ, the similarities between adjacent sub-tensor block data are compared, and the sub-tensor block data are merged based on the similarity between the sub-tensor block data to realize the reorganization of the original tensor block data, that is: i,i+1 >δ, then Otherwise, X i =Y i and X i+1 =Y i+1 Still maintained as two independent sub-tensor blocks.

[0067] Assume that the reorganized tensor block data is represented as It can be considered as data reorganized based on the heterogeneity of the original data, that is, the structure within each block of data is relatively uniform, and the structural differences between blocks are relatively large.

[0068] Step 2: Step 1 is only valid for complete atmospheric data. When the reorganized tensor block data When it is a sparse tensor, a mask tensor is also needed to identify whether the data at a certain time and space is missing, thereby ensuring the effectiveness of the decomposition model, as follows:

[0069] Each block of data is still sparsely distributed data in essence. Combined with the distribution structure of specific humidity data in the atmosphere [min, mid], (mid, max], for each block of data Y i , for each specific value (Y i ) pjk Assign different weights, expressed as follows:

[0070]

[0071] And α+β=1.

[0072] Among them, p, j, k represent the sub-tensor block Y i Index in three dimensions: longitude, latitude, and time.

[0073] After this operation, the data on each sparse sub-tensor block can be assigned different weights according to its different structures in the entire atmospheric value distribution.

[0074] For the weight tensor Weight((Y i ) pjk ), in order to maximize the influence of representative sample values ​​on the final sparse data feature estimation, a mask tensor is constructed The definition is as follows:

[0075]

[0076] Among them, (Q i ) pjk Represents the mask tensor Q′ p The value at position (p,j,k).

[0077] Step 3: Based on the distribution structure pattern of the specific humidity data, the characteristic components of the atmospheric environment multidimensional signal on each sub-tensor block data are extracted from different scales, specifically:

[0078] For the reorganized sub-tensor block data Combined with the constructed mask tensor The weighted tensor CP decomposition is performed as follows:

[0079]

[0080] in, for The rth feature in longitude, latitude and time dimensions respectively, For data information that is not captured, is the weight coefficient set of the rth feature; * represents the scalar multiplication of the elements at corresponding positions of the two tensors.

[0081] Step 4: Based on step 3, further perform significance test on feature component sets at different scales, select significant components and perform feature reconstruction; specifically:

[0082] Based on the weight coefficient set of each sparse block data Construct the corresponding feature component set Statistical indicators such as Cumulative Variance Contribution (CVC) are as follows:

[0083]

[0084] Among them, CVC (i,r) Represents the sub-tensor block Y i The characteristic components on In the set of all feature components The cumulative variance contribution rate of .

[0085] Based on this indicator and the given cumulative variance contribution rate threshold γ, the extracted feature components are screened. That is: CVC (i,r) >γ, then the sub-tensor block Y i The rth group of characteristic components is filtered as the final set of feature estimation components. Assume that the sub-tensor block Y i The set of feature components after the above screening is recorded as The eigenvalue set is denoted as

[0086] Step 5: By performing weighted summation on the feature reconstruction data of each block data and the original sub-tensor block data, the multi-dimensional signal reconstruction of the extremely sparse atmospheric field model can be completed.

[0087] The above-selected feature component set and the eigenvalue set The tensor reconstruction is performed as follows: and with the corresponding tensor block Y i The weighted summation is performed as follows:

[0088]

[0089] Among them, I represents the size and Q i A tensor with all positions equal to 1.

[0090] Each sub-tensor block data By performing the above operations in sequence, the original sparse atmospheric environment reconstruction can be obtained.

[0091] The present invention is further described below in conjunction with embodiments.

[0092] The reconstruction method of the present invention is verified by experiments. The 2.5°X 2.5° atmospheric reanalysis daily average specific humidity data from January 1, 1948 to December 31, 2010 released by NOAA (http: / / www.cdc.noaa.gov / Composites) is selected as the experimental data, that is, the specific humidity data set is stored as a (longitude×latitude×time) tensor 90% of the data is removed from Shum to construct sparse data, and the proposed method is used on sparse data to verify the feature extraction of data with different sparsity. The maximum iteration step is set to 2000, and the initial value of the iteration is set to a random number. Relative error, correlation coefficient and common statistical indicators such as maximum value, minimum value, median value, mean value and standard deviation are selected as modeling evaluation indicators.

[0093] The original data is the daily average data. Assuming that only the structural heterogeneity of the data in the time dimension is considered, the integrity of the data in space is maintained, and only the time dimension is divided. In the absence of prior knowledge, the data is evenly divided into blocks. Therefore, the original data is divided into 12 blocks along the time dimension, as shown in Table 1.

[0094] Table 1 Uniform division of data

[0095]

[0096] After the original data is evenly divided into blocks, each block still has the characteristics of high dimension and large data volume. Directly applying the correlation coefficient judgment in the data space will lead to deviations in the results. Based on the similarity between the sub-tensor blocks, the time feature principal component of each block is extracted, and the correlation coefficient of the time principal component of adjacent data is calculated, such as Figure 2 As shown in the figure, we can see that the correlation coefficient is between 0.3662 and 0.8617, indicating that the characteristic principal component structure in different time periods is quite different, that is, the data is significantly heterogeneous in the time dimension.

[0097] Based on the above-mentioned correlation calculation of the characteristic principal components in different time periods, for adjacent block data with relatively large correlation, it can be considered that their structures in the feature space are very similar. From the characteristics of tensor decomposition, it can be seen that the original data space is composed of linear / nonlinear configurations of components in each dimension in the feature space, so it can be considered that similar data in the feature space are also similar in the data space. Taking the correlation coefficient of 0.65 as the threshold, that is, if the correlation coefficient of adjacent block data is greater than 0.65, it is considered that the adjacent block data is structurally similar, so the adjacent blocks can be merged and processed as a whole. If it is less than 0.65, it is considered that the structural differences between the two adjacent blocks are relatively large, so the adjacent blocks should be processed separately. Through the merging operation, the original data can be reorganized into five sub-tensors. As shown in Table 2.

[0098] Table 2 Merging of Shum block data

[0099]

[0100] A mask tensor is constructed for each sub-tensor block, and the weight coefficient set of the features of each sub-tensor block is solved by weighted tensor CP decomposition, and then the variance contribution rate and cumulative variance contribution rate are solved. Based on these two indicators and the given variance contribution rate threshold ε and cumulative variance contribution rate threshold γ, the extracted feature components are screened. The given variance contribution rate threshold ε and cumulative variance contribution rate threshold γ are 0.97.

[0101] The results are as follows Figure 3 As shown in the figure, we can see that when the number of feature components reaches 70, the cumulative variance contribution rate reaches 0.97002, which is greater than the set threshold of 0.97. When the number of features increases from 70 to 80, although the variance contribution rate still increases, the difference in the cumulative variance contribution rate is less than 0.004; therefore, 70 is selected as the optimal feature number for extremely sparse Shum.

[0102] Figure 4 The reconstruction model is shown when the Shum data is missing 90% and the number of features is 140. The original image is as follows Figure 4 As shown in (a), when 90% of the data is missing, the effect is Figure 4 As shown in (b), the interpolation effect when R = 140 is as follows Figure 4 As shown in (c), it can be seen that the structural information of the reconstructed model is similar to the original data. Its accuracy evaluation is shown in Table 3. Its relative error is 0.1482, indicating that this method has a high accuracy in interpolation of extremely sparse atmospheric field models. In addition, the minimum and maximum values ​​of the interpolated results are 25.2875 and 19.7415 smaller than the original data, respectively, and have the same direction of deviation relative to the original data; but the median and average values ​​are almost the same as the original data, and the standard deviation is 2.2687 smaller than the original data, indicating that the interpolated data is smoother than the original data; the correlation coefficient of the interpolation result is only 0.0293 lower than the original data, which further proves the feasibility of this method in reconstructing extremely sparse atmospheric data.

[0103] Table 3 Statistics of Shum’s feature estimation results with different sparsity

[0104] Sparsity 0% 90% Relative error 0.0000 0.1482 R 0.0000 140.0000 Minimum 523.6875 498.4000 Maximum 596.0035 576.2620 Median 99.6585 99.7022 average value 80.8977 99.7022 Standard Deviation 107.2150 104.9463 Correlation coefficient 1.3253 1.2960

[0105] To verify the advantages of this method for feature extraction of extremely sparsely sampled high-order specific humidity data, we compared the performance of this method with the classic interpolation method, ordinary kriging for spatiotemporal data (KrigingST), which is widely used for feature estimation of spatiotemporal sparse data. The same data (Shum with 90% missing data) was selected as input data, and the number of features R was also selected as 70.

[0106] In the space-time kriging method, it assumes second-order (or intrinsic) stationarity in space and time and estimates the variogram. The parameters required in this method are complex, especially for the calculation and fitting of the variogram involved in model selection and specific parameter settings, which are usually selected empirically. In order to detect the impact of changes in the function model on the feature estimation results, different fitting models were selected, such as spherical (sph) and exponential (exp), and all other parameters were set to the same.

[0107] The comparison indicators are shown in Table 4.

[0108] Table 4 Comparison between KrigingST and this method

[0109]

[0110] It can be seen that the relative error ratio of this method is the lowest, which is 0.1482, and the correlation coefficient is the highest, which is 1.2960, indicating that this method has relatively high accuracy. Although the running time shows that the complexity of this method is much higher than that of space-time kriging, when the feature number R decreases, the running time is greatly reduced when the accuracy decreases is not very obvious. In addition, the standard deviation of space-time kriging is lower, which means that this method is more robust. However, due to the relatively complex parameters of this method, the interpolation results are easily affected by different parameters, which makes it difficult to determine the optimal parameters when the accuracy is given. However, the method proposed in this method has only one parameter, namely the feature number R, which can derive the optimal solution under the given constraints.

Claims

1. A high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account temporal and spatial heterogeneity, characterized by: The following steps are involved: Step 1: Reorganize the original high-dimensional specific humidity data into multiple sub-tensor blocks Recorded as the original sub-tensor block data; Step 2: Block each sub-tensor Y obtained in step 1 i As a whole, for each sub-tensor block Y i Performing operations S100 to S400 in sequence to complete the reconstruction of high-dimensional specific humidity data; S100: Based on the distribution structure of specific humidity data in the atmosphere, cluster the sub-tensor block data, select representative samples, and construct a mask tensor based on them; S200: based on the distribution structure of specific humidity data, extract the characteristic components of each sub-tensor block data from different scales; S300: performing a significance test on the feature component sets at different scales extracted in S200, screening the significant components and performing feature reconstruction to obtain feature reconstruction data; S400: Combining the mask tensor, performing weighted summation on the feature reconstruction data of each sub-tensor block data and the original sub-tensor block data.

2. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account the temporal and spatial heterogeneity according to claim 1 is characterized by: The step 1 specifically includes: Step 1.1: Organize the original high-dimensional specific humidity data into a high-order tensor, and divide the high-order tensor into multiple sub-tensor blocks with the same dimension size; Step 1.2: Aggregate locally similar sub-tensor blocks through spatiotemporal heterogeneity measures to form locally similar sub-tensor blocks.

3. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account the temporal and spatial heterogeneity according to claim 2 is characterized by: The step 1.1 specifically includes: Organize the raw high-dimensional specific humidity data into high-order tensors; Partition high-order tensors based on attribute values; The high-order tensor is split according to the spatial and temporal dimensions; the temporal dimension is split according to the data update interval, and the spatial dimension is split into regular blocks of the same size.

4. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account the temporal and spatial heterogeneity according to claim 2 is characterized by: The step 1.2 specifically includes: Assume that the decomposed sub-tensor blocks For adjacent sub-tensor block data, high-dimensional similarity calculation is performed based on the spatiotemporal heterogeneity measure to obtain the similarity between adjacent sub-tensor block data: ρ i,i+1 =SC(X i ,X i+1 ) (2) The similarity ρ between adjacent sub-tensor block data i,i+1 Compare with the given similarity threshold δ: If ρ i,i+1 >δ, then Otherwise X i =Y i , X i+1 =Y i+1 .

5. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account temporal and spatial heterogeneity according to claim 1 is characterized by: S100 specifically includes: Combined with the distribution structure of specific humidity data in the atmosphere [min, mid], (mid, max], for the sub-tensor block Y i , the sub-tensor block Y i The data on (Y i ) pjk , where p, j, k represent the sub-tensor block Y i The indexes in the three dimensions of longitude, latitude and time are assigned different weights according to their different structures in the distribution of the entire atmospheric value: And α+β=1; Construct the sub-tensor block Y in this way i A set of weight coefficients for the data; For the weight tensor Weight((Y i ) pjk ), construct the mask tensor Among them, (Q i ) pjk Denotes the mask tensor Q′ p The value at position (p, j, k).

6. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account the temporal and spatial heterogeneity according to claim 1, characterized in that: S200 specifically includes: According to the sub-tensor block And the mask tensor, perform weighted tensor CP decomposition to obtain: in, for The rth feature in longitude, latitude and time dimensions respectively, For data information that is not captured, is the weight coefficient set of the rth feature; * represents the scalar multiplication of the elements at the corresponding positions of the two tensors, Q i is the mask tensor.

7. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account temporal and spatial heterogeneity according to claim 1, characterized in that: S300 specifically includes: According to formula (6), calculate the sub-tensor block Y i The rth group of eigencomponents on In the set of all feature components The cumulative variance contribution rate of is: In the formula, i=1, 2, ..., M is a set of weight coefficients of each sub-tensor block data; Based on the calculated cumulative variance contribution rate and the given cumulative variance contribution rate threshold γ, only when CVC (i,r) >γ, the subtensor block Y i The rth group of characteristic components is the significant component; Get the sub-tensor block Y i The set of significant components after the above screening is The eigenvalue set is denoted as The significant components are combined and the eigenvalue set Perform tensor reconstruction to obtain feature reconstruction data 8. The high-dimensional specific humidity data reconstruction method based on extremely sparse sampling taking into account temporal and spatial heterogeneity according to claim 1, characterized in that: S400 specifically includes: Based on the following formula, the feature reconstruction data of the sub-tensor block data and the original sub-tensor block data are weighted and summed respectively: Restructure the data for the feature.

Citation Information

Patent Citations

  • Structured sparsity-based compression tensor acquisition and reconstruction system

    CN105721869A

  • Short-term traffic forecasting method based on multi-task multi-view learning model

    CN109598939A