An interpolation method for sparse and biased data
By building two adjacency matrices of different information transmission rules, training the initial model and the intermediate model, the problem of low accuracy when dealing with sparse and biased data is solved, and a high-precision three-dimensional distribution description of polluted areas is realized.
Patent Information
- Application Number
- CN202510220715.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-27
AI Technical Summary
When existing interpolation methods process sparse and biased data, it is difficult to accurately describe the three-dimensional distribution of polluted areas, resulting in low interpolation accuracy.
By constructing two adjacency matrices of different information transmission rules, training the initial model and the intermediate model respectively, dynamically capturing the concentration information of the sample points, and finally obtaining a high-precision prediction model.
It improves the interpolation accuracy and can more accurately describe the three-dimensional distribution of polluted areas, which is suitable for processing sparse and biased data.
Smart Images

Figure CN119719758B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an interpolation method for sparse and biased data, belonging to the technical field of data processing. Background Art
[0002] Currently, the long-term enrichment of waste discharged from high-energy-consuming and highly polluting industrial equipment in the soil has caused serious soil pollution. It is of great practical significance to clarify the spatial distribution law of typical pollutants in the site and accurately delimit the pollution boundary and scope.
[0003] Due to exploration costs and land use restrictions, the actually available drilling data often have sparsity and bias, resulting in difficulties for traditional methods to accurately describe the three-dimensional distribution of the polluted area. Currently, traditional spatial interpolation methods mainly include design-based (such as IDW) and model-based (such as Kriging) statistical methods. However, IDW depends on a specific sampling design, and the accuracy of its interpolation results is easily affected by the sample size and soil sample design. The Kriging model describes the spatial autocorrelation between two sampling points by establishing a variogram and then predicts unknown points. However, the smoothing effect of this method often fails to capture the pollutant changes at sparse borehole positions smaller than the sampling distance. In addition, affected by the location of the pollution source, the properties of the pollutants themselves, soil characteristics (such as soil texture, porosity, permeability coefficient, etc.) and human activities (such as production processes, functional area distribution, pollution leakage), it is difficult to construct a suitable semi-variogram, ultimately resulting in a reduction in the accuracy of the interpolation results. Summary of the Invention
[0004] The present invention provides an interpolation method for sparse and biased data, which can solve the problem of low interpolation accuracy of existing interpolation methods.
[0005] The present invention provides an interpolation method for sparse and biased data, and the method includes:
[0006] S1. Construct a first adjacency matrix and a second adjacency matrix between points according to the sampling information of sampling points in the research area; the information transfer rules of the first adjacency matrix and the second adjacency matrix are different;
[0007] S2. Construct an initial model, train the initial model using the sampling information of the sampling points and the first adjacency matrix to obtain an intermediate model, and train the intermediate model using the sampling information of the sampling points and the second adjacency matrix to obtain a prediction model;
[0008] S3. Input the coordinates of the points to be interpolated in the research area and the sampling information of the sampling points into the prediction model to obtain the predicted values of the points to be interpolated.
[0009] Optionally, the sampling points are divided into training nodes and test nodes;
[0010] The information transfer rule of the first adjacency matrix is as follows: Only information transfer between training nodes and information transfer to test nodes are allowed; test nodes are only allowed to receive information from training nodes.
[0011] The information transfer rule of the second adjacency matrix is as follows: Information transfer between training nodes is allowed, and information transfer to test nodes and themselves is also allowed; test nodes are allowed to receive information from training nodes, and information transfer to training nodes and themselves is also allowed.
[0012] Optionally, the S1 specifically includes:
[0013] Construct a weighted adjacency matrix between points according to the sampling information of the sampling points in the research area;
[0014] Divide the weighted adjacency matrix into a first adjacency matrix and a second adjacency matrix according to different information transfer rules.
[0015] Optionally, the construction of the weighted adjacency matrix between points according to the sampling information of the sampling points in the research area is specifically:
[0016] Construct a distance matrix between points according to the sampling information of the sampling points in the research area, and convert the distance matrix into a weighted adjacency matrix.
[0017] Optionally, the construction of the initial model in the S2 specifically includes:
[0018] Determine an initial weight matrix and an initial edge mask matrix according to the sampling information of the sampling points;
[0019] Construct an initial model according to the initial weight matrix and the initial edge mask matrix.
[0020] Optionally, the training of the initial model with the sampling information of the sampling points and the first adjacency matrix in the S2 to obtain an intermediate model specifically includes:
[0021] Perform function activation processing and symmetrization processing on the initial edge mask matrix to obtain an updated edge mask matrix;
[0022] Determine an updated weight matrix according to the sampling information of the sampling points, the updated edge mask matrix, the first adjacency matrix, and the initial weight matrix; the updated edge mask matrix and the updated weight matrix form an intermediate model.
[0023] Optionally, determining the updated weight matrix according to the sampling information of the sampling points, the updated edge mask matrix, the first adjacency matrix, and the initial weight matrix specifically includes:
[0024] Multiply the updated edge mask matrix by the first adjacency matrix to obtain a first masked adjacency matrix;
[0025] Multiply the first masked adjacency matrix by the initial weight matrix, and multiply the multiplication result by the information matrix of the sampling points to obtain an updated weight matrix.
[0026] Optionally, in S2, training the intermediate model using the sampling information of the sampling points and the second adjacency matrix to obtain a prediction model, specifically including:
[0027] Multiply the updated edge mask matrix by the second adjacency matrix to obtain a second masked adjacency matrix;
[0028] Multiply the second masked adjacency matrix by the updated weight matrix, and multiply the multiplication result by the information matrix of the sampling points to obtain a trained weight matrix; the updated edge mask matrix and the trained weight matrix form a prediction model.
[0029] Optionally, after S3, the method further includes:
[0030] S4. Input the coordinates of the sampling points in the research area into the prediction model to obtain the predicted values of the sampling points, and determine the residual values of the sampling points according to the sampling information and predicted values of the sampling points;
[0031] S5. Use the residual values to correct the predicted values of the points to be interpolated.
[0032] Optionally, S5 specifically includes:
[0033] Determine the residual predicted values of the points to be interpolated according to the residual values;
[0034] Use the residual predicted values to correct the predicted values of the points to be interpolated to obtain the corrected predicted results of the points to be interpolated.
[0035] The beneficial effects that the present invention can produce include:
[0036] The interpolation method for sparse and biased data provided by the present invention constructs two adjacency matrices with different information transfer rules through the spatial autocorrelation of sampling points, and sequentially trains the model based on these two adjacency matrices to dynamically capture the concentration information of sample points. The trained prediction model has high interpolation accuracy. Description of the Drawings
[0037] Figure 1 It is a flowchart of the interpolation method for sparse and biased data provided by an embodiment of the present invention;
[0038] Figure 2Schematic diagram of information transfer of nodes in the first convolutional layer provided by an embodiment of the present invention during the training process;
[0039] Figure 3 Schematic diagram of information transfer of nodes in the first convolutional layer provided by an embodiment of the present invention during the testing process;
[0040] Figure 4 First adjacency matrix of nodes in the first convolutional layer provided by an embodiment of the present invention during the training process;
[0041] Figure 5 Schematic diagram of information transfer of nodes in the second convolutional layer provided by an embodiment of the present invention during the training process;
[0042] Figure 6 Schematic diagram of information transfer of nodes in the second convolutional layer provided by an embodiment of the present invention during the testing process;
[0043] Figure 7 Second adjacency matrix of nodes in the second convolutional layer provided by an embodiment of the present invention during the training process;
[0044] Figure 8 Schematic diagram of the creation process of the three-dimensional grid of interpolation points provided by an embodiment of the present invention. Detailed implementation manners
[0045] The present invention will be described in detail below with reference to embodiments, but the present invention is not limited to these embodiments.
[0046] An embodiment of the present invention provides an interpolation method for sparse biased data, as Figure 1 shown, the method includes:
[0047] S1. Construct a first adjacency matrix and a second adjacency matrix between points according to the sampling information of the sampling points in the research area; the information transfer rules of the first adjacency matrix and the second adjacency matrix are different.
[0048] Specifically include:
[0049] (1) Construct a weighted adjacency matrix between points according to the sampling information of the sampling points in the research area.
[0050] Specifically, construct a distance matrix between points according to the sampling information of the sampling points in the research area, and convert the distance matrix into a weighted adjacency matrix.
[0051] In practical applications, the weighted adjacency matrix can be constructed according to the following steps.
[0052] Step 1: Determination of the research area. Based on the grid sampling method (e.g., grid interval, length × width is 3×3 m or length × width × height is 3×3×3 m), the spatial positions of the sampling points are arranged (the point positions are at the center of the grid), and the sample collection and analysis are carried out strictly in accordance with the requirements of "Technical Specifications for Soil Environmental Monitoring (HJ / T 166-2004)".
[0053] Step 2: Sampling and collection of pollutant data (pollution elements such as arsenic, beryllium, cadmium, chromium, copper, lead, mercury, nickel, zinc, tin, cyanide, fluoride, polycyclic aromatic hydrocarbons, etc.; point coordinates such as X, Y, Z; concentration values; ID numbers and other relevant data), forming the sampling information of the sampling points in the research area.
[0054] Step 3: Construction of the weighted adjacency matrix. Calculate the distance between any two point positions, and construct a distance matrix between the point positions. Convert this distance matrix into a weighted adjacency matrix, where the weight is the square of the reciprocal of the distance, and the weight of a point position to itself is "0". Example: The distance matrix is , calculate the square of the reciprocal of the values in the distance matrix, that is is 0.25, and then the weighted adjacency matrix obtained is . Corresponding number data can also be obtained as .
[0055] (2) According to different information transfer rules, divide the weighted adjacency matrix into the first adjacency matrix and the second adjacency matrix.
[0056] Among them, the sampling points are divided into training nodes and test nodes;
[0057] The information transfer rule of the first adjacency matrix is: Only information transfer between training nodes and information transfer from training nodes to test nodes are allowed; Test nodes are only allowed to receive information from training nodes;
[0058] The information transfer rule of the second adjacency matrix is: Information transfer between training nodes is allowed, and information transfer to test nodes and to themselves is also allowed; Test nodes are allowed to receive information from training nodes, and information transfer to training nodes and to themselves is also allowed.
[0059] Specifically, for the information transfer rule of the first adjacency matrix: The initial value of the training node comes from the actual sampling value, so the training node can transfer concentration information to adjacent nodes (including training nodes and test nodes). However, self-loops will weaken the information transferred by other training nodes, so information should not be transferred to itself. In addition, the initial value of the test node is set to 0, which is an assumed value. Therefore, the test node can only receive concentration information from adjacent nodes (training nodes), cannot transfer concentration information to adjacent nodes (including training nodes and test nodes), nor can it transfer to itself.
[0060] The second adjacent matrix information transfer rule: The initial values of both the training nodes and the test nodes come from the aggregated values obtained after the previous layer's training. Therefore, both the training nodes and the test nodes can transfer or receive the concentration information of adjacent nodes, and can also transfer the concentration information to themselves, so as to ensure that the aggregated values from the previous layer are not ignored.
[0061] Reference Figures 2 to 7 As shown, for example, the training nodes are: X1, X2, X4, X5, X6. The test nodes are: X3, X7. Figure 2 and Figure 3 In, the values of X1, X2, X4, X5, X6 are taken as the real observed values. The initial values of X3 and X7 are 0. Figure 5 and Figure 6 represent the aggregated values obtained after the previous layer's training. The lines with arrows represent the information flow. Figure 2 and Figure 5 represent the training process. Figure 3 and Figure 6 represent the test process. Figure 4 and Figure 7 represent the change of the weights of the weighted adjacent matrix. Figure 2 , Figure 3 and Figure 4 represent the information transfer of the previous layer (i.e., the first convolutional layer). Figure 5 , Figure 6 and Figure 7 represent the information transfer of the current layer (i.e., the second convolutional layer).
[0062] S2. Construct the initial model, train the initial model using the sampling information of the sampling points and the first adjacent matrix to obtain an intermediate model, and then train the intermediate model using the sampling information of the sampling points and the second adjacent matrix to obtain a prediction model.
[0063] The construction of the initial model in S2 specifically includes:
[0064] First, determine the initial weight matrix and the initial edge mask matrix according to the sampling information of the sampling points; then construct the initial model according to the initial weight matrix and the initial edge mask matrix.
[0065] Construct the network structure of the initial model: an adaptive layer (i.e., the mask layer), and the mask layer is composed of a weight matrix and an edge mask matrix. Both the weight matrix and the edge mask matrix are created by the torch.nn package in Python. The dimension of the former is determined by the sampling information of the input sampling points (i.e., the concentration information) and the set output dimension, and the dimension of the latter is determined by the input concentration information. Specifically, initialize the weight matrix and the edge mask matrix in the network structure according to the sampling point concentration information and the node dimension to obtain the initial weight matrix and the initial edge mask matrix, and both are normalized by the torch.nn package in Python.
[0066] Then, set the hyperparameters of the initial model. The hyperparameters include Epoch, learning rate, weight decay, node dimension, etc.
[0067] In S2, use the sampling information of the sampling points and the first adjacency matrix to train the initial model to obtain an intermediate model, which specifically includes:
[0068] (1) Perform function activation processing and symmetrization processing on the initial edge mask matrix to obtain an updated edge mask matrix.
[0069] (2) Determine the updated weight matrix according to the sampling information of the sampling points, the updated edge mask matrix, the first adjacency matrix, and the initial weight matrix; the updated edge mask matrix and the updated weight matrix form the intermediate model.
[0070] Specifically: First, multiply the updated edge mask matrix by the first adjacency matrix to obtain the first masked adjacency matrix; then multiply the first masked adjacency matrix by the initial weight matrix, and multiply the multiplication result by the information matrix of the sampling points to obtain the updated weight matrix.
[0071] In practical applications, use the updated edge mask matrix to mask the first adjacency matrix, that is, multiply the updated edge mask matrix by the first adjacency matrix to obtain the first masked adjacency matrix; then the first masked adjacency matrix can be normalized, and the normalized first masked adjacency matrix is multiplied by the initial weight matrix, and the multiplication result is multiplied by the information matrix of the sampling points to obtain the updated weight matrix.
[0072] It should be noted that the above steps can be looped according to the set number of mask layers, and the updated weight matrix obtained after each loop is saved. The multiple updated weight matrices obtained by looping are regularized using the F.dropout function in the PyTorch framework to obtain the updated weight matrix of the final first convolutional layer.
[0073] In S2, use the sampling information of the sampling points and the second adjacency matrix to train the intermediate model to obtain a prediction model, which specifically includes:
[0074] First, multiply the updated edge mask matrix by the second adjacency matrix to obtain the second masked adjacency matrix; then multiply the second masked adjacency matrix by the updated weight matrix, and multiply the multiplication result by the information matrix of the sampling points to obtain the trained weight matrix, that is, the weight matrix of the second convolutional layer; the updated edge mask matrix and the trained weight matrix form the prediction model.
[0075] In practical applications, the intermediate model and the prediction model can be cyclically trained according to specific Epoch values to obtain the final prediction model. After training, calculate the loss and accuracy results of the prediction model and draw the corresponding curves. Use the leave-one-out cross-validation method to verify the results of model training, calculate the loss and accuracy results of the prediction model, and draw the corresponding curves. When the model training and verification results conform to the actual situation, the model can be used for prediction. When the results of model training and verification are not satisfactory, reset the hyperparameters and retrain the model.
[0076] It should be noted that, according to the depth value (Z) of the sampling points, the sampling points can be divided into multiple data sets (for example, when Z = 14m, it is divided into 14 data sets, such as: 0 - 1m is the first sampling data set, 1 - 2m is the second sampling data set,..., 13 - 14m is the 14th sampling data set). According to the leave-one-out verification method, mark the training and verification data points for each data set, and perform normalization processing on the sampling data. Then, perform model training in sequence according to the divided data to obtain the prediction model corresponding to the depth value.
[0077] S3. Input the coordinates of the points to be interpolated and the sampling information of the sampling points in the research area into the prediction model to obtain the predicted values of the points to be interpolated.
[0078] Reference Figure 8 As shown, based on the coordinate range of the sampling points, create three-dimensional grid point coordinates (i.e., the coordinates of the points to be interpolated), set the initial value "0" and the corresponding number for the coordinates of the points to be interpolated. Figure 8 In the depth range of 0 - 1m, V1, V2, V3, V4, and V5 are sampling points; V6 and V7 are points to be interpolated.
[0079] Input the sampling information of the sampling points and the coordinates of the points to be interpolated into the trained prediction model, and the predicted values of the points to be interpolated can be obtained, that is, the predicted concentration values of the points to be interpolated.
[0080] After S3, the method further includes:
[0081] S4. Input the coordinates of the sampling points in the research area into the prediction model to obtain the predicted values of the sampling points, and determine the residual values of the sampling points according to the sampling information and the predicted values of the sampling points.
[0082] Input the sampling point data into the trained prediction model to obtain the predicted values of the sampling point concentrations. Then, calculate the residual values of the sampling points according to the true values and the predicted values of the sampling points.
[0083] S5. Use the residual values to correct the predicted values of the points to be interpolated.
[0084] Specifically, it includes:
[0085] First, determine the residual prediction value of the point to be interpolated according to the residual value; then use the residual prediction value to correct the prediction value of the point to be interpolated to obtain the corrected prediction result of the point to be interpolated.
[0086] The residual values of the sampling points can be input into the trained prediction model to obtain the residual prediction value of the point to be interpolated. They can also be input into Kriging or other interpolation methods to obtain the residual prediction value of the point to be interpolated. The embodiments of the present invention do not limit this.
[0087] Add the residual prediction value of the point to be interpolated and the prediction value of the point to be interpolated obtained previously to obtain the corrected prediction result of the point to be interpolated.
[0088] The advantages of the present invention are mainly as follows:
[0089] Advantage 1: The dual-mode adaptive learning sample point feature mechanism deeply mines the feature information of sparse and biased data. First, the transfer rule of sample point information is optimized, different adjacency matrices are constructed based on spatial autocorrelation, and the sample point concentration information is dynamically captured. Secondly, the sample point structure is dynamically learned, a dynamic mask matrix and a full-parameter learning method are constructed, and the weight value of the "edge" (the non-linear relationship between sample points) is dynamically optimized in a data-driven manner, making full use of the limited sample point spatial distribution data to adapt to the characteristic that the spatial structure of pollutants is dynamically changing.
[0090] Advantage 2: The residual mechanism improves the expansibility of the model. Using the prediction value of the point to be interpolated and the residual value result of the sampling point obtained by the prediction model of the present invention, interpolation is performed again based on the residual value result of the sampling point, and finally, combined with the prediction result obtained by the first interpolation, it is the final interpolation prediction result. The present invention uses the residual mechanism to reduce the influence of outliers on the model and can combine with the results of other interpolation methods to obtain the final prediction result, improving the accuracy of the final prediction result.
[0091] Advantage 3: Without auxiliary variable information, only using the three-dimensional point position and the pollutant concentration information of the sampling points, a high-precision three-dimensional spatial distribution pattern of pollutants can be achieved.
[0092] As described above, these are only several embodiments of the present application, and do not impose any form of limitation on the present application. Although the present application is disclosed with preferred embodiments as above, it is not intended to limit the present application. Any person skilled in the art, without departing from the scope of the technical solution of the present application, making some changes or modifications using the disclosed technical content is equivalent to equivalent implementation cases and all belong to the scope of the technical solution.
Claims
1. A sparse biased data interpolation method for interpolation prediction of spatial distribution of soil pollutants, characterized in that: The method comprises: S1. Construct a distance matrix between points according to the sampling information of the sampling points in the study area, and convert the distance matrix into a weighted adjacency matrix; divide the weighted adjacency matrix into a first adjacency matrix and a second adjacency matrix according to different information transmission rules; wherein the sampling information of the sampling points includes the coordinates of the sampling points and the sampling concentration values of the pollutant elements at the sampling points; the sampling points are divided into training nodes and test nodes; the information transmission rule of the first adjacency matrix is: training nodes are only allowed to transmit concentration information to each other and to test nodes; test nodes are only allowed to receive concentration information from training nodes; the information transmission rule of the second adjacency matrix is: training nodes are allowed to transmit concentration information to each other, and to test nodes and themselves; test nodes are allowed to receive concentration information from training nodes, and to transmit concentration information to training nodes and themselves; S2, determining an initial weight matrix and an initial edge mask matrix according to sampling information of the sampling points, and constructing an initial model according to the initial weight matrix and the initial edge mask matrix, training the initial model using the sampling information of the sampling points and the first adjacency matrix to obtain an intermediate model, and training the intermediate model using the sampling information of the sampling points and the second adjacency matrix to obtain a prediction model; S3. Input the coordinates of the to-be-interpolated point and the sampling information of the sampling point in the study area into the prediction model to obtain the predicted value of the to-be-interpolated point; wherein the predicted value of the to-be-interpolated point is the predicted concentration value of the pollutant element at the to-be-interpolated point.
2. The method according to claim 1, characterized in that The step S2 uses the sampling information of the sampling points and the first adjacency matrix to train the initial model to obtain an intermediate model, which specifically includes: Performing function activation processing and symmetrization processing on the initial edge mask matrix to obtain an updated edge mask matrix; An updated weight matrix is determined according to sampling information of sampling points, the updated edge mask matrix, the first adjacency matrix and the initial weight matrix; the updated edge mask matrix and the updated weight matrix constitute an intermediate model.
3. The method according to claim 2, characterized in that Determining an updated weight matrix according to the sampling information of the sampling points, the updated edge mask matrix, the first adjacency matrix and the initial weight matrix specifically includes: Multiplying the updated edge mask matrix by the first adjacency matrix to obtain a first mask adjacency matrix; The first mask adjacency matrix is multiplied by the initial weight matrix, and the multiplication result is multiplied by the information matrix of the sampling points to obtain an updated weight matrix.
4. The method according to claim 2, characterized in that: The step S2 uses the sampling information of the sampling points and the second adjacency matrix to train the intermediate model to obtain a prediction model, which specifically includes: Multiplying the updated edge mask matrix by the second adjacency matrix to obtain a second mask adjacency matrix; The second mask adjacency matrix is multiplied by the updated weight matrix, and the multiplication result is multiplied by the information matrix of the sampling points to obtain a trained weight matrix; the updated edge mask matrix and the trained weight matrix constitute a prediction model.
5. The method according to claim 1, characterized in that After S3, the method further includes: S4, inputting the coordinates of the sampling points in the study area into the prediction model to obtain the predicted values of the sampling points, and determining the residual values of the sampling points according to the sampling information and the predicted values of the sampling points; S5. Correct the predicted value of the point to be interpolated using the residual value.
6. The method according to claim 5, characterized in that The S5 specifically includes: Determine the residual prediction value of the to-be-interpolated point according to the residual value; The predicted value of the point to be interpolated is corrected using the residual predicted value to obtain a corrected prediction result of the point to be interpolated.
Citation Information
Patent Citations
Scattered data interpolation model training method, interpolation method and device
CN115983370A
Soil heavy metal concentration prediction method based on distance relation graph convolutional network
CN118887030A