Radioactive geophysical exploration data processing method and system based on physical field attenuation model
Through data processing methods based on the physical field attenuation model, including abnormal point extraction, diffusion impact model construction, Kriging interpolation and deep learning smoothing processing, the problem of image parts distortion caused by insufficient fitting of mutation points and data inhomogeneity in the prior art is solved, and a more accurate and intuitive radioactive geophysical data processing effect is achieved.
Patent Information
- Application Number
- CN202510228118.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing radioactive physicological data processing methods have poor fitting to the mutation points in the data during gridding, resulting in the high-value anomalies being removed and cannot reflect the true situation of the anomalies. At the same time, the generated image pieces are distorted due to the unevenness of the point data, which cannot reflect the distribution characteristics of the radioactive field well.
The data processing method based on the physical field attenuation model is adopted to determine the abnormal points, build the diffusion impact model of the abnormal points, use the Kriging interpolation method to generate a smooth basic field, and superimpose the abnormal point diffusion field to the normal data field to form a comprehensive field. Finally, the integrated field is smoothed using a deep learning-based contour smoothing model.
The accurate extraction and natural attenuation of abnormal points are achieved, and the problem of abnormal points being ignored or excessive smoothing is avoided. The generated contour map is more continuous and intuitive, and can effectively reflect the distribution characteristics of the radioactive field.
Smart Images

Figure CN119738889B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radioactive geophysical exploration data processing, and in particular relates to a radioactive geophysical exploration data processing method and system based on a physical field attenuation model. Background Art
[0002] In the late data collection of the radioactive geophysical survey project currently being carried out, it is necessary to grid the point data on the survey line and then generate a contour map, which is then submitted as a result map after processing. Several mainstream data processing methods such as distance weighted average method, trend surface fitting method, Kriging method, least squares method, etc. are not ideal and have the following problems:
[0003] 1. Although some methods have a high overall fit when gridding, they have a poor fit for the mutation points in the data. All algorithms have a certain degree of smoothness, which causes the high-value anomaly parts to be eliminated when gridding, and cannot reflect the true situation of the anomaly;
[0004] 2. Due to the unevenness of point data collection and the inherent characteristics of various interpolation methods, the generated maps have varying degrees of distortion and cannot well reflect the distribution characteristics of the radioactive field. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a radioactive geophysical exploration data processing method and system based on a physical field attenuation model to solve the problems existing in the above prior art.
[0006] In a first aspect, to achieve the above-mentioned purpose, the present invention provides a method for processing radioactive geophysical exploration data based on a physical field attenuation model, comprising the following steps:
[0007] Identify anomalies in radioactive geophysical data;
[0008] The inverse distance weighted method is used to construct the diffusion impact model of outliers;
[0009] Use Kriging interpolation method to generate smoothed base fields of normal data;
[0010] The abnormal point diffusion field is superimposed on the normal data field to form a comprehensive field.
[0011] The synthetic field is smoothed using a deep learning-based contour smoothing model.
[0012] Optionally, the process of determining anomalies in radioactive geophysical data includes:
[0013] By calculating the mean and standard deviation of radioactive geophysical data, the points beyond the range of 3 times the standard deviation of the mean are regarded as outliers;
[0014] The coordinates and field value information of the outliers are extracted and stored as separate datasets.
[0015] Optionally, the process of constructing a diffusion impact model of anomalies using an inverse distance weighted method includes:
[0016] Use the field value of the outlier point as the diffusion source;
[0017] Calculating a weight according to the distance between the outlier point and the target point, wherein the weight is represented by an inverse distance weighted function;
[0018] The diffusion influences of multiple outliers are linearly superimposed to generate the diffusion field values of the outliers, and then the diffusion influence model of the outliers is obtained.
[0019] Optionally, the process of generating a smoothed base field of normal data using the kriging interpolation method includes:
[0020] Perform spatial correlation analysis on normal data fields and construct semivariogram model;
[0021] Based on the semivariogram model, calculate the interpolation weights;
[0022] Generates a smooth basis field of the normal data field using interpolation weights.
[0023] Optionally, the process of superimposing the abnormal point diffusion field onto the normal data field to form a comprehensive field includes:
[0024] Superimpose the diffusion field value of the abnormal point with the normal data field value point by point;
[0025] The superimposed field values are smoothed to form a comprehensive field.
[0026] Optionally, the process of smoothing the synthetic field using a contour smoothing model based on deep learning includes:
[0027] Build a deep learning model for contour smoothing;
[0028] Use the contour data of the comprehensive field as training data to train the deep learning model;
[0029] The trained deep learning model is used to smooth the contour lines of the comprehensive field.
[0030] In a second aspect, the present invention further provides a radioactive geophysical exploration data processing system based on a physical field attenuation model, which is used to implement a radioactive geophysical exploration data processing method based on a physical field attenuation model, and the system comprises:
[0031] Anomaly identification module, used to analyze radioactive geophysical exploration data and extract anomalies;
[0032] A model building module for generating a diffusion impact model of outliers based on the inverse distance weighted method;
[0033] Interpolation module, used to generate smoothed base fields of normal data using Kriging interpolation method;
[0034] Data fusion module, used to superimpose the abnormal point diffusion field onto the normal data field to form a comprehensive field;
[0035] The smoothing module is used to smooth the contour lines of the comprehensive field based on the deep learning model.
[0036] Optionally, the model building module includes a weight calculation unit, and the weight calculation unit is used to calculate the diffusion weight according to the distance between the abnormal point and the target point.
[0037] In a third aspect, the present invention further provides a computer terminal device, comprising:
[0038] one or more processors;
[0039] A memory, coupled to the processor, for storing one or more programs;
[0040] When the one or more programs are executed by the one or more processors, the one or more processors implement a radioactive geophysical exploration data processing method based on a physical field attenuation model.
[0041] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements a radioactive geophysical exploration data processing method based on a physical field attenuation model.
[0042] Compared with the prior art, the present invention has the following advantages and technical effects:
[0043] The present invention provides a radioactive geophysical data processing method and system based on a physical field attenuation model. First, the three-fold mean standard is used to quickly screen the abnormal points in the radioactive geophysical data, and the coordinates and field value information of the abnormal points are extracted to ensure the accuracy of the abnormal point extraction. Then, the diffusion influence model of the abnormal point is constructed by the inverse distance weighted method, and the diffusion field value of the abnormal point is generated by using the field value of the abnormal point and its distance weight relationship, so as to realize the spatial natural attenuation of the influence of the abnormal point, and avoid the problem of the abnormal point being ignored or over-smoothed in the traditional method. Further, the smoothed basic field of normal data is generated by combining the Kriging interpolation method, and more scientific and realistic basic field value data is generated by modeling the spatial correlation of the data and accurately calculating the interpolation weight. Subsequently, the diffusion field of the abnormal point is superimposed on the normal field point by point, and the characteristics of the abnormal point are retained by smoothing to form a comprehensive field, which solves the interpolation distortion problem caused by uneven point data. Finally, the contour line smoothing model based on deep learning is used to smooth and optimize the contour lines generated by the comprehensive field data, which effectively improves the continuity and visual effect of the contour lines, while retaining the local detail characteristics of the data. Based on the above technical means, the present invention overcomes the limitations of traditional interpolation methods in outlier processing and contour line generation, and provides an innovative solution for the accurate processing and intuitive expression of radioactive geophysical exploration data.
[0044] The present invention uses the triple mean standard to quickly screen outliers to ensure accurate extraction of outliers; the inverse distance weighted method (IDW) is used to construct an outlier diffusion model so that the influence of outliers can be naturally attenuated in space and gradually integrated into the overall data field. The Kriging interpolation method is used to generate a smoothed basic field of normal data, making full use of the spatial correlation of the data to generate a more realistic and scientific interpolation result. After superimposing the outlier diffusion field onto the normal data field, the continuity of the field value is ensured by smoothing, while retaining the characteristics of the outliers, solving the problem of traditional interpolation methods ignoring or over-smoothing outliers. The contour smoothing model based on deep learning is used to smooth the comprehensive field, significantly reducing the jagged or abrupt phenomenon of contour lines, and generating a smoother and more intuitive contour map. This method overcomes the problem of improper handling of outliers by traditional interpolation methods (such as outliers being simply ignored or over-smoothed), while avoiding the interpolation distortion caused by uneven point distribution. In summary, the present invention not only retains the true characteristics of abnormal points in the data field, but also improves the smoothness of the overall data field through the steps of precise identification, scientific modeling, effective fusion and smooth optimization, providing an innovative solution for the accurate processing and analysis of radioactive geophysical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0046] Figure 1 It is the overall technical roadmap of the embodiment of the present invention;
[0047] Figure 2 This is a schematic diagram of determining outliers by using the 3-times mean method according to an embodiment of the present invention;
[0048] Figure 3 A schematic diagram showing a comparison between a predicted value and an actual value of Kriging interpolation of a normal value according to an embodiment of the present invention;
[0049] Figure 4 Schematic diagram of the distribution of prediction variance according to an embodiment of the present invention;
[0050] Figure 5 This is a schematic diagram of the results after being processed by a deep neural network model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0051] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0052] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0053] This paper proposes an anomaly influence diffusion and smoothing interpolation method based on the physical field decay model (Anomaly Influence Diffusion and Smoothing Interpolation Method Based on Physical Field Decay Model). The introduction and processing of anomalies is an important and complex step. In order to reasonably express the influence of anomalies on the overall field without destroying the characteristics of the normal data field, an anomaly influence diffusion and smoothing interpolation method based on the physical field decay model is proposed. The following is the overall framework and core idea of the method:
[0054] 1. Core idea:
[0055] The influence of outliers is seen as a diffusion phenomenon similar to that of physical fields. Outliers can be likened to a "source" that affects the surrounding data fields, similar to the effects of physical fields such as gravity fields and electric fields. The influence characteristics of outliers have the following characteristics:
[0056] The local impact is significant: areas closer to the outlier point are more affected.
[0057] Gradual attenuation over long distances: As the distance increases, the influence of outliers gradually weakens and eventually merges into the normal data field.
[0058] Linear superposition effect: The influence of multiple abnormal points can be superimposed, and this superposition effect conforms to the superposition principle of physical fields.
[0059] Therefore, the core of this method is to naturally incorporate the influence of outliers in the interpolation process through the outlier point spread function and field value fitting strategy, while maintaining the smoothness of the overall field and the significant characteristics of local outliers.
[0060] 2. Methods and steps:
[0061] (1) Define outliers and their characteristics:
[0062] a) Outlier determination: A simple and clear statistical method, the 3x mean standard, is used to quickly determine outliers. The coordinates and field values of the outliers are extracted and stored as separate data sets.
[0063] b) Characteristics of outliers: Outlier field values may have extremely high or low values, far exceeding the distribution range of normal data.
[0064] The field value of the outlier point will significantly affect the data field in the surrounding area, but the impact range is limited.
[0065] (2) Diffusion modeling of outlier impact:
[0066] a) Diffusion model based on physical fields: The impact of outliers on surrounding field values can be modeled through distance functions, similar to the effects of physical fields such as gravity and electric field strength.
[0067] Use the inverse distance weighted (IDW) diffusion function to calculate the impact of outliers on the field values of surrounding coordinates: or ;
[0068] Among them, d is the distance between the outlier point and the target point, p is the weight decay exponent, Controls the spread of the Gaussian distribution, is a smoothing parameter to prevent the denominator from being zero, The distance between the outlier point pair and its The field value of the normal point affects the size.
[0069] b) Linear superposition effect: For each target point, the diffusion effect of all abnormal points is superimposed:
[0070] ;
[0071] in, The coordinates after superimposing the diffusion effect of all abnormal points are The gamma value of the point is the field value weight of the outlier point, is the diffusion function of the i-th outlier point to the target point.
[0072] (3) Fusion of abnormal point impact and normal field:
[0073] a) Kriging interpolation of normal field:
[0074] Kriging interpolation is performed using a normal data field to produce a smooth base field.
[0075] Kriging interpolation can effectively capture the spatial distribution law of normal data and ensure the smoothness of the overall field.
[0076] b) Superposition of abnormal point fields:
[0077] The diffuse field of the outlier is superimposed on the normal field to form an overall field that includes the influence of the outlier:
[0078] ;
[0079] in, The coordinates are The field value of the normal point after being affected by the abnormal point, The coordinates are The field value of the normal point when it is not affected by the abnormal point field, The coordinates of the outlier point are The field value at the normal point affects the size.
[0080] c) Keep the original characteristics of the outliers:
[0081] During the superposition process, ensure that the original field values of the outliers are not smoothed out and retain their precise values.
[0082] (4) Data smoothing:
[0083] In order to optimize the effect of the contour lines after interpolation, the field data is further smoothed and the contour lines are post-processed through a deep learning network.
[0084] Verification of the overall algorithm effect. Use MSE to test the interpolation fit of the entire field. For the IDW-based anomaly point introduction model, grid traverse the parameters within a certain range to find the parameters with the smallest change in the abnormal area ratio (compared to the original data) as the final parameters. For the deep learning model, use Loss to control network learning. The overall technical roadmap of the algorithm can be found in Figure 1 .
[0085] The core innovation of this invention is that it proposes an interpolation method based on the principle of physical field attenuation based on the inverse distance weighted interpolation (IDW) outlier influence diffusion model and the contour smoothing network model based on deep learning. This method simulates the diffusion process of the influence of outliers on the surrounding areas by designing a specific influence function, so that the force of the outliers can gradually decay and smoothly integrate into the field value of normal data. Compared with traditional interpolation methods, the interpolation method based on the principle of physical field attenuation can handle outliers more naturally, avoiding sudden changes or over-smoothing of the strong influence of outliers on the surrounding areas.
[0086] 3.1 Limitations of traditional methods:
[0087] Traditional interpolation methods (such as linear interpolation, nearest neighbor interpolation, and polynomial interpolation) have limitations when dealing with outliers: linear interpolation smoothes outliers, nearest neighbor interpolation limits its influence, and polynomial interpolation may overfit. These methods usually ignore or remove outliers and are difficult to retain their true characteristics.
[0088] 3.2 Innovation of the outlier introduction model:
[0089] The present invention combines IDW interpolation with a deep learning model, giving the interpolation process more flexibility and accuracy. The IDW interpolation method gradually attenuates the influence of outliers through the relationship between weight and distance, and naturally considers the influence of outliers on surrounding areas during the interpolation process. The deep learning model smoothes the contour lines by learning the complex relationship between outliers and their surrounding data, further enhancing the ability to model the local influence of outliers on data field values.
[0090] Compared with traditional interpolation methods, the IDW and deep learning attenuation model of the present invention can more effectively handle the impact of outliers and avoid the information loss problem caused by simple smoothing. By treating outliers as "sources" and combining the physical field attenuation principle, not only the impact characteristics of outliers are retained, but also their impact can be smoothly transitioned to the surrounding area, thereby achieving a natural fusion of outliers and normal data field values. This innovative method provides a more accurate, flexible and actual physical phenomenon-compliant interpolation strategy, which is particularly suitable for processing data sets containing outliers.
[0091] The outlier point influence diffusion and smooth interpolation method based on the physical field attenuation model proposed in the present invention combines multiple classical theories and modern interpolation techniques, aiming to solve the impact of outliers in the data on the field value and the smoothing problem of contour images. In order to better understand the background and innovation of the present invention, this section will introduce the theories and methods related to the subject of the present invention. Based on the methods proposed by existing scholars, the application and limitations of these theoretical models in solving similar problems will be analyzed to provide a basis for the innovation of the present invention.
[0092] 4.1 Kriging field value interpolation:
[0093] In the present invention, the gamma value of normal data (usually the statistical characteristics of the data, such as standard deviation, variance, etc.) and coordinate information are first used for Kriging interpolation to build a spatial model to infer the field value of the normal area. This process provides a basis for the diffusion effect of outliers.
[0094] In order to deal with the interpolation problem of normal data, the present invention uses the Kriging Interpolation method. Kriging interpolation is an interpolation method based on spatial autocorrelation, and its core assumption is that there is a spatial dependency between data points. The Kriging method estimates the value of unknown points by minimizing the error variance, and is particularly suitable for processing geographic data and field value data with spatial correlation.
[0095] The core of the Kriging interpolation method lies in the construction of the variation function and parameter estimation. It can handle the randomness and trend of data and is suitable for the interpolation of complex geological or geographical fields. It uses spatial autocorrelation to construct an interpolation model that minimizes errors.
[0096] The influence model of outliers based on inverse distance weighting:
[0097] In nature, many physical phenomena (such as gravitational fields, electric fields, etc.) follow a certain attenuation law. Anomalies can be regarded as a "source" in a physical field, and their influence on the surrounding area gradually decays. The field value generated by the anomaly is not limited to itself, but will have an impact on the surrounding area, similar to the distribution of electric fields or gravitational fields. By introducing a physical field attenuation model in the interpolation process, this influence can be simulated, and the effect of the anomaly can be naturally transitioned and integrated into the field of normal data. The present invention uses an improved IDW interpolation method and a weighted smoothing method based on deep learning to introduce anomalies.
[0098] The definition of outliers uses the 3 times mean standard, which is a common practice based on statistics to identify points in the data that deviate significantly from the normal range. This method calculates the mean and standard deviation of the data and defines points that are beyond the mean by 3 times the mean as outliers. Because outliers can be caused by accidental or systematic errors, it is important to treat them specially during the interpolation process.
[0099] Assume that the addition of an outlier will affect the field values in the surrounding area. Just like the effect of physical fields such as gravity or electric fields, an outlier can be regarded as a "source" that affects the "field" around it. To achieve this, an attenuation model similar to that of a physical field can be used to gradually diffuse its effect through an influence function.
[0100] In the present invention, the influence of outliers is simulated by a model similar to the attenuation of physical fields. As a "source", outliers will spread their influence in all directions, and the intensity of the influence will gradually weaken as the distance increases. The IDW (Inverse Distance Weighting) interpolation method is used to simulate this attenuation process. The IDW interpolation method assumes that the closer the known points are to the target point, the greater the influence on the interpolation result. By setting an influence function, the gradual attenuation of the field values of the surrounding areas by outliers can be simulated, and finally the influence of outliers can be naturally integrated into the field of normal data.
[0101] Inverse Distance Weighting (IDW) is a simple interpolation method based on the principle of "the closer the distance, the greater the influence". IDW assumes that the influence of each point to be estimated is inversely proportional to its distance, and the commonly used weight function is a negative power function of the distance. The advantage of IDW interpolation is that it is simple and intuitive, but because it is sensitive to distance, it may cause a certain degree of "cliff" phenomenon in the interpolation result, so it is necessary to reasonably select the power parameter when applying it.
[0102] The expression of inverse distance weighted interpolation is:
[0103]
[0104] in, is the field value of the point to be estimated, is the field value at a known point, is the distance between the estimated point and the known point, is a power parameter (usually ), which controls the effect of distance on interpolation.
[0105] Deep learning has significant advantages in spatial data analysis, especially when dealing with data containing noise, outliers and irregularities. The weighted smoothing network based on deep learning can adaptively adjust the smoothing process of data by introducing learning capabilities, while retaining important outlier features, thereby better processing complex spatial data surfaces.
[0106] The way deep neural networks smooth data is to convert noisy input data into smooth output data through regression learning. Specifically, the neural network learns the mapping relationship between input data (such as noisy contour coordinates) and ideal output data (i.e. smoothed contour coordinates) to perform the smoothing operation. The following is a detailed description of this process.
[0107] 1. Generation of training data:
[0108] Input data (noisy contour): First, generate noisy circular contour data, which contains noise to simulate abnormal or irregular data that may exist in reality.
[0109] Output data (smoothed contour): The ideal output data is a circular contour without noise. This data represents the result after smoothing. The goal is to let the neural network recover this ideal output from the noisy input data.
[0110] For example, if the input contour contains slight perturbations (such as jitter or offset), the output should be a smooth circular contour that removes these perturbations.
[0111] 2. Neural network structure:
[0112] The network adopts a simple fully connected layer architecture, including:
[0113] l Input layer: receives the coordinates of each data point (such as x, yx, yx, y).
[0114] Hidden layer: Use nonlinear activation functions (such as ReLU) to increase the expressive power of the model and capture complex patterns in the data.
[0115] l Output layer: Outputs smooth coordinates with the same dimensions as the input data. The goal is to remove noise through the network and restore a smooth curve.
[0116] 3. Network training:
[0117] (1) Training objective: The network is trained by minimizing the error between the predicted value and the true smoothed value (usually using the mean square error loss function). The training process continuously adjusts the network weights so that the network can accurately predict the corresponding smoothed output when given input data.
[0118] (2) Optimization process: The network uses the Adam optimizer (adaptive momentum estimation) to optimize parameters and find the best network weights so that the output results are as close to the real smooth data as possible.
[0119] 4. Data smoothing process:
[0120] (1) Input noisy data: When processing each contour, the network receives the x and y coordinates of each contour, which usually contain a certain amount of noise.
[0121] (2) Predicting smooth output: The network processes the input data and predicts the smoothed result through the trained weights. Specifically, the neural network gradually removes the noise in the input data through nonlinear transformations in multiple hidden layers and restores a smooth curve.
[0122] (3) Output smoothing results: The smoothing results of each contour are generated through the output layer of the network, and these smoothed coordinates are updated back to the original Shapefile data.
[0123] Kriging interpolation effect verification:
[0124] The validation mechanism of the Kriging algorithm combines K-fold cross validation and multiple validation indicators. By dividing the data set into K subsets, training and validation are performed in a cycle to ensure that each sample is validated once. The main validation indicators include RMSE (root mean square error), MAE (mean absolute error) and R² (coefficient of determination) to measure the prediction accuracy and fit of the model.
[0125] In addition, standardized residuals and confidence interval coverage are validation tools specific to kriging, the former used to assess prediction uncertainty and the latter to check the reliability of confidence intervals.
[0126] Through the comprehensive evaluation of these indicators, model parameters can be optimized, model assumptions can be diagnosed, and the prediction accuracy of the model and the reliability of uncertainty estimation can be ensured. The validation mechanism of the Kriging algorithm includes the following key parts:
[0127] 1. K-fold cross validation principle:
[0128] First, the dataset is randomly divided into K subsets (default K=5), and K-1 subsets are used for training each time, and the remaining 1 subset is used for validation. This cycle is repeated K times to ensure that each sample is used as a validation set.
[0129] 2. Verification indicators:
[0130] a) RMSE (root mean square error): measures the deviation between the predicted value and the actual value. Formula:
[0131] ;
[0132] in, is the number of points of the interpolated data, is the predicted value of the interpolated data point, is the actual value of the interpolated data point;
[0133] b) MAE (Mean Absolute Error): Provides a direct measure of error, unlike RMSE which amplifies the impact of large errors. Formula:
[0134] ;
[0135] in, is the predicted value of the interpolated data point, is the actual value of the interpolated data point;
[0136] c) R² (coefficient of determination): measures the ability of the model to explain the variability of the data. Formula:
[0137] ;
[0138] in, is the predicted value of the interpolated data point, is the actual value of the interpolated data point.
[0139] Kriging-specific validation mechanism:
[0140] a) Standardized residual: It measures the degree of standardization of the model prediction error and reflects the prediction uncertainty of the Kriging algorithm. If the model accurately estimates the prediction uncertainty, the standardized residual should approximate the standard normal distribution. Formula:
[0141] ;
[0142] in, is the predicted value of the interpolated data point, is the actual value of the interpolated data point, is the standard deviation of the predicted values of the interpolated data points;
[0143] b) Confidence interval coverage: Check whether the actual value falls within the range of ±1.96 times the predicted standard deviation. In theory, the confidence interval should cover 95% of the actual values (assuming that the prediction variance is estimated accurately).
[0144] Application of verification results:
[0145] ①Optimize model parameters by comparing verification indicators (such as range, sill value, etc.) under different parameter settings.
[0146] ②The distribution of standardized residuals can diagnose the rationality of model assumptions.
[0147] ③Confidence interval coverage is used to evaluate the reliability of uncertainty estimates.
[0148] Anomaly point introduction verification mechanism based on inverse distance weighting:
[0149] For verification of the introduction of outliers, the magnitude of the change in the field value of each data point can be checked through volatility, the proportion of outliers, and visual verification to ensure that the change is within a reasonable range; the proportion of outliers in the data is monitored to ensure that it is consistent with expectations. The impact of outliers is checked through graphical means to ensure that their impact does not cause unnatural fluctuations. The following verification mechanisms are evaluated:
[0150] Volatility Verification:
[0151] Volatility verification is designed to detect whether the change in field value before and after the introduction of anomalies is too drastic, so as to ensure the smoothness and rationality of the data. This can be done by calculating the change in field value of each original data point before and after the introduction of anomalies, and ensuring that this change is within a predetermined reasonable range.
[0152] Assume that the field value of the original data is , the field value after introducing the abnormal point is , then the field value change of each data point It can be defined as:
[0153] ;
[0154] in, is the original data point The field value, is the field value after the outlier point is introduced.
[0155] For all data points, the overall maximum change or root mean square (RMSE) change can be calculated:
[0156] ;
[0157] ;
[0158] in is the total number of data points. The smaller the RMSE, the smoother the field value changes. If the changes of some points exceed the predetermined maximum threshold, the changes are considered unreasonable and the abnormal point introduction mechanism may need to be adjusted. It is the maximum allowable field value change after the normal point is affected by the abnormal point field.
[0159] Verification of abnormal point ratio:
[0160] The purpose of outlier ratio verification is to ensure that the ratio of introduced outliers is consistent with expectations, and to avoid introducing too many or too few outliers, which would have an unreasonable impact on the data.
[0161] First, we need to define the standard of outliers. Assume that the field value of the original data follows a normal distribution with a mean of , the standard deviation is , here according to the field value classification standard - anomalies are defined as points that exceed three standard deviations from the mean or ;
[0162] Percentage of abnormal points It can be expressed as:
[0163] ;
[0164] in is the total number of data points, and the number of abnormal points is the field value exceeding the threshold The number of data points, where is the field value mean of normal data points, is the standard deviation of the field value of normal data points.
[0165] In order to verify whether the proportion of outliers after introduction is within a reasonable range, the proportion of outliers after introduction can be calculated. , and the target ratio (Percentage of abnormal data points after interpolation) for comparison:
[0166] ;
[0167] if Exceeds the predetermined tolerance fluctuation range , it means that the outlier introduction mechanism is too strong or insufficient, and it may be necessary to adjust the weight coefficient or the introduction mechanism. is the proportion of outliers in the total number of data points after the introduction of outliers, is the target ratio of allowable outliers to the total number of data points.
[0168] Visual verification:
[0169] Visual verification uses graphical means to check the impact of outliers on field values and observe whether they effectively affect the surrounding area without causing unnatural fluctuations. This verification mainly relies on visualization methods such as contour maps. Draw contour lines of field values before and after the introduction of outliers, check the smoothness of the contour lines and whether the outliers cause unreasonable curve changes. Ideally, outliers should have a local impact on field values without causing global fluctuations.
[0170] Through these verification mechanisms, the data quality after the introduction of outliers can be ensured, excessive interference with the original data can be avoided, and it can be ensured that the introduction of outliers is reasonable and meets the predetermined goals.
[0171] The training update mechanism of the weighted smoothing network based on deep learning is as follows:
[0172] 1. Loss function calculation:
[0173] In the forward propagation, the mean square error (MSE) between the predicted value and the target value is calculated and multiplied by the normalized weight of each sample, so that the loss function can reflect the weighted contribution of different data points, helping the network to pay more attention to the influence range of outliers. Using weighted mean square error (MSE) as the loss function Loss makes the network sensitive to the weighted differences of different points:
[0174] ;
[0175] in, is the number of samples, is the true value, is the predicted value of the network output. The loss of each sample point will be multiplied by its weighted influence, highlighting the impact of outliers in the overall loss and enhancing the model's sensitivity to outliers. The weighted mean square error (MSE) is used as the loss function.
[0176] 2. Back propagation and parameter update:
[0177] During the gradient propagation of the weighted loss, the gradient value is weighted according to the weight, thereby adjusting the weight and bias of each layer. The network updates the weights through stochastic gradient descent (SGD) and momentum to preserve the influence of outliers as much as possible during back propagation. Based on the gradient calculation of the weighted loss function, the update formula of the weight matrix is:
[0178] ;
[0179] in, is the learning rate, The momentum term is the weight change at the previous moment, and the convergence process is accelerated by the momentum term. In the back propagation, the network will adjust the update amount according to the weighted influence of each sample point, so that the characteristics of the outliers are preserved as much as possible when the weights are updated. is the weight change after network update, is the desired partial derivative of the loss function with respect to the weight w.
[0180] Data processing part:
[0181] 5.1 Outlier Classification:
[0182] This module defines outliers based on the mean of the data - using 3 times the gamma background mean, it is able to quickly and effectively identify outliers that deviate significantly from the overall trend of the data by directly comparing the data to the mean. By comparing the data value to 3 times its mean, it can highlight values that are far from the center of the data, especially when the data is relatively evenly distributed or does not rely on the standard deviation.
[0183] The data processing in the present invention focuses on the deviation of data points from the average level, rather than the details of the distribution (such as the analysis of normal distribution or skewed distribution), and the standard deviation may be affected by extreme values, resulting in inaccurate results. The definition of 3 times the mean can better reflect the "normal" range of the data itself, avoiding over-reliance on local fluctuations.
[0184] It can be observed that there are very few points that are more than 3 times the mean, indicating that 3 times the mean can effectively mark the extreme outliers in the market value. There is a small amount of long tail on the right side. Although the frequency is low, there are still some points that are higher than 3 times the mean. These points are considered outliers. Figure 2 From the above, it can be seen that 3 times the mean is a reasonable outlier threshold for this data set. Most of the data points are below 3 times the mean, so outliers that deviate far from the mean can be effectively identified.
[0185] Field level division:
[0186] According to the gamma field value classification standard in the Nuclear Industry Standard of the People's Republic of China "Specifications for the Measurement of Total Ground Gamma" (see Table 1), the standard number of the above-mentioned industry standard "Specifications for the Measurement of Total Ground Gamma" is EJ / T 831-1994, and the gamma value of the data after the outliers are removed is classified into six levels: low field, low field, normal field, high field, high field and abnormal field. The interpolated data and the data after the outliers are added are divided into field levels according to the data divided by this level.
[0187] Table 1
[0188]
[0189] Remove the mean of the data with outliers , standard deviation According to the field level division standard, the original data field value area (including abnormal points) is divided into: low field area ( )8 points, low area ( )2241 points, normal area ( )24158 points, high area ( )4793 points, high field area ( )557 points, abnormal area ( )517 points. Pay attention to the abnormal area point ( ) accounts for 1.60%.
[0190] Kriging normal field value interpolation:
[0191] 5.3.1 Kriging interpolation:
[0192] Kriging interpolation is used for normal value data (data other than the abnormal field). The specific formula of Kriging interpolation is:
[0193]
[0194] in, is the actual gamma value, Yes The estimated gamma value of . is the weight coefficient, which uses the weighted sum of the data of all known points in space to estimate the value of the unknown point. However, the weight coefficient is not simply the inverse of the distance, but can satisfy the point The estimated value at With the true value The optimal set of coefficients with the smallest difference is:
[0195]
[0196] At the same time, the conditions for unbiased estimation are met:
[0197]
[0198] For estimated value With the true value In the process of data gridding, Krigings interpolation takes into account the spatial correlation properties of the object being described, making the interpolation result more scientific and closer to the actual situation, and can give the interpolation error (Kriging variance), making the reliability of the interpolation clear at a glance.
[0199] Kriging interpolation effect analysis:
[0200] The effect of using MATLAB software to perform Kriging interpolation on normal values is shown in Figure 3 , Figure 3 The comparison between the predicted value and the actual value is shown. The data points are basically distributed along the ideal 45-degree diagonal line (dashed line), indicating that the model is relatively accurate in predicting most data. The dispersion of data points is relatively small in the low-value area, while more discrete points appear in the high-value area, which may mean that there is a certain deviation in the prediction of the high-value area. According to the interpolated average determination coefficient MeanR²=0.88186, the overall fit is high and the model captures the trend of most data.
[0201] The mean square error of the predicted value is Mean MSE=2.6614, which is relatively low, but there are still some areas with large deviations. We can focus on these high variance areas for further optimization. Figure 4 The distribution of the prediction variance is shown. It can be seen that the variance is high near 0, showing a skewed peak state, indicating that the error of most predictions is small. The variance distribution is wide, and the existence of extreme variance values indicates that the interpolation accuracy of the model has decreased in some areas, especially the part with variance above 10, which may correspond to outliers or areas with uneven data distribution.
[0202] Because the interpolated data increases the density of the original data and fills the space gaps of the original data (excluding outliers), the background value of the field value division is based on the statistical value of the interpolated data. , standard deviation According to the field level division standard, the original data field value area (including abnormal points) is divided into: low field area ( )44 points, low area ( )58709 points, normal area ( )743232 points, high area ( )170174 points, high field area ( )16461 points, abnormal area ( )11380 points. Also pay attention to the abnormal area point at this time ( ) accounts for 1.14%, and the original data abnormal area point ( ) accounts for 1.60%.
[0203] Introduction model of outliers:
[0204] 5.4.1 Outlier point introduction model based on IDW interpolation:
[0205] It is understood that the addition of outliers will have an impact on the field values in the surrounding area. Just like the effects of physical fields such as gravity or electric fields, outliers can be regarded as a "source" that will affect the "field" around them. To achieve this, the present invention uses an attenuation model similar to the physical field, gradually diffusing its effects through the influence function, including the influence of outliers on outliers within the influence range and the influence of outliers on normal points, introducing a distance-based attenuation weight mechanism, and considering the mutual influence between outliers and the influence of outliers on normal points.
[0206] Inverse Distance Weighting (IDW) is used to simulate the influence of foreign points on the original data points. The basic idea of IDW is that the closer the foreign point is to the target data point, the greater its influence on the target data point. The specific process is as follows:
[0207] 1. Inverse distance weighted principle:
[0208] The inverse distance weighting method assumes that the external points closer to the target point have a greater influence on the target point, and the external points farther away have a smaller influence. The influence is usually determined by the inverse proportional relationship of the distance, and the weighting coefficient is usually defined as a negative power of the distance (such as the inverse of the square of the distance).
[0209] Impact calculation: Given an external point , the distance from the original point The Euclidean distance is:
[0210] ;
[0211] The inverse relationship between influence and distance is adjusted by a weighting factor:
[0212] ;
[0213] in is the distance decay exponent, A small value (to prevent division by zero), usually a very small value (such as eps).
[0214] Weighted average calculation: The new field value of the target point (field_orig) is obtained by weighted averaging the field value of the foreign point (field_foreign):
[0215] ;
[0216] in, is the new field value, is the original field value, is the weight, is the field value of the external point.
[0217] 2. Methods and steps:
[0218] Input parameters are shown in Table 2:
[0219] Table 2
[0220]
[0221] Calculation process:
[0222] a) Iterate over each original data point (original_data).
[0223] b) For each original data point, traverse all foreign points (foreign_points).
[0224] c) Calculate the Euclidean distance between the original data point and each foreign point.
[0225] d) If the distance is less than or equal to the set threshold threshold_distance, calculate the influence weight of the foreign point on the original data point and accumulate the influence value.
[0226] e) Calculate the new field value of the original data point through the weighted influence of all external points and update the result.
[0227] 3. Weighted calculation of impact:
[0228] a) The impact value is weighted summed by multiplying the field value of the external point by its weight.
[0229] b) The weight w is calculated in such a way that closer external points have a greater impact on the field value of the original point.
[0230] c) The final result is to add the weighted influence value to the field value of the original data point to obtain a new field value.
[0231] This method can effectively simulate the impact of outliers or foreign points on original data points by adjusting the impact range (controlled by threshold_distance) and impact strength (controlled by p). Through weighted averaging, the impact of outliers is reasonably extended to their surrounding areas, ensuring the smoothness of the data.
[0232] For each original point, the weighted field value is given by:
[0233] ;
[0234] in is the distance from the original point The number of foreign points less than threshold_distance, It is The foreign point The weight of the original point, It is The original point and The Euclidean distance between the external points, It is The field value of an external point.
[0235] When introducing outliers, this method can accurately control the influence range and intensity of the outliers by adjusting p and threshold_distance, thereby ensuring the rationality and smoothness of the interpolation results.
[0236] 4. Control of the proportion of abnormal point areas:
[0237] At the same time, the control of the proportion of abnormal point areas is intended to ensure that the proportion of introduced abnormal points is consistent with expectations, avoiding the introduction of too many or too few abnormal points, which would have an unreasonable impact on the data.
[0238] First, we need to define the standard of outliers. Assume that the field value of the original data follows a normal distribution with a mean of , the standard deviation is , here according to the field value classification standard - anomalies are defined as points that exceed three standard deviations from the mean or ;
[0239] Percentage of abnormal points It can be expressed as:
[0240] ;
[0241] in is the total number of data points, and the number of abnormal points is the field value exceeding the threshold The number of data points.
[0242] In order to verify whether the proportion of outliers after introduction is within a reasonable range, the proportion of outliers after introduction can be calculated. , and the target ratio (Percentage of abnormal data points after interpolation) for comparison:
[0243] ;
[0244] if Exceeds the predetermined tolerance fluctuation range , it means that the outlier introduction mechanism is too strong or insufficient, and the weight coefficient or the introduction mechanism needs to be adjusted.
[0245] 5.4.2 Abnormal points introduce fluctuation adjustment mechanism:
[0246] In order to ensure that the fluctuation of the abnormal area ratio is within a reasonable range, it is necessary to determine whether the effect of introducing abnormal points is appropriate by calculating the fluctuation of the abnormal area ratio.
[0247] The fluctuation adjustment mechanism calculates the fluctuation of the abnormal area ratio and dynamically adjusts the k value to ensure that the fluctuation is within the expected range. Fluctuation detection calculates the field value change and compares it with the threshold to ensure that the data change does not exceed the reasonable range. Dynamically adjust the k value to enhance or weaken the effect of introducing abnormal points, thereby ensuring that the fluctuation of the abnormal area ratio is moderate, which can effectively introduce abnormal points without causing excessive fluctuations.
[0248] Specifically:
[0249] (1) Fluctuation of abnormal area proportion:
[0250] Given the target anomaly ratio q_target, calculate the actual anomaly ratio q_actual and find its fluctuation:
[0251] ;
[0252] q_actual: The actual abnormal area ratio, which means that after the introduction of abnormal points, the area exceeds three times the standard deviation ( ) range.
[0253] ;
[0254] q_target: the proportion of abnormal areas in normal data, is the number of outliers, is the number of new data after the outliers are introduced.
[0255] (2) Volatility adjustment rules:
[0256] ① Fluctuation is too small: If , indicating that the effect of introducing outliers is insufficient. At this time, it is necessary to increase the k value to enhance the effect of introducing outliers. The allowed target fluctuation size is set.
[0257] ;
[0258] in, is the new outlier point ratio coefficient, is the old outlier point proportion coefficient;
[0259] ② Excessive fluctuations: If , indicating that the outliers are too strong. In this case, we need to reduce the k value to weaken the influence of outliers.
[0260] (3) Fluctuation detection formula:
[0261] After each parameter combination optimization, it is necessary to detect whether the field value change exceeds a given threshold. Fluctuation detection is performed by calculating the field value change of each original data point.
[0262] Field value change calculation (field_change). For each data point, calculate the field value change before and after the introduction of the outlier point:
[0263] ;
[0264] : Field value of the original data point.
[0265] : Field value of the data point after IDW influence.
[0266] ③ Determine whether the change exceeds the threshold:
[0267] If the field value change of a data point exceeds the given threshold, it is considered that the change of the point is beyond the reasonable range and a fluctuation warning is generated. The formula is as follows:
[0268] ;
[0269] is the given field value change threshold, To warn of volatility;
[0270] (4) Specific calculation of dynamic adjustment of k value:
[0271] In order to make the fluctuations conform to the expected range, the proportional coefficient k is dynamically adjusted according to the size of the fluctuations during the optimization process. The specific method of dynamic adjustment is:
[0272] If the fluctuation is too small ( ), increase the k value and enhance the influence of outliers:
[0273] ;
[0274] If the fluctuation is too large ( ), reduce the k value and weaken the influence of outliers:
[0275] ;
[0276] In this way, the effect of introducing anomalies can be controlled and the fluctuation of the proportion of anomalies can be kept within a reasonable range.
[0277] (5) Specific implementation steps:
[0278] Step 1: Calculate the impact of outliers on original data points through the IDW interpolation method to obtain the affected data.
[0279] Step 2: Calculate the abnormal point ratio q_actual based on the new data, compare it with the target ratio q_target, and calculate the fluctuation:
[0280] ;
[0281] Step 3: Determine whether the k value needs to be adjusted based on the fluctuation value to ensure that the fluctuation of the abnormal area ratio is within a reasonable range.
[0282] Step 4: Optimize by dynamically adjusting the k value (increasing or decreasing) to ensure that the abnormal points finally introduced can effectively increase the proportion of abnormal areas without causing the abnormal areas to be too drastic and affecting the smoothness of the data.
[0283] Step 5: Use the fluctuation_check function to verify whether the field value change of each data point exceeds the threshold to ensure that the change is reasonable.
[0284] Optimization of model parameters by introducing outliers:
[0285] A sliding window is used to use grid search to optimize the relevant parameters of IDW interpolation (distance decay exponent p, threshold threshold_distance, proportional coefficient k) so that the proportion of abnormal points in the data fluctuates moderately and the best one is selected from multiple parameter combinations.
[0286] 1. Parameter initialization and setting:
[0287] •p_range: The search range of the distance decay exponent p, with values from 1 to 5 and a step size of 0.5.
[0288] •threshold_range_start and threshold_range_end: The starting and ending points of the threshold range, indicating the maximum distance affected, the default is 100-500.
[0289] •window_width: The width of the sliding window, that is, the range of the threshold_distance parameter calculated each time, the default is 100.
[0290] • window_step: Window movement step, which indicates the distance the window moves backward from current_start each time. The default value is 50.
[0291] •k_range: The search range of the scaling factor k, with a step size of 0.1.
[0292] •threshold_step_size: The step size of threshold_distance within the window, that is, the step size of the threshold change within each window, the default is 5.
[0293] 2. Sliding window traversal:
[0294] The code uses a sliding window traversal method to explore different values of threshold_distance. The specific process is as follows:
[0295] (1) Initialize the window range:
[0296] a) current_start indicates the starting position of the current sliding window.
[0297] b) current_end indicates the end position of the current window, ensuring that it does not exceed threshold_range_end.
[0298] c) Calculate the threshold_range in the current window, where the step size of threshold_range is threshold_step_size. If the number of parameters in the window is insufficient, fill it with zeros to make the window size consistent with the predetermined size.
[0299] (2) Grid Search:
[0300] a) In each window, for each combination of p, threshold_distance and k, calculate the influenced data (influenced_data).
[0301] b) Analyze the data through the analyzeLevels function to obtain the abnormal point ratio q_actual.
[0302] c) Calculate the fluctuation of the abnormal point ratio, and compare it with the target fluctuation m_target.
[0303] d) If the current fluctuation is less than the minimum fluctuation value of the current window (min_fluctuation) and the fluctuation is within the allowed range, update the optimal parameters of the window.
[0304] (3) Dynamically adjust k:
[0305] a) During the search process, if the fluctuation is too small ( ), increase the proportional coefficient k, otherwise reduce the proportional coefficient to ensure that the fluctuation range can be stabilized at the expected value.
[0306] b) Dynamically adjust the value of k, but make sure it stays between the minimum and maximum values of k_range.
[0307] 3. Store and update optimal parameters:
[0308] (1) The optimal parameters of each window are stored in the best_params_window structure, including:
[0309] a)p: optimal distance decay exponent.
[0310] b)threshold_distance: optimal threshold.
[0311] c) k: optimal proportionality coefficient.
[0312] d)min_fluctuation: the minimum fluctuation value of the current window.
[0313] e)q_actual: the actual abnormal point ratio of the current window.
[0314] (2) Store the optimal parameters of each window in best_params_list.
[0315] 4. Window movement:
[0316] After each window search is completed, current_start is updated and the window is slid backwards by window_step. The window slides forward 50 units each time, that is, the overlap is 50 units.
[0317] 5. Output optimal parameters:
[0318] (1) After searching all windows, output the optimal parameters for each window.
[0319] (2) Finally, the optimal parameters in all windows are selected, and the median of the proportion of qualified areas q_actual and the corresponding parameter combination are selected as the final optimal parameters.
[0320] 6. Return results:
[0321] Finally, the optimal parameters (best_params) of the entire process are returned, which is the parameter combination corresponding to the median of the area proportion value q_actual in the entire sliding window search.
[0322] The core of the traversal logic lies in the search of sliding windows. In each window, the optimal parameters are found through grid search (a combination of p, threshold_distance, and k). In each window, the fluctuation of the proportion of abnormal points is used as the optimization target, the proportional coefficient k is adjusted dynamically, and the parameters of each window will overlap with the previous window, so as to ensure a more detailed search process.
[0323] 5.4.4 Analysis of the effect of introducing abnormal points:
[0324] The optimal parameter combination was obtained by introducing the influence of outliers using MATLAB software: the search range of the distance decay index p is p=3, the influence range threshold_distance=250, the optimal proportional coefficient k=0.1000, the minimum fluctuation value of the current window min_fluctuation=0.0015 (the difference between the proportion of abnormal areas and normal data), and the actual proportion of abnormal points in the current window q_actual=0.0129.
[0325] Contour smoothing:
[0326] Use MATLAB to generate contour images in shp format. The special SHP (Shapefile) file is a commonly used file format in geographic information systems (GIS). It contains the geometric shapes of spatial data (such as points, lines, polygons) and attribute data (such as place names, area numbers, etc.). The key operations in SHP file processing include: ① Processing NaN values: Null values (such as missing data) in SHP files need to be processed. The common method is to interpolate or delete records containing NaN values. ② Maintain the original attribute information: During the spatial data processing process, the integrity of the data must be maintained to ensure that the original attributes (such as area numbers, labels, etc.) are not tampered with or lost.
[0327] Smoothing of contour images is a common image processing task, which is usually used to remove noise, improve image quality, and make contour lines smoother and continuous for easy analysis and visualization. To achieve this goal, the present invention adopts four common smoothing methods: Gaussian smoothing, Savitzky-Golay filtering, cubic spline interpolation, and a contour smoothing model based on deep learning. The method mainly adopts a contour smoothing model based on deep learning.
[0328] 6.4 Contour line smoothing model based on deep learning:
[0329] In order to smooth the generated contour lines and ensure that they do not intersect, retain local features and are closed within the allowed range, the contour lines can be post-processed through a deep learning network. A convolutional neural network (CNN) can be selected to handle this task. The following are the specific design ideas and steps.
[0330] 6.4.1 Input and Output Format:
[0331] 1. Input format: The input data is a Shapefile (input_shp), which contains the contour information of multiple contour lines. Each contour consists of several (x,y) coordinate points. Each point in the data needs to be of double type. Shapefile files usually have multiple fields, including coordinates (X,Y) and other attribute data.
[0332] 2. Output format: The output data is also a Shapefile (output_shp), which contains the smoothed contour lines. The coordinates (X, Y) of each contour are smoothed by the deep learning model, and other fields remain unchanged.
[0333] 6.4.2 Network structure design:
[0334] The network structure is designed as follows, as shown in Table 3:
[0335] Input layer: The input data is (x, y) coordinates, so the number of nodes in the input layer is 2, indicating that each point has two features: the horizontal coordinate and the vertical coordinate.
[0336] Hidden layers: Two fully connected layers, containing 64 and 32 neurons respectively, followed by ReLU activation function:
[0337] The first layer has 64 neurons, which are used to learn high-level representations of input features.
[0338] The second layer has 32 neurons, which further refine the feature representation.
[0339] Output layer: The output layer also contains 2 neurons, corresponding to the smoothed horizontal and vertical coordinates (x', y').
[0340] Activation function: Using ReLU (Rectified Linear Unit) activation function can effectively avoid the gradient vanishing problem and accelerate training.
[0341] Regression layer: The regression layer is used for regression tasks to map the output of the network into continuous smooth coordinate values.
[0342] Table 3
[0343]
[0344] 6.4.3 Training process:
[0345] 1. Generate training data:
[0346] (1) 100 training samples were simulated, each with 50 points representing a circular contour. The radius of the circle was randomly perturbed to simulate contours of different shapes.
[0347] (2) Add noise to each training example (by randomly generating small perturbations) to simulate noisy contour data.
[0348] (3) The target output (ideal smoothing result) is a circular contour without noise.
[0349] 2. Training methods, see Table 4:
[0350] (1) Using the Adam optimizer, the maximum number of training epochs is 30 (each epoch represents a pass through all training data).
[0351] (2) Each training session uses 32 samples as a mini-batch for gradient descent updates.
[0352] (3) The learning rate is initialized to 0.001 and will be automatically adjusted during training.
[0353] Table 4
[0354]
[0355] 3. Training objectives:
[0356] The goal of the network is to transform the noisy contour point X into a smooth ideal output Y, that is, to remove noise and make the contour smoother.
[0357] 4. Output and post-processing:
[0358] The trained network will generate smooth and satisfactory contour lines. The following are the post-processing steps:
[0359] (1) Remove intersections: Ensure that contour lines do not intersect by detecting intersection points in the image.
[0360] (2) Closed repair: If the endpoints of the contour lines are not closed, morphological operations (such as dilation, erosion, etc.) can be used to repair the unclosed areas.
[0361] (3) Local feature preservation: Ensure that the local features of the contour lines are preserved during the smoothing process.
[0362] 6.4.4 Specific steps: contour generation and post-processing:
[0363] 1. Contour line generation: Use the analyzeLevels function to generate preliminary contour line data (contour_lines.shp).
[0364] 2. Post-processing of contour lines: Input contour_lines.shp into the deep learning network for smoothing, remove intersections and cracks, and ensure that the contour lines are closed and smooth.
[0365] 3. Read Shapefile data:
[0366] (1) Use shaperead to read the input Shapefile file and obtain all contour data.
[0367] (2) If the Shapefile is empty or fails to read, an error message is thrown.
[0368] 4. Data preprocessing:
[0369] (1) Ensure that the coordinate data (X, Y) of each contour is of double type to avoid problems caused by inconsistent data types.
[0370] (2) Remove invalid data (NaN and Inf) to ensure that the coordinates of each contour are valid.
[0371] (3) Make sure each contour has enough points (at least 3 points) to perform smoothing.
[0372] 5. Training Model:
[0373] Call the train_smoothing_model function to train a deep learning model. The model is based on noisy training data and learns how to restore noisy data to a smooth contour.
[0374] 6. Process each contour:
[0375] (1) For each contour, extract its coordinates, remove invalid values, and ensure that the coordinates are of double type.
[0376] (2) Input the coordinates into the trained deep learning model to obtain the smoothed coordinates and update the contour. The specific process can be seen in Table 5.
[0377] Table 5
[0378]
[0379] 7. Save and visualize the results:
[0380] (1) Save the smoothed contour data in Shapefile format.
[0381] (2) Use the visualize_contours function to draw a smoothed contour map to see the smoothing effect.
[0382] 6.5 Analysis of contour smoothing effect:
[0383] Figure 5 It is the result of processing by the deep neural network model. After the contour image is processed by the deep learning model, the noise and outliers are effectively removed, making the contour curve smoother and more uniform, eliminating the irregular fluctuations and sharp corners in the original data. The model maintains the overall shape of the contour, while finely smoothing the local details to avoid unnatural transitions. The processed image has clear boundaries, smooth and natural transition areas, no jagged or jumping phenomena, and the overall visual effect is more harmonious and has a higher visualization quality. It is suitable for scientific analysis and display, and is overall better than the first three methods.
[0384] The contour lines processed by the deep learning model have an overall trend that is closer to the original data distribution, eliminating irregular fluctuations and sharp corners in the original data. The model maintains the overall shape of the contour, while finely smoothing the local details to avoid unnatural transitions. In general, compared with the three traditional methods, if you need to strike a balance between maintaining details and smoothing, Savitzky-Golay filtering is a better choice, but it has breakpoints in some places and even has a phenomenon of short-distance straight line connections, and there are local unevenness. Compared with traditional Gaussian smoothing, Savitzky-Golay filtering and spline interpolation methods, the contour smoothing method based on the deep learning model has stronger adaptability and denoising capabilities, and can effectively process complex nonlinear data and retain important local details. It avoids the parameter setting and possible over-smoothing problems in traditional methods by automatically learning data features, and can more accurately retain key changes in the data while denoising, especially when processing data with noise or complex boundaries, showing higher accuracy and adaptability.
[0385] 7.1 Kriging interpolation of normal data:
[0386] Comprehensive analysis shows that Kriging interpolation performs well in most areas, especially in places where data are densely distributed. It can effectively capture the overall trend of the data, with an R² value close to 1 and a low mean square error (MSE).
[0387] However, in areas with highly uneven data distribution, especially in areas with high variance, the model still has some prediction bias. The reasons for this bias may include outliers in high-value areas or complexity in areas with dense data. To further improve the accuracy of the model, it is recommended to make detailed adjustments in areas with uneven data distribution and optimize interpolation parameters, such as using local weighted interpolation in high variance areas or adjusting parameter settings in the Kriging model. In addition, enhancing the modeling of the impact of outliers will help improve the model's adaptability to complex areas and improve overall prediction accuracy.
[0388] 7.2 IDW method for introducing outliers:
[0389] By introducing the physical field diffusion model, the influence of outliers is naturally integrated into the overall field, which not only retains the unique characteristics of outliers, but also ensures the global smoothness and continuity of the data field. By combining Kriging interpolation, inverse distance weighted interpolation and deep learning optimization, an efficient, flexible and easy-to-interpret outlier processing solution is provided, which is suitable for a variety of scientific and engineering application scenarios. It has the following advantages:
[0390] (1) Intuitive physical model: The impact of outliers is analogized to the diffusion of physical fields, making the method easy to understand and explain.
[0391] (2) Balance between local and global: The global field is dominated by normal data, and outliers only affect local areas, ensuring the smoothness of the overall field.
[0392] (3) Significance of outliers: The original characteristics of outliers are retained while avoiding the problem of outliers being over-smoothed or directly eliminated.
[0393] (4) Flexibility and scalability: The method can be adapted to different scenarios and data types, for example, by adjusting the parameters of the diffusion function to control the scope of influence.
[0394] Combined with the previous optimization process and objectives, there are several reasonable reasons for choosing the median of q_actual and its corresponding parameter combination as the final optimal parameters:
[0395] (1) Reflects the overall trend of the proportion of abnormal points:
[0396] During the optimization process, the goal is to minimize the fluctuation of the outlier ratio q_actual in all windows and ensure that the outlier ratio of each window is close to the target q_target. The optimal parameter combination of each window will affect the outlier ratio, but since the parameter adjustment of each window may be affected by local data fluctuations or outliers, directly selecting the optimal result of a certain window may not represent the global optimal. By selecting the median of q_actual, the overall trend can be better reflected, avoiding over-reliance on the results of a single window, thereby achieving a more balanced optimization.
[0397] (2) Balancing local fluctuations and global consistency:
[0398] During the optimization process, although we hope that the fluctuation is as small as possible, for some local special cases, the q_actual values of some windows may deviate greatly from the target value (q_target). Using the median can prevent these local anomalies from affecting the overall results, ensuring that the selected parameter combination can not only ensure the smoothness of the overall data, but also maintain a reasonable fluctuation range. In this way, it can avoid that some window parameter combinations appear to be too superior due to local extreme values, and thus select a more stable intermediate value.
[0399] (3) Resist the influence of outliers and ensure the robustness of the results:
[0400] In best_params_list, the parameter combination of each window is optimized by different parameters such as p, threshold_distance and k, which may cause large fluctuations due to abnormal data in some windows. If you simply choose the window parameters with the smallest fluctuation, the impact of these outliers may be ignored. Since the median itself is highly robust to outliers, using the median can effectively avoid the excessive impact of extreme parameter combinations on the final result.
[0401] 7.3 Contour line smoothing:
[0402] In terms of contour line smoothing, different smoothing methods are suitable for different needs and data characteristics.
[0403] Compared with traditional smoothing methods (such as Gaussian smoothing, Savitzky-Golay filtering, and spline interpolation), the quality of smoothing results can be significantly improved. By automatically learning the complex features in the data, the deep learning model can adaptively remove noise and retain key local details in the original data, avoiding the over-smoothing or loss of important features that may occur in traditional methods. In particular, when processing data containing a lot of noise, complex shapes, and nonlinear changes, the deep learning method shows higher accuracy and reliability, ensuring that the smoothing results are more realistic and natural.
[0404] But here are some suggestions:
[0405] 1. Considerations when choosing a smoothing method: For scenarios with simple data and less noise, traditional smoothing methods can continue to be used because they are computationally efficient and easy to implement. However, when the data being processed is more complex, contains a lot of noise, or has highly nonlinear changes, it is recommended to use a smoothing method based on deep learning to ensure the accuracy of the results and the retention of details.
[0406] 2. Deep learning model optimization: To further improve the smoothing effect, it is recommended to enhance the data when training the model and add more diverse samples, especially for data with different shapes and noise levels. In addition, self-supervised learning or transfer learning methods can be introduced to use pre-trained models for fine-tuning to adapt to different types of contour data.
[0407] 3. Adaptive adjustment: The advantage of deep learning models lies in their adaptive ability. Therefore, it is recommended to adjust the network structure and training strategy for different contour data sets in practical applications. In particular, in scenarios where outliers are prominent, the weights of outliers can be increased to ensure that important high-value point features are not lost during the smoothing process.
[0408] 4. Efficiency and computing resources: Although deep learning methods provide higher accuracy and adaptability, their computational cost is relatively high. In practical applications, it is necessary to balance the accuracy of the model with the computing resources, especially when processing large-scale data. It is possible to consider improving computing efficiency by compressing the model or using hardware acceleration.
[0409] In summary, the contour smoothing method based on deep learning shows significant advantages when the data complexity is high, but in application, the method should be flexibly selected according to the specific situation, and the model should be properly optimized and adjusted to maximize its effectiveness in actual scenarios. Example
[0410] like Figure 1 As shown, this embodiment provides a method for processing radioactive geophysical exploration data based on a physical field attenuation model, including:
[0411] Identify anomalies in radioactive geophysical data;
[0412] The inverse distance weighted method is used to construct the diffusion impact model of outliers;
[0413] Use Kriging interpolation method to generate smoothed base fields of normal data;
[0414] The abnormal point diffusion field is superimposed on the normal data field to form a comprehensive field.
[0415] The synthetic field is smoothed using a deep learning-based contour smoothing model.
[0416] As an implementation method in this embodiment, the process of determining abnormal points in radioactive geophysical exploration data includes:
[0417] By calculating the mean and standard deviation of radioactive geophysical data, the points beyond the range of 3 times the standard deviation of the mean are regarded as outliers;
[0418] The coordinates and field value information of the outliers are extracted and stored as separate datasets.
[0419] As an implementation method in this embodiment, the process of constructing the diffusion influence model of the outlier point by using the inverse distance weighted method includes:
[0420] Use the field value of the outlier point as the diffusion source;
[0421] Calculating a weight according to the distance between the outlier point and the target point, wherein the weight is represented by an inverse distance weighted function;
[0422] The diffusion influences of multiple outliers are linearly superimposed to generate the diffusion field values of the outliers, and then the diffusion influence model of the outliers is obtained.
[0423] As an implementation method in this embodiment, the process of generating a smoothed basic field of normal data using the Kriging interpolation method includes:
[0424] Perform spatial correlation analysis on normal data fields and construct semivariogram model;
[0425] Based on the semivariogram model, calculate the interpolation weights;
[0426] Generates a smooth basis field of the normal data field using interpolation weights.
[0427] As an implementation method in this embodiment, the process of superimposing the abnormal point diffusion field onto the normal data field to form a comprehensive field includes:
[0428] Superimpose the diffusion field value of the abnormal point with the normal data field value point by point;
[0429] The superimposed field values are smoothed to form a comprehensive field.
[0430] As an implementation method in this embodiment, the process of smoothing the comprehensive field using the contour smoothing model based on deep learning includes:
[0431] Build a deep learning model for contour smoothing;
[0432] Use the contour data of the comprehensive field as training data to train the deep learning model;
[0433] The trained deep learning model is used to smooth the contour lines of the comprehensive field.
[0434] Based on this, a radioactive geophysical data processing method based on a physical field attenuation model provided by an embodiment of the present invention first uses the triple mean standard to quickly screen the outliers in the radioactive geophysical data, and extracts the coordinates and field value information of the outliers, thereby ensuring the accuracy of the outlier extraction. Then, the outlier diffusion influence model is constructed by the inverse distance weighted method, and the diffusion field value of the outlier is generated by using the field value of the outlier and its distance weight relationship, thereby realizing the spatial natural attenuation of the outlier influence and avoiding the problem of outliers being ignored or over-smoothed in the traditional method. Further, the smoothed basic field of normal data is generated by combining the Kriging interpolation method, and more scientific and realistic basic field value data is generated by modeling the spatial correlation of the data and accurately calculating the interpolation weights. Subsequently, the outlier diffusion field and the normal field are superimposed point by point, and the characteristics of the outliers are retained by smoothing to form a comprehensive field, thereby solving the interpolation distortion problem caused by uneven point data. Finally, the contour smoothing model based on deep learning is used to smooth and optimize the contour lines generated by the comprehensive field data, which effectively improves the continuity and visual effect of the contour lines while retaining the local detail characteristics of the data. Based on the above technical means, the present invention overcomes the limitations of traditional interpolation methods in outlier processing and contour line generation, and provides an innovative solution for the accurate processing and intuitive expression of radioactive geophysical exploration data. Example
[0435] In this embodiment, a computer terminal device is provided, including:
[0436] one or more processors;
[0437] A memory, coupled to the processor, for storing one or more programs;
[0438] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the above embodiments.
[0439] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the method in the above embodiment is implemented.
[0440] In this embodiment, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method in the above embodiment.
[0441] The above program can be run in the processor, or it can be stored in the memory (or computer-readable medium), which includes permanent and non-permanent, removable and non-removable media. Information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0442] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more blocks can be implemented by different modules corresponding to different steps.
[0443] This embodiment provides such a device or system. The system is called a radioactive geophysical exploration data processing system based on a physical field attenuation model, including:
[0444] Anomaly identification module, used to analyze radioactive geophysical exploration data and extract anomalies;
[0445] A model building module for generating a diffusion impact model of outliers based on the inverse distance weighted method;
[0446] Interpolation module, used to generate smoothed base fields of normal data using Kriging interpolation method;
[0447] Data fusion module, used to superimpose the abnormal point diffusion field onto the normal data field to form a comprehensive field;
[0448] The smoothing module is used to smooth the contour lines of the comprehensive field based on the deep learning model.
[0449] As an implementation manner in this embodiment, the model building module includes a weight calculation unit, and the weight calculation unit is used to calculate the diffusion weight according to the distance between the abnormal point and the target point.
[0450] As an implementation in this embodiment, the model building module includes a weight calculation unit, which is used to calculate the diffusion weight according to the distance between the abnormal point and the target point;
[0451] As an implementation in this embodiment, the interpolation module includes a semivariance function construction unit, which is used to generate a spatial correlation model;
[0452] As an implementation method in this embodiment, the smoothing processing module includes a deep learning training unit, which is used to train a smoothing model using contour line data and output a smoothed contour line map.
[0453] The system or device is used to implement the functions of the method in the above-mentioned embodiment. Each module in the system or device corresponds to each step in the method, which has been explained in the method and will not be repeated here.
[0454] Through the above implementation, the problem of radioactive geophysical data processing based on the physical field attenuation model in the related technology is solved, thereby ensuring that the present invention overcomes the limitations of traditional interpolation methods in outlier processing and contour line generation, and provides an innovative solution for the accurate processing and intuitive expression of radioactive geophysical data.
[0455] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A radioactive geophysical exploration data processing method based on a physical field attenuation model, characterized in that: The following steps are involved: Identify anomalies in radioactive geophysical data; The inverse distance weighted method is used to construct the diffusion impact model of outliers. The specific formula is: Among them, F(x, y) is the gamma value of the point with coordinates (x, y) after superimposing the diffusion effect of all abnormal points, ω i is the field value weight of the outlier point, f i (d) is the diffusion function of the i-th outlier point to the target point; or Among them, d is the distance between the outlier point and the target point, p is the weight decay exponent, σ is the diffusion range of the Gaussian distribution, and ∈ is the smoothing parameter to prevent the denominator from being zero; Generate a smoothed base field of normal data using the Kriging interpolation method; Superimpose the abnormal point diffusion field onto the normal data field to form a comprehensive field; The synthetic field is smoothed using a deep learning-based contour smoothing model.
2. The method according to claim 1, characterized in that: The process of identifying anomalies in radioactive geophysical data includes: By calculating the mean and standard deviation of radioactive geophysical data, the points beyond the range of 3 times the standard deviation of the mean are regarded as outliers; The coordinates and field value information of the outliers are extracted and stored as separate datasets.
3. The method according to claim 1, characterized in that Using the inverse distance weighted method, the process of constructing the diffusion impact model of outliers includes: Use the field value of the outlier point as the diffusion source; Calculating a weight according to the distance between the outlier point and the target point, wherein the weight is represented by an inverse distance weighted function; The diffusion influences of multiple outliers are linearly superimposed to generate the diffusion field values of the outliers, and then the diffusion influence model of the outliers is obtained.
4. The method according to claim 1, characterized in that The process of generating a smoothed base field for normal data using the kriging interpolation method involves: Perform spatial correlation analysis on normal data fields and construct semivariogram model; Based on the semivariogram model, calculate the interpolation weights; Generates a smooth basis field of the normal data field using interpolation weights.
5. The method according to claim 1, characterized in that The process of superimposing the abnormal point diffusion field onto the normal data field to form a comprehensive field includes: Superimpose the diffusion field value of the abnormal point with the normal data field value point by point; The superimposed field values are smoothed to form a comprehensive field.
6. The method according to claim 1, characterized in that The process of smoothing the synthetic field using the contour smoothing model based on deep learning includes: Build a deep learning model for contour smoothing; Use the contour data of the synthetic field as training data to train the deep learning model; The trained deep learning model is used to smooth the contour lines of the comprehensive field.
7. A radioactive geophysical exploration data processing system based on a physical field attenuation model, characterized in that: The system comprises: Anomaly identification module, used to analyze radioactive geophysical data and extract anomalies; The model building module is used to generate the diffusion impact model of outliers based on the inverse distance weighted method. The specific formula is: Among them, F(x, y) is the gamma value of the point with coordinates (x, y) after superimposing the diffusion effect of all abnormal points, ω i is the field value weight of the outlier point, f i (d) is the diffusion function of the i-th outlier point to the target point; or Among them, d is the distance between the outlier point and the target point, p is the weight decay exponent, σ is the diffusion range of the Gaussian distribution, and ∈ is the smoothing parameter to prevent the denominator from being zero; Interpolation module, used to generate smoothed base fields of normal data using Kriging interpolation method; Data fusion module, used to superimpose the abnormal point diffusion field onto the normal data field to form a comprehensive field; The smoothing module is used to smooth the contour lines of the comprehensive field based on the deep learning model.
8. The system according to claim 7, characterized in that The model building module includes a weight calculation unit, and the weight calculation unit is used to calculate the diffusion weight according to the distance between the abnormal point and the target point.
9. A computer terminal device, characterized in that: include: one or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the radioactive geophysical exploration data processing method based on the physical field attenuation model as described in any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the radioactive geophysical exploration data processing method based on the physical field attenuation model as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Radiation field inversion reconstruction method based on curved surface average curvature and diffusion equation
CN113205596A
Mine transient electromagnetic three-dimensional display method based on multi-interpolation method
CN113341467A