Soil heavy metal content prediction method, device, equipment and storage medium
By clustering and filtering the soil data set, combining entropy weight distance and recursive neural network optimization, a soil heavy metal content prediction model was established, solving the problems of inaccurate prediction and low efficiency in the existing technology, and achieving more efficient soil heavy metal content detection.
Patent Information
- Application Number
- CN202210884365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-25
AI Technical Summary
In the prior art, the prediction of soil heavy metal content is not accurate enough and inefficient, so it is impossible to understand the soil heavy metal pollution in real time.
By obtaining the initial soil data set, clustering and filtering, and using entropy weight distance detection strategy and recursive neural network optimization, a soil heavy metal content prediction model is established, including data cleaning, clustering, entropy weight distance filtering and recursive neural network training.
It improves the accuracy and efficiency of soil heavy metal content prediction, shortens network training time, and achieves faster and more accurate soil heavy metal content detection.
Smart Images

Figure CN115394370B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soil heavy metal detection, and particularly to a method, device, equipment and storage medium for predicting soil heavy metal content. Background Art
[0002] With the development of modernization, the emissions of heavy metal pollution are continuously increasing, resulting in the increasingly serious problem of soil heavy metal pollution. The excessive heavy metal content in the soil will not only reduce the yield of crops, but also enter the human society and endanger human health. Therefore, it is necessary to detect soil heavy metals to understand the heavy metal content in the soil in real time.
[0003] The existing soil heavy metal content is detected and analyzed manually, and the prediction is inaccurate and inefficient. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method, device, equipment and storage medium for predicting soil heavy metal content, aiming to solve the technical problem of inaccurate prediction of soil heavy metal content in the prior art.
[0005] To achieve the above object, the present invention provides a method for predicting soil heavy metal content, the method comprising the following steps:
[0006] Obtain an initial soil data set;
[0007] Cluster the initial soil data set to obtain target clustering clusters;
[0008] Obtain a reference soil data set through the target clustering clusters;
[0009] Filter the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set;
[0010] Optimize a recurrent neural network through the target soil data set to obtain a target recurrent neural network;
[0011] Train the target recurrent neural network through the target soil data set to obtain a soil heavy metal content prediction model;
[0012] Predict the soil heavy metal content through the soil heavy metal content prediction model.
[0013] Optionally, the filtering the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set includes:
[0014] Obtain the attributes of each data in the reference soil data set and the dimension of the reference soil data set;
[0015] Combine the attributes of the each data to obtain an attribute set;
[0016] Obtain the value sets and value probabilities of each attribute in the attribute set;
[0017] Calculate the entropy value based on the value sets and value probabilities of each attribute;
[0018] Obtain the outlier attributes and non-outlier attributes in the reference soil dataset based on the entropy value and the dimension of the reference soil dataset;
[0019] Set the attribute weights of the outlier attributes and non-outlier attributes respectively;
[0020] Obtain the values of the data objects in the reference soil dataset on the corresponding attributes;
[0021] Calculate the entropy-weighted distance based on the values on the corresponding attributes and the attribute weights;
[0022] When the entropy-weighted distance is greater than the preset distance threshold, regard the data corresponding to the entropy-weighted distance as outlier data;
[0023] Remove the outlier data from the reference soil dataset to obtain the target soil dataset.
[0024] Optionally, optimizing the recurrent neural network with the target soil dataset to obtain the target recurrent neural network includes:
[0025] Obtain the quantity and actual values of the target soil dataset;
[0026] Obtain the predicted values of the recurrent neural network;
[0027] Calculate the mean square error function based on the quantity, the actual values, and the predicted values;
[0028] Set the weights in the target soil dataset;
[0029] Calculate the weight decay term according to the weights;
[0030] Obtain the Jacobian matrix and Hessian matrix of the target soil dataset;
[0031] Calculate the number of effective weights based on the Jacobian matrix, the Hessian matrix, and the quantity;
[0032] Calculate the regularization parameter based on the number of effective weights, the weight decay term, the quantity, and the mean square error function;
[0033] Calculate the objective function based on the regularization parameter, the weight decay term, and the mean square error function;
[0034] Optimize the structure of the recurrent neural network according to the objective function and the regularization parameter to obtain the structure of the target recurrent neural network;
[0035] Obtain the network weights and thresholds of the structure of the target recurrent neural network;
[0036] Update the network weights and the thresholds through the adaptive GA-PSO algorithm to optimize the network error of the recurrent neural network and obtain the target recurrent neural network.
[0037] Optionally, the adaptive GA-PSO algorithm includes an adaptive genetic algorithm and an adaptive particle swarm algorithm;
[0038] The updating of the network weights and the thresholds through the adaptive GA-PSO algorithm to obtain the target recurrent neural network includes:
[0039] Initialize the population through the adaptive genetic algorithm, where the population includes a preset number of individuals, and each individual is a set of parameters of network weights and thresholds;
[0040] Calculate the individual fitness value of each individual;
[0041] Obtain the maximum value of the individual fitness value and the average value of the individual fitness;
[0042] Set the crossover base probability, the mutation base probability, the crossover constant, and the mutation constant, where the crossover base probability is less than the crossover constant, and the mutation base probability is less than the mutation constant;
[0043] Perform crossover and mutation on the individuals through the individual fitness value, the maximum value, the average value, the crossover base probability, the crossover constant, the mutation base probability, and the mutation constant;
[0044] Obtain the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value according to the individual fitness value;
[0045] Calculate and obtain the update probability through the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value;
[0046] Judge whether the individual performs crossover and mutation according to the update probability;
[0047] When the update probability is equal to the preset probability value, output the current population;
[0048] Use the adaptive particle swarm algorithm to calculate the current population to update the network weights and the thresholds to obtain the target recurrent neural network.
[0049] Optionally, using the adaptive particle swarm optimization algorithm to calculate the current population to update the network weights and the threshold to obtain a target recurrent neural network, including:
[0050] Initializing the current population through the adaptive particle swarm optimization algorithm, and setting the initial position and the initial velocity of the current population;
[0051] Calculating the individual fitness values in the current population to obtain the individual extreme value and the global extreme value;
[0052] Setting the maximum number of iterations and obtaining the current number of iterations;
[0053] Setting the position parameter according to the maximum number of iterations;
[0054] Setting the first inertia weight, the second inertia weight, the first learning factor, the second learning factor, and a random value, where the first inertia weight is less than the second inertia weight, and the first learning factor is less than the second learning factor;
[0055] Calculating the current inertia weight through the first inertia weight, the second inertia weight, the position parameter, and the current number of iterations;
[0056] Calculating the current learning factor according to the first learning factor, the second learning factor, the maximum number of iterations, and the current number of iterations;
[0057] Calculating the compression factor through the current learning factor;
[0058] Calculating the particle update velocity according to the compression factor, the current inertia weight, the current learning factor, the random value, the individual extreme value, the global extreme value, the initial velocity, and the initial position;
[0059] Performing local optimization and global optimization training according to the particle update velocity, and when the training requirements are met, determining to update the network weights and update the threshold;
[0060] Optimizing the network error of the recurrent neural network according to the updated network weights and the updated threshold to obtain a target recurrent neural network.
[0061] Optionally, clustering the initial soil dataset to obtain target clustering clusters, including:
[0062] Selecting a target point from the initial soil dataset and using the target point as the initial clustering center;
[0063] Calculating the first distance from each point in the initial soil dataset to the initial clustering center;
[0064] Screen the initial cluster centers through the first distance to obtain target cluster centers;
[0065] Partition the initial soil dataset based on the target cluster centers to obtain target clusters.
[0066] Optionally, the obtaining of the initial soil dataset includes:
[0067] Obtain the original soil dataset;
[0068] Perform data cleaning on the original soil dataset to obtain abnormal data in the original soil dataset;
[0069] Remove the abnormal data from the original soil dataset to obtain the initial soil dataset.
[0070] In addition, to achieve the above object, the present invention also proposes a device for predicting soil heavy metal content, and the device for predicting soil heavy metal content includes:
[0071] An acquisition module, configured to acquire an initial soil dataset;
[0072] A clustering module, configured to cluster the initial soil dataset to obtain target clusters;
[0073] The acquisition module is further configured to obtain a reference soil dataset through the target clusters;
[0074] A filtering module, configured to filter the reference soil dataset through an entropy weight distance detection strategy to obtain a target soil dataset;
[0075] An optimization module, configured to optimize a recurrent neural network through the target soil dataset to obtain a target recurrent neural network;
[0076] A training module, configured to train the target recurrent neural network through the target soil dataset to obtain a soil heavy metal content prediction model;
[0077] A prediction module, configured to predict the soil heavy metal content through the soil heavy metal content prediction model.
[0078] In addition, to achieve the above object, the present invention also proposes a device for predicting soil heavy metal content, and the device for predicting soil heavy metal content includes: a memory, a processor, and a soil heavy metal content prediction program stored on the memory and executable on the processor, and the soil heavy metal content prediction program is configured to implement the steps of the soil heavy metal content prediction method as described above.
[0079] In addition, to achieve the above object, the present invention also provides a storage medium, on which a soil heavy metal content prediction program is stored. When the soil heavy metal content prediction program is executed by a processor, the steps of the soil heavy metal content prediction method described above are implemented.
[0080] The present invention obtains an initial soil data set; performs clustering on the initial soil data set to obtain target clustering clusters; obtains a reference soil data set through the target clustering clusters; filters the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set; optimizes a recursive neural network through the target soil data set to obtain a target recursive neural network; trains the target recursive neural network through the target soil data set to obtain a soil heavy metal content prediction model; and predicts the soil heavy metal content through the soil heavy metal content prediction model, thereby accelerating the network training convergence speed and improving the soil heavy metal content prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 is a schematic structural diagram of a soil heavy metal content prediction device in a hardware operating environment related to the solution of an embodiment of the present invention;
[0082] Figure 2 is a schematic flowchart of the first embodiment of the soil heavy metal content prediction method of the present invention;
[0083] Figure 3 is a schematic overall flowchart of an embodiment of the soil heavy metal content prediction method of the present invention;
[0084] Figure 4 is a schematic flowchart of the second embodiment of the soil heavy metal content prediction method of the present invention;
[0085] Figure 5 is a schematic flowchart of the third embodiment of the soil heavy metal content prediction method of the present invention;
[0086] Figure 6 is a schematic flowchart of the fourth embodiment of the soil heavy metal content prediction method of the present invention;
[0087] Figure 7 is a schematic block diagram of the first embodiment of the soil heavy metal content prediction device of the present invention.
[0088] The implementation, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0089] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0090] Refer to Figure 1 ,Figure 1 Schematic diagram of the structure of a soil heavy metal content prediction device for the hardware operating environment involved in the solution of the embodiment of the present invention.
[0091] As shown in Figure 1 , the soil heavy metal content prediction device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0092] Those skilled in the art can understand that Figure 1 the structure shown in
[0093] does not constitute a limitation on the soil heavy metal content prediction device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 1 As shown in
[0094] In Figure 1 , in the soil heavy metal content prediction device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the soil heavy metal content prediction device of the present invention may be arranged in the soil heavy metal content prediction device. The soil heavy metal content prediction device calls the soil heavy metal content prediction program stored in the memory 1005 through the processor 1001 and executes the soil heavy metal content prediction method provided by the embodiment of the present invention.
[0095] The embodiment of the present invention provides a method for predicting soil heavy metal content. Referring to Figure 2 , Figure 2Schematic flowchart of the first embodiment of the method for predicting soil heavy metal content according to the present invention.
[0096] In this embodiment, the method for predicting soil heavy metal content includes the following steps:
[0097] Step S10: Obtain an initial soil data set.
[0098] It should be noted that the execution subject of this embodiment is a soil heavy metal content prediction device, and it can also be other devices that can achieve the same or similar functions. This embodiment does not limit this, and this embodiment is described by taking the soil heavy metal content prediction device as an example.
[0099] In specific implementation, the initial soil data set refers to the soil data collected in a preset area and preliminarily processed. The preset area can be selected according to user needs, and this embodiment does not limit this.
[0100] Further, the step of obtaining the initial soil data set specifically includes: obtaining an original soil data set; performing data cleaning on the original soil data set to obtain abnormal data in the original soil data set; and removing the abnormal data from the original soil data set to obtain the initial soil data set.
[0101] The original soil data set can be directly collected. Since there may be dirty data or useless data in the original soil data set, the original soil data is cleaned to obtain abnormal data in the original soil data set. The abnormal data includes dirty data and useless data. Then, the abnormal data is removed from the original soil data set to obtain the cleaned data set, that is, the initial soil data set.
[0102] Step S20: Cluster the initial soil data set to obtain target clustering clusters.
[0103] It should be noted that the K-means (k-means clustering algorithm) algorithm can be used to cluster the initial soil data set. By using the K-means algorithm to cluster the initial soil data set, the clustering centers are selected from the initial soil data set, so as to obtain the target clustering clusters. The target clustering clusters refer to several clustering centers obtained after clustering the data in the initial soil data set.
[0104] Step S30: Obtain a reference soil data set through the target clustering clusters.
[0105] In specific implementation, after obtaining the target clustering clusters, the data corresponding to the target clustering clusters can be used as the reference soil data set.
[0106] Step S40: Filter the reference soil dataset through an entropy weight distance detection strategy to obtain a target soil dataset.
[0107] It should be understood that the entropy weight distance detection strategy refers to detecting the degree of outliers of each data point in the reference soil dataset, and filtering the outliers in the reference soil dataset through the degree of outliers to obtain a target soil dataset. The target soil dataset refers to a non-outlier data set, and soil data with similar attributes can be obtained.
[0108] Step S50: Optimize the recurrent neural network through the target soil dataset to obtain a target recurrent neural network.
[0109] In a specific implementation, the recurrent neural network refers to an Elman neural network. The Elman neural network includes an input layer, an output layer, a context layer, and a hidden layer. The number of neurons in the input layer is the same as the dimension of the data features, and the number of neurons in the output layer is the same as the dimension of the output data labels. The number of neurons in the hidden layer is not fixed. When the number of neurons in the hidden layer is too large, the network training speed will be greatly reduced; when the number of neurons in the hidden layer is too small, the learning amplitude of the network will be reduced, so that learning cannot be completed. The main function of the context layer is to use the hidden layer state at the previous time point as the input of the hidden layer at the next time point through connection memory, which is equivalent to playing a role of state feedback. The calculation formula for the number of hidden layer neurons is as follows in Equation 1:
[0110]
[0111] In Equation 1, the number of neurons in the hidden layer is k, α and β are the number of neuron nodes in the input layer and the output layer respectively, and γ is a random number between 1 and 10.
[0112] It should be noted that optimizing the recurrent neural network refers to using the Bayesian regularization method to optimize the structure of the recurrent neural network, and using the GA-PSO algorithm to optimize the network error of the recurrent neural network, so as to obtain a target recurrent neural network. The target recurrent neural network refers to the optimized recurrent neural network.
[0113] Step S60: Train the target recurrent neural network through the target soil dataset to obtain a soil heavy metal content prediction model.
[0114] In this embodiment, after obtaining the target recurrent neural network, the target soil dataset can be input into the target recurrent neural network for training, which can accelerate the network training convergence speed, so as to obtain a soil heavy metal content prediction model. An optimized soil heavy metal content prediction model can be obtained through the optimized recurrent neural network and the filtered target soil dataset.
[0115] Step S70: Predict the soil heavy metal content through the soil heavy metal content prediction model.
[0116] It should be understood that after obtaining the soil heavy metal content prediction model, the soil heavy metal content can be predicted through the soil heavy metal content prediction model, improving the accuracy and precision of the prediction.
[0117] As Figure 3 shown, Figure 3 This is a schematic diagram of the overall process of the soil heavy metal content prediction method in this embodiment. After obtaining the initial soil data set, the initial soil data set is clustered by the K-means algorithm to obtain the target clustering clusters. The corresponding data sets can be obtained through the target clustering clusters, and the data sets are processed using the entropy weight distance-based detection method to eliminate the outliers in the data sets, obtaining the processed data, i.e., the target soil data set. The structure of the recurrent neural network is selected, and the processed data is input into the recurrent neural network. The individual fitness of the processed data is calculated using the adaptive genetic algorithm, and individuals are selected using the roulette wheel rule. Adaptive crossover and mutation operations are performed on the individuals, and the update probability is calculated simultaneously. Whether to continue the crossover and mutation is judged through the update probability value. If the update probability value meets the preset value, the crossover and mutation are stopped, and the current population is output. The current population is initialized using the adaptive particle swarm algorithm to obtain the initial velocity and initial position, and the individual fitness value is calculated to obtain the individual extreme value and the global extreme value. The individual extreme value and the global extreme value are updated, and at the same time, the particle velocity and position are updated to obtain the particle update velocity. Local optimization and global training are performed through the particle update velocity. When the training requirements are met, the updated network weights and updated thresholds are determined, thereby updating the network weights and thresholds, and optimizing the network error of the recurrent neural network to obtain the optimized recurrent neural network, i.e., the target recurrent neural network. The soil heavy metal content prediction model is obtained by training the processed data according to the target recurrent neural network, and thus the heavy metals in the soil are detected according to the soil heavy metal content prediction model, improving the detection accuracy.
[0118] In this embodiment, by obtaining the initial soil data set; clustering the initial soil data set to obtain the target clustering clusters; obtaining the reference soil data set through the target clustering clusters; filtering the reference soil data set through the entropy weight distance detection strategy to obtain the target soil data set; optimizing the recurrent neural network through the target soil data set to obtain the target recurrent neural network; training the target recurrent neural network through the target soil data set to obtain the soil heavy metal content prediction model; predicting the soil heavy metal content through the soil heavy metal content prediction model, the network training convergence speed is accelerated, thereby improving the prediction accuracy of the soil heavy metal content.
[0119] Reference Figure 4, Figure 4 It is a schematic flowchart of the second embodiment of the method for predicting soil heavy metal content of the present invention.
[0120] Based on the above first embodiment, step S40 of the method for predicting soil heavy metal content in this embodiment specifically includes:
[0121] Step S401: Obtain the attributes of each data in the reference soil dataset and the dimension of the reference soil dataset.
[0122] It should be noted that the reference soil dataset carries the attributes of each data, and the dimension of the reference soil dataset can be obtained.
[0123] Step S402: Combine the attributes of the each data to obtain an attribute set.
[0124] In this embodiment, when the attributes of the reference soil dataset are obtained, the attributes can be combined to obtain an attribute set A, and the attribute set A = {A1, A2... A d} where d is the dimension of the dataset.
[0125] Step S403: Obtain the value set of each attribute in the attribute set and the value probability of each attribute.
[0126] It can be understood that when the attribute set is obtained, the value set S(A i ) of each attribute in the attribute set can be obtained, and the value probability P(x j ) of each attribute can be obtained. P(x j ) is the proportion of the jth data under the attribute A i .
[0127] Step S404: Calculate the entropy value through the value set of each attribute and the value probability of each attribute.
[0128] In a specific implementation, when the value set S(A i ) of each attribute and the value probability P(x j ) of each attribute are obtained, the information entropy value can be calculated, and the calculation process is as follows in Equation 2:
[0129]
[0130] In Equation 2, H(A i ) is the entropy value, S(A i ) is the value set of the attribute A i , P(x j ) is the proportion of the jth data under the attribute A i , that is, the attribute value probability.
[0131] Step S405: Obtain the outlier attributes and non-outlier attributes in the reference soil dataset based on the entropy value and the dimension of the reference soil dataset.
[0132] When the calculated entropy value H(A i ) is obtained, it is possible to determine whether each point in the reference soil dataset is an outlier through H(A i ) and the dimension d, and obtain the outlier attributes or non-outlier attributes of each point. The calculation is as follows in Equation 3:
[0133]
[0134] In Equation 3, H(A i ) is the entropy value, d is the data dimension. When the entropy value H(A i ) satisfies Equation 3 above, the attribute A i is called an outlier attribute. When the entropy value H(A i ) does not satisfy Equation 3 above, the attribute A i is called a non-outlier attribute.
[0135] Step S406: Set the attribute weights of the outlier attributes and the non-outlier attributes respectively.
[0136] It can be understood that after obtaining the outlier attributes and non-outlier attributes in the reference soil dataset, the attribute weights of the outlier attributes and non-outlier attributes can be set, and the setting is as follows in Equation 4:
[0137]
[0138] In Equation 4, q > 1. When the data point is an outlier attribute, the attribute weight is set to q. When the data point is a non-outlier attribute, the attribute weight is set to 1.
[0139] Step S407: Obtain the values of the data objects in the reference soil dataset on the corresponding attributes.
[0140] Step S408: Calculate the entropy weight distance through the values on the corresponding attributes and the attribute weights.
[0141] In a specific implementation, it is possible to obtain the values of the data objects in the reference soil dataset on the corresponding attributes, and calculate the entropy weight distance through the values, so as to judge the outlier points according to the entropy weight distance. The calculation process of the values on the corresponding attributes and the attribute weights is as follows in Equation 5:
[0142]
[0143] f Ai (a) and f Ai (b) are the values of object a and object b on attribute A i respectively, and θ iis the attribute weight, θ i The values are 1, 2, 3, 4, 5... d. It can be seen that when the attribute weight value of the outlier attribute is relatively large, the outlier degree of the outlier can be better represented, so as to improve the ability to distinguish outliers from inliers.
[0144] Step S409: When the entropy weight distance is greater than the preset distance threshold, use the data corresponding to the entropy weight distance as outlier data.
[0145] In a specific implementation, the preset distance threshold can be set according to requirements, such as 0.5, 0.6, 0.8, etc., and this embodiment does not limit this. When the entropy weight distance is greater than the preset distance threshold, outliers can be obtained. Use the data corresponding to the entropy weight distance greater than the preset distance threshold as outlier data.
[0146] Step S410: Remove the outlier data from the reference soil dataset to obtain the target soil dataset.
[0147] It can be understood that after obtaining the outlier data, the outlier data can be removed to obtain the data after removing the outliers, that is, the target soil dataset. By removing the outliers, a more accurate data fitting effect can be achieved.
[0148] In this embodiment, by obtaining the attributes of each data in the reference soil dataset and the dimension of the reference soil dataset; combining the attributes of the each data to obtain an attribute set; obtaining the value set of each attribute and the probability of each attribute value in the attribute set; calculating the entropy value through the value set of each attribute and the probability of each attribute value; obtaining the outlier attributes and non-outlier attributes in the reference soil dataset through the entropy value and the dimension of the reference soil dataset; respectively setting the attribute weights of the outlier attributes and the non-outlier attributes; obtaining the values of the data objects in the reference soil dataset on the corresponding attributes; calculating the entropy weight distance through the values on the corresponding attributes and the attribute weights; when the entropy weight distance is greater than the preset distance threshold, using the data corresponding to the entropy weight distance as outlier data; removing the outlier data from the reference soil dataset to obtain the target soil dataset. Through the above method, the outlier data can be removed from the reference soil dataset, and a more accurate data fitting effect can be achieved by removing the outlier data.
[0149] Reference Figure 5 , Figure 5 is the flow chart of the third embodiment of the soil heavy metal content prediction method of the present invention.
[0150] Based on the above first embodiment, step S50 of the soil heavy metal content prediction method in this embodiment specifically includes:
[0151] Step S501: Obtain the quantity and actual value of the target soil dataset.
[0152] It should be noted that the quantity of the target soil dataset is the total number N of the trained target soil datasets, and the actual value refers to the actual value of the trained target soil datasets.
[0153] Step S502: Obtain the predicted value of the recurrent neural network.
[0154] In a specific implementation, the predicted value of the recurrent neural network is Predicted value which is calculated by the recurrent neural network model.
[0155] Step S503: Calculate the mean square error function from the quantity, the actual value, and the predicted value.
[0156] In this embodiment, the calculation process of the mean square error function is as follows in Equation 6:
[0157]
[0158] In Equation 6, E1 is the mean square error function, N is the quantity of the target soil dataset, is the actual value of the target soil dataset, is the predicted value of the target soil dataset.
[0159] Step S504: Set the weights in the target soil dataset.
[0160] Step S505: Calculate the weight decay term according to the weights.
[0161] It should be understood that the weight in the target soil dataset is ω i , and after setting the weights, the weight decay term can be calculated according to the weights, the quantity of the target soil number set, and the weights. The calculation process is as follows in Equation 7:
[0162]
[0163] In Equation 7, E2 is the weight decay term, ω i is the i-th weight.
[0164] Step S506: Obtain the Jacobian matrix and Hessian matrix of the target soil dataset.
[0165] It should be noted that the Jacobian matrix of the target soil dataset is J, and the Hessian matrix is The Hessian matrix is the matrix of the objective function at the point ω min , and since the calculation amount of this matrix is too large, the Gauss-Newton approximation method is used for replacement, and the following Equation 8 can be obtained:
[0166]
[0167] In Equation 8, Q is the Hessian matrix, the Jacobian matrix is J, and I N is the identity matrix.
[0168] Step S507: Calculate the number of effective weight values through the Jacobian matrix, the Hessian matrix, and the quantity.
[0169] In a specific implementation, when the Hessian matrix and the Jacobian matrix are obtained according to the above Equation 8, the number of effective weight values can be calculated based on the Jacobian matrix, the Hessian matrix, and the quantity of the target soil dataset The calculation process is as follows in Equation 9:
[0170]
[0171] In Equation 9, is the number of effective weight values, N ω is the quantity of the target soil dataset, and Q is the Hessian matrix. By calculating through the above Equation 9, the number of effective weight values is obtained, and the number of effective weight values indicates how many parameters play a role in reducing the total error function during network training.
[0172] Step S508: Calculate the regularization parameter through the number of effective weight values, the weight decay term, the quantity, and the mean square error function.
[0173] In a specific implementation, calculate the regularization parameter through the number of effective weight values the weight decay term E2, the quantity of the target soil dataset, and the mean square error function E1. The calculation process is as follows in Equation 10:
[0174]
[0175] In Equation 10, a and b are the regularization parameters, is the number of effective weight values, E1 is the mean square error function, E2 is the weight decay term, and N is the quantity of the target soil dataset. By calculating through the above Equation 10, the regularization parameters a and b can be obtained.
[0176] Step S509: Calculate the objective function through the regularization parameter, the weight decay term, and the mean square error function.
[0177] It should be noted that the calculation process of the objective function is as follows in Equation 11:
[0178] J ω = aE2 + bE1 (Equation 11)
[0179] In Equation 11, J ωLet \(J\) be the objective function, \(a\) and \(b\) be the regularization parameters, \(E1\) be the mean squared error function, and \(E2\) be the weight decay term. Through Equation (11) above, the objective function is calculated. The objective function is the objective function with a "penalty term" added.
[0180] Step S510: Optimize the structure of the recurrent neural network according to the objective function and the regularization parameter to obtain the structure of the target recurrent neural network.
[0181] In this embodiment, the structure of the recurrent neural network can be optimized by the objective function and the regularization parameter to obtain the optimized structure of the recurrent neural network, that is, the structure of the target recurrent neural network.
[0182] Step S511: Obtain the network weights and thresholds of the structure of the target recurrent neural network.
[0183] It should be understood that after obtaining the structure of the target recurrent neural network, the network weights and thresholds of the structure of the target recurrent neural network can be set, and the network weights and thresholds of the structure of the target recurrent neural network can be obtained.
[0184] Step S512: Update the network weights and the thresholds through the adaptive GA-PSO algorithm to optimize the network error of the recurrent neural network and obtain the target recurrent neural network.
[0185] It should be noted that the adaptive GA-PSO algorithm includes an adaptive genetic algorithm and an adaptive particle swarm algorithm. The network weights and thresholds can be calculated and updated through the adaptive genetic algorithm and the adaptive particle swarm algorithm, so as to optimize the network error of the recurrent neural network and obtain the target recurrent neural network.
[0186] Further, updating the network weights and thresholds through the adaptive genetic algorithm includes: initializing a population through the adaptive genetic algorithm, where the population includes a preset number of individuals, and the individuals are a set of parameters of network weights and thresholds; calculating the individual fitness value of each individual; obtaining the maximum value of the individual fitness values and the average value of the individual fitness; setting the crossover base probability, mutation base probability, crossover constant, and mutation constant, where the crossover base probability is less than the crossover constant, and the mutation base probability is less than the mutation constant; performing crossover and mutation on the individuals through the individual fitness value, the maximum value, the average value, the crossover base probability, the crossover constant, the mutation base probability, and the mutation constant; obtaining the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value according to the individual fitness value; calculating an update probability through the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value; determining whether the individuals perform crossover and mutation according to the update probability; when the update probability is equal to a preset probability value, outputting the current population; and using the adaptive particle swarm algorithm to calculate the current population to update the network weights and the thresholds to obtain a target recurrent neural network.
[0187] It should be understood that the population can be initialized through the adaptive genetic algorithm. The population includes a preset number of individuals, and the preset number can be set according to requirements. The individuals are a set of parameters of network weights and thresholds. After initializing the population, the individual fitness value of each individual can be calculated, and the individuals can be selected using the roulette wheel rule. In the initial stage of population iteration, the crossover and mutation probabilities of the population individuals are appropriately reduced to increase the proportion of individuals with high fitness values in the population. In the later stage of population iteration, the differences between individuals have been greatly reduced. To avoid falling into a local optimal solution, the crossover and mutation probabilities of the population individuals can be appropriately increased. To accelerate the search for the global optimal solution, the crossover and mutation operations on individuals with high fitness values should be reduced, and the crossover and mutation probabilities of individuals with lower fitness values should be increased.
[0188] In a specific implementation, the maximum value of the individual fitness value can be obtained according to the individual fitness value, and the average value of the individual fitness can be calculated. At the same time, the crossover base probability, mutation base probability, crossover constant, and mutation constant are set. The crossover constant and mutation constant are random values. Among them, the crossover base probability is greater than 0 and less than the crossover constant, the crossover constant is less than 1, the mutation base probability is greater than 0 and less than the mutation constant, and the mutation constant is less than 1. The crossover probability and genetic probability can be calculated through the average value of the individual fitness, the maximum value of the individual fitness value, the crossover base probability, the mutation base probability, the crossover constant, and the mutation constant, so as to perform crossover and mutation on the individuals. The calculation process is as shown in Equation 12 below:
[0189]
[0190] In Equation 12, P ci is the crossover probability, P mi is the mutation probability, c1 and c2 are the basic crossover probabilities, c3 is the crossover constant, 0 < c1 < c2 < c3 < 1, c4 and c5 are the basic mutation probabilities, c6 is the mutation constant, 0 < c4 < c5 < c6 < 1, f i is the individual fitness value, f max is the maximum value of the individual fitness value, f avg is the average value of the individual fitness. As can be seen from Equation 11 above, when the fitness value of an individual is less than the average value, the probabilities of crossover and mutation can be appropriately higher than the basic crossover probability and the basic mutation probability. When the fitness value of an individual is too high, the probabilities of crossover and mutation will be appropriately reduced. This can avoid the situation where the fitness value of an individual continuously decreases during the iteration process. In the process of natural evolution, it is necessary to ensure that excellent genes are inherited to the next generation. Therefore, in the adaptive genetic algorithm, it is also necessary to ensure that the genes of individuals with higher fitness values are inherited to the next generation, so as to reduce the risk of falling into the local optimal solution. Therefore, the update probability can be calculated, and it is determined whether an individual needs to perform crossover and mutation according to the update probability. Then, the first individual fitness value and the second individual fitness value can be obtained according to the individual fitness value. The first individual fitness value is the fitness value of the new individual, and the second individual fitness value is the fitness value of the old individual. And the number of consecutive decreases in the individual fitness value is obtained according to the individual fitness, and the update probability is calculated through the first individual fitness value, the second individual fitness value, and the number of consecutive decreases. The calculation process is as follows in Equation 13:
[0191]
[0192] In Equation 13, P is the update probability, f n is the fitness value of the new individual, f o is the fitness value of the old individual, T is the number of consecutive decreases in the individual fitness value, and the initial value is 0. As can be seen from Equation 13 above, P gradually decreases as the number of consecutive decreases increases. When the number of consecutive decreases in the fitness value is too large, the value of the update probability P is 0. At this time, the individual will not perform crossover and mutation operations, and the excellent genes of the individual can be retained, thus avoiding the risk of falling into the local optimum. It can be judged whether an individual performs crossover and mutation according to the update probability. The preset probability value can be set according to requirements, such as 0, 0.1, etc. This embodiment does not limit this, and this embodiment is described by taking 0 as an example. When the update probability is 0, the current population is output, and the adaptive particle swarm algorithm is used to calculate the current population, so as to update the network weights and thresholds to obtain the target recurrent neural network.
[0193] Further, the steps of calculating the current population by the adaptive particle swarm optimization algorithm and updating the network weights and thresholds to obtain the target recurrent neural network specifically include: initializing the current population by the adaptive particle swarm optimization algorithm, setting the initial position and the initial velocity of the current population; calculating the individual fitness values in the current population to obtain the individual extreme value and the global extreme value; setting the maximum number of iterations and obtaining the current number of iterations; setting the position parameter according to the maximum number of iterations; setting the first inertia weight, the second inertia weight, the first learning factor, the second learning factor, and a random value, where the first inertia weight is less than the second inertia weight, and the first learning factor is less than the second learning factor; calculating the current inertia weight by the first inertia weight, the second inertia weight, the position parameter, and the current number of iterations; calculating the current learning factor by the first learning factor, the second learning factor, the maximum number of iterations, and the current number of iterations; calculating the compression factor by the current learning factor; calculating the particle update velocity by the compression factor, the current inertia weight, the current learning factor, the random value, the individual extreme value, the global extreme value, the initial velocity, and the initial position; performing local optimization and global optimization training according to the particle update velocity, and when the training requirements are met, determining to update the network weights and update the thresholds; optimizing the network error of the recurrent neural network according to the updated network weights and the updated thresholds to obtain the target recurrent neural network.
[0194] It should be noted that in the traditional particle swarm algorithm, the value of the inertia weight seriously affects the motion state of the particles and the global optimization ability of the algorithm. When the inertia weight is large, it is difficult for the particles to change their motion state, and the search position space will also increase accordingly. At this time, the local optimization ability of the algorithm is weak, and the convergence speed of the population will also decrease; when the inertia weight is small, the particles can easily change their motion state, but the search position space will be reduced, which easily leads to the particles being unable to obtain the current optimal solution, and the convergence speed of the population will also be reduced. In this embodiment, the Sigmoid activation function is introduced into the PSO (Particle Swarm Optimization) algorithm to adjust the change of the inertia weight value. The Sigmoid activation function curve grows relatively slowly at both ends and grows relatively fast at the center of the function. The current population can be initialized through the adaptive particle swarm algorithm, the initial position and initial velocity of the current population can be set, and the individual fitness values in the current population can be calculated to obtain the individual extreme value and the global extreme value. When the individual extreme value and the global extreme value are obtained, the maximum number of iterations can be set and the current number of iterations can be obtained at the same time. The position parameter is set through the maximum number of iterations, and the position parameter is set to half of the maximum number of iterations of the population. The first inertia weight, the second inertia weight, the first learning factor, the second learning factor, and the random value are set. The first inertia weight is less than the second inertia weight. The first inertia weight is the minimum value of the inertia weight, and the second inertia weight is the maximum value of the inertia weight. To further ensure that the value of the current inertia weight ω is within (ω min , ω max ), the current inertia weight can be constructed based on the Sigmoid activation function. The current inertia weight is calculated through the first inertia weight, the second inertia weight, the position parameter, and the current number of iterations. The calculation process is as follows in Equation 14:
[0195]
[0196] In Equation 14 above, ω is the current inertia weight, ω min is the first inertia weight, ω max is the second inertia weight, t is the position parameter, k is the current number of iterations. Through Equation 14, the current inertia weight can be calculated. In a specific implementation, the minimum inertia weight value is 0.3, and the maximum inertia weight value is 0.9. According to Equation 14, in the initial stage of population iteration, the current inertia weight ω of individuals with fewer iterations is larger. At this time, the focus is on the global optimization ability. In the middle stage of population iteration, the number of iterations is half of the maximum number of iterations, and the current inertia weight ω is (ω min + ω max ) / 2. At this time, the individual is in the transition stage between local optimization and global optimization. In the later stage of population iteration, the current inertia weight ω of the individual approaches ω min, the individual focuses on local optimization to complete the entire optimization process.
[0197] In a specific implementation, the current learning factor can be calculated based on the first learning factor, the second learning factor, the position parameter, and the current iteration number. The calculation process is as follows in Equation 15:
[0198]
[0199] In Equation 15, c min is the first learning factor, c max is the second learning factor. The first learning factor is the minimum value of the learning factor, with a value of 2. The second learning factor is the maximum value of the learning factor, with a value of 3.5. The first learning factor is less than the second learning factor. c1 = c2 is the current learning factor, M is the maximum number of iterations, k is the current iteration number. When the current learning factor c1 is relatively large, the individual will linger in the local range too much. When the current learning factor c2 is relatively large, the individual is prone to falling into the local minimum. In order to effectively control the speed of the individual and enable the population to achieve a balance between local optimization and global optimization.
[0200] In this embodiment, after calculating the current learning factor, the compression factor can be calculated based on the current learning factor. The calculation process is as follows in Equation 16:
[0201]
[0202] In Equation 16, a is the compression factor, C is the sum of the current learning factors, C = c1 + c2, c1 = c2 ≥ 2. After calculating the compression factor, the particle update speed can be calculated through the compression factor, the current inertia weight, the current learning factor, the random value, the individual extreme value, the global extreme value, the initial speed, and the initial position. The calculation process is as follows in Equation 17:
[0203] v i+1 = a{ωv i + c1 × rand() × (pbest i - x i ) + c2 × rand() × (gbest i - x i )}
[0204] (Equation 17)
[0205] In Equation 17, v i+1 is the particle update speed, a is the compression factor, ω is the current inertia weight, v i is the initial speed, x i is the initial position, c1 and c2 are the current learning factors, rand() is the random value, the random value is between 0 and 1, pbest i is the individual extreme value, gbesti is the global extreme value, which is calculated by Equation 17 above to obtain the particle update velocity. Through the adjustment of the compression factor a, the convergence of PSO is ensured, and the boundary limit of the example velocity is eliminated. Local optimization and global optimization training are performed through the example update velocity to determine whether the individual meets the training requirements. When the training requirements are met, the updated network weights and updated thresholds are obtained, so that the network error can be calculated according to the updated network weights and updated thresholds, and the network error of the recursive neural network is optimized to obtain the optimized recursive neural network, that is, the target recursive neural network. The target soil data can be trained according to the target recursive neural network to obtain a soil heavy metal content prediction model, and the soil heavy metal content can be detected according to the obtained model to improve the detection accuracy.
[0206] In this embodiment, the number and actual value of the target soil data set are obtained; the predicted value of the recursive neural network is obtained; the mean square error function is calculated through the number, the actual value, and the predicted value; the weights in the target soil data set are set; the weight decay term is calculated according to the weights; the Jacobian matrix and Hessian matrix of the target soil data set are obtained; the effective number of weights is calculated through the Jacobian matrix, the Hessian matrix, and the number; the regularization parameter is calculated through the effective number of weights, the weight decay term, the number, and the mean square error function; the objective function is calculated through the regularization parameter, the weight decay term, and the mean square error function; the structure of the recursive neural network is optimized according to the objective function and the regularization parameter to obtain the structure of the target recursive neural network; the network weights and thresholds of the structure of the target recursive neural network are obtained; the network weights and the thresholds are updated through the adaptive GA-PSO algorithm to optimize the network error of the recursive neural network to obtain the target recursive neural network. The structure and network error of the recursive neural network can be optimized according to the adaptive GA-PSO algorithm to obtain an optimized recursive neural network, so that the sample data can be trained according to the optimized recursive neural network to accelerate the convergence speed of network training.
[0207] Reference Figure 6 , Figure 6 is a schematic flowchart of the fourth embodiment of the soil heavy metal content prediction method of the present invention.
[0208] Based on the above first embodiment, step S20 of the soil heavy metal content prediction method in this embodiment specifically includes:
[0209] Step S201: Select a target point from the initial soil data set and use the target point as the initial clustering center.
[0210] It should be noted that the target point can be a point randomly selected from the initial soil dataset. The initial clustering center refers to the first clustering center, and a randomly selected point is used as the first clustering center.
[0211] Step S202: Calculate the first distance from each point in the initial soil dataset to the initial clustering center.
[0212] In a specific implementation, the first distance refers to the distance from each point in the initial soil dataset to the initial clustering center, and the calculation process is as follows in Equation 18:
[0213]
[0214] In Equation 18, D(x) is the first distance, x i is a point in the initial soil dataset D, and P r is the initial clustering center.
[0215] Step S203: Screen the initial clustering center through the first distance to obtain the target clustering center.
[0216] It should be understood that after calculating the first distance D(x) between each point in the initial soil dataset and the initial clustering center, a new point is selected as the new clustering center according to the first distance D(x). Points with larger distances are screened from the first distance, and the points with larger distances are used as the new clustering center. Then, the distances between each point in the initial soil dataset and the new clustering center are calculated again, and points with smaller distances are selected from the new clustering center for calculation until k clustering centers, that is, the target clustering centers, are selected.
[0217] Step S204: Divide the initial soil dataset based on the target clustering center to obtain the target clustering clusters.
[0218] In this embodiment, after obtaining the target clustering center, the initial soil dataset D is divided into k clustering clusters, denoted as and at the same time, the Euclidean distances from each point x i in the initial soil dataset D to each clustering center p j {j = 1, 2,... k} in the target clustering center are calculated, denoted as d ij , and the dataset x ij corresponding to the minimum value in d i is classified into the clustering cluster C j , and the updated clustering center p j in the clustering cluster C j is recalculated, and the calculation process is as follows in Equation 19:
[0219]
[0220] In Equation 19, p j is the updated cluster center, C j is the cluster, x is the point in the initial soil dataset. If all the updated clusters p j remain unchanged, then output the updated clusters C = {C1, C2... C k}, that is, the target clusters. When calculating the Euclidean distances from each point x i in the initial soil dataset D to each cluster center p j {j = 1, 2,... k} in the target cluster centers, the amount of work is extremely large. The triangle property can be introduced to greatly reduce the amount of calculation. Calculate the distance between the cluster centers p j and p j+1 , denoted as d(p j , p j+1 ). The distance from the data x i to the cluster center p j is denoted as d(x j , p j ). The distance from the data x i to the cluster center p j+1 is denoted as d(x j , p j+1 ). When 2d(x j , p j ) ≤ d(p j , p j+1 ), then d(x j , p j ) ≤ d(p j , p j+1 ). Thus, the amount of calculation of the Euclidean distances from each point x i in the initial soil dataset D to each cluster center p j {j = 1, 2,... k} in the target cluster centers is reduced.
[0221] In this embodiment, by selecting a target point from the initial soil dataset and using the target point as the initial cluster center; calculating the first distance from each point in the initial soil dataset to the initial cluster center; screening the initial cluster center through the first distance to obtain the target cluster center; and partitioning the initial soil dataset based on the target cluster center to obtain the target clusters. By performing clustering partitioning on the initial soil dataset through the k - means algorithm, the points in the initial soil dataset can be quickly clustered, facilitating subsequent filtering of outliers.
[0222] Refer to Figure 7 , Figure 7 which is the structural block diagram of the first embodiment of the soil heavy metal content prediction device of the present invention.
[0223] As Figure 7 shown, the soil heavy metal content prediction device proposed in the embodiment of the present invention includes:
[0224] An acquisition module 10, configured to acquire an initial soil data set.
[0225] A clustering module 20, configured to cluster the initial soil data set to obtain target clustering clusters.
[0226] The acquisition module 10 is further configured to obtain a reference soil data set through the target clustering clusters.
[0227] A filtering module 30, configured to filter the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set.
[0228] An optimization module 40, configured to optimize a recurrent neural network through the target soil data set to obtain a target recurrent neural network.
[0229] A training module 50, configured to train the target recurrent neural network through the target soil data set to obtain a soil heavy metal content prediction model.
[0230] A prediction module 60, configured to predict the soil heavy metal content through the soil heavy metal content prediction model.
[0231] In this embodiment, by acquiring an initial soil data set; clustering the initial soil data set to obtain target clustering clusters; obtaining a reference soil data set through the target clustering clusters; filtering the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set; optimizing a recurrent neural network through the target soil data set to obtain a target recurrent neural network; training the target recurrent neural network through the target soil data set to obtain a soil heavy metal content prediction model; and predicting the soil heavy metal content through the soil heavy metal content prediction model, the network training convergence speed is accelerated, thereby improving the prediction accuracy of the soil heavy metal content.
[0232] In one embodiment, the filtering module 30 is further configured to obtain the attributes of each data in the reference soil dataset and the dimension of the reference soil dataset; combine the attributes of each data to obtain an attribute set; obtain the value set and the value probability of each attribute in the attribute set; calculate the entropy value through the value set of each attribute and the value probability of each attribute; obtain the outlier attributes and non-outlier attributes in the reference soil dataset based on the entropy value and the dimension of the reference soil dataset; respectively set the attribute weights of the outlier attributes and the non-outlier attributes; obtain the values of the data objects in the reference soil dataset on the corresponding attributes; calculate the entropy weight distance through the values on the corresponding attributes and the attribute weights; when the entropy weight distance is greater than a preset distance threshold, use the data corresponding to the entropy weight distance as outlier data; remove the outlier data from the reference soil dataset to obtain a target soil dataset.
[0233] In one embodiment, the optimization module 40 is further configured to obtain the quantity and the actual value of the target soil dataset; obtain the predicted value of the recurrent neural network; calculate the mean square error function through the quantity, the actual value, and the predicted value; set the weights in the target soil dataset; calculate the weight decay term according to the weights; obtain the Jacobian matrix and the Hessian matrix of the target soil dataset; calculate the number of effective weights through the Jacobian matrix, the Hessian matrix, and the quantity; calculate the regularization parameter through the number of effective weights, the weight decay term, the quantity, and the mean square error function; calculate the objective function through the regularization parameter, the weight decay term, and the mean square error function; optimize the structure of the recurrent neural network according to the objective function and the regularization parameter to obtain the structure of the target recurrent neural network; obtain the network weights and thresholds of the structure of the target recurrent neural network; update the network weights and the thresholds through the adaptive GA-PSO algorithm to optimize the network error of the recurrent neural network to obtain the target recurrent neural network.
[0234] In one embodiment, the optimization module 40 is further configured to initialize a population through an adaptive genetic algorithm. The population includes a preset number of individuals, where each individual is a set of parameters of network weights and thresholds. Calculate the individual fitness value of each individual; obtain the maximum value of the individual fitness value and the average value of the individual fitness; set the crossover base probability, mutation base probability, crossover constant, and mutation constant, where the crossover base probability is less than the crossover constant, and the mutation base probability is less than the mutation constant; perform crossover and mutation on the individuals through the individual fitness value, the maximum value, the average value, the crossover base probability, the crossover constant, the mutation base probability, and the mutation constant; obtain the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value according to the individual fitness value; calculate an update probability through the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value; determine whether the individual performs crossover and mutation according to the update probability; when the update probability is equal to a preset probability value, output the current population; use an adaptive particle swarm algorithm to calculate the current population to update the network weights and the thresholds to obtain a target recurrent neural network.
[0235] In one embodiment, the optimization module 40 is further configured to initialize the current population through an adaptive particle swarm algorithm, and set the initial position and initial velocity of the current population; calculate the individual fitness value in the current population to obtain the individual extreme value and the global extreme value; set the maximum number of iterations and obtain the current number of iterations; set a position parameter according to the maximum number of iterations; set a first inertia weight, a second inertia weight, a first learning factor, a second learning factor, and a random value, where the first inertia weight is less than the second inertia weight, and the first learning factor is less than the second learning factor; calculate the current inertia weight through the first inertia weight, the second inertia weight, the position parameter, and the current number of iterations; calculate the current learning factor according to the first learning factor, the second learning factor, the maximum number of iterations, and the current number of iterations; calculate a compression factor through the current learning factor; calculate the particle update velocity according to the compression factor, the current inertia weight, the current learning factor, the random value, the individual extreme value, the global extreme value, the initial velocity, and the initial position; perform local optimization and global optimization training according to the particle update velocity, and when the training requirements are met, determine the updated network weights and the updated thresholds; optimize the network error of the recurrent neural network according to the updated network weights and the updated thresholds to obtain a target recurrent neural network.
[0236] In one embodiment, the clustering module 20 is further configured to select target points from the initial soil dataset, and use the target points as initial clustering centers; calculate the first distances from each point in the initial soil dataset to the initial clustering centers; screen the initial clustering centers through the first distances to obtain target clustering centers; and divide the initial soil dataset based on the target clustering centers to obtain target clustering clusters.
[0237] In one embodiment, the acquisition module 10 is further configured to acquire an original soil dataset; perform data cleaning on the original soil dataset to obtain abnormal data in the original soil dataset; and remove the abnormal data from the original soil dataset to obtain an initial soil dataset.
[0238] In addition, to achieve the above object, the present invention further provides a soil heavy metal content prediction device, which includes: a memory, a processor, and a soil heavy metal content prediction program stored on the memory and executable on the processor. The soil heavy metal content prediction program is configured to implement the steps of the soil heavy metal content prediction method as described above.
[0239] Since this soil heavy metal content prediction device adopts all the technical solutions of the above all embodiments, it at least has all the beneficial effects brought by the technical solutions of the above embodiments, which will not be elaborated herein one by one.
[0240] In addition, an embodiment of the present invention further provides a storage medium, on which a soil heavy metal content prediction program is stored. When the soil heavy metal content prediction program is executed by a processor, it implements the steps of the soil heavy metal content prediction method as described above.
[0241] Since this storage medium adopts all the technical solutions of the above all embodiments, it at least has all the beneficial effects brought by the technical solutions of the above embodiments, which will not be elaborated herein one by one.
[0242] It should be understood that the above is only for illustration and does not constitute any limitation to the technical solutions of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not make any restrictions in this regard.
[0243] It should be noted that the above-described work process is only illustrative and does not constitute a limitation to the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no restrictions are made here.
[0244] In addition, for the technical details not described in detail in this embodiment, reference may be made to the soil heavy metal content prediction method provided in any embodiment of the present invention, which will not be elaborated here.
[0245] In addition, it should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.
[0246] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.
[0247] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM) / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0248] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for predicting the heavy metal content in soil, characterized in that, The soil heavy metal content prediction method includes: Obtain an initial soil data set; Cluster the initial soil data set to obtain target clustering clusters; Obtain a reference soil data set through the target clustering clusters; Filter the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set; Optimize a recursive neural network through the target soil data set to obtain a target recursive neural network; Train the target recursive neural network through the target soil data set to obtain a soil heavy metal content prediction model; Predict the soil heavy metal content through the soil heavy metal content prediction model; The step of optimizing a recursive neural network through the target soil data set to obtain a target recursive neural network includes: Obtain the quantity and actual values of the target soil data set; Obtain the predicted values of the recursive neural network; Calculate a mean square error function through the quantity, the actual values, and the predicted values; Set the weights in the target soil data set; Calculate a weight decay term according to the weights; Obtain the Jacobian matrix and Hessian matrix of the target soil data set; Calculate the number of effective weights through the Jacobian matrix, the Hessian matrix, and the quantity; Calculate a regularization parameter through the number of effective weights, the weight decay term, the quantity, and the mean square error function; Calculate an objective function through the regularization parameter, the weight decay term, and the mean square error function; Optimize the structure of the recursive neural network according to the objective function and the regularization parameter to obtain the structure of the target recursive neural network; Obtain the network weights and thresholds of the structure of the target recursive neural network; Update the network weights and the thresholds through an adaptive GA-PSO algorithm to optimize the network error of the recursive neural network to obtain a target recursive neural network, where the formulas for calculating individual crossover and mutation in the adaptive GA-PSO algorithm are: , P ci is the crossover probability, P mi is the mutation probability, c 1 and c 2 are the basic crossover probabilities, c 3 is the crossover constant, 0 < c 1 < c 2 < c 3 < 1, c 4 and c 5 are the basic mutation probabilities, c 6 is the mutation constant, 0 < c 4 < c 5 < c 6 < 1, f i is the individual fitness value, f max is the maximum of the individual fitness values, f avg is the average of the individual fitness values; among them, the formula for calculating the current inertia weight in the adaptive GA-PSO algorithm is: , ω is the current inertia weight, ω min is the first inertia weight, ω max is the second inertia weight, t is the position parameter, and k is the current iteration number.
2. The soil heavy metal content prediction method according to claim 1, characterized in that, The step of filtering the reference soil data set through an entropy weight distance detection strategy to obtain a target soil data set includes: Obtain the attributes of each data in the reference soil data set and the dimension of the reference soil data set; Combine the attributes of the each data to obtain an attribute set; Obtain the value sets of each attribute in the attribute set and the value probabilities of each attribute; Calculate entropy values through the value sets of the each attribute and the value probabilities of the each attribute; Obtain the outlier attributes and non-outlier attributes in the reference soil data set through the entropy values and the dimension of the reference soil data set; Respectively set the attribute weights of the outlier attributes and the non-outlier attributes; Obtain the values of the data objects in the reference soil data set on the corresponding attributes; Calculate an entropy weight distance through the values on the corresponding attributes and the attribute weights; When the entropy weight distance is greater than a preset distance threshold, regard the data corresponding to the entropy weight distance as outlier data; Remove the outlier data from the reference soil data set to obtain a target soil data set.
3. The soil heavy metal content prediction method according to claim 1, characterized in that The adaptive GA-PSO algorithm includes an adaptive genetic algorithm and an adaptive particle swarm algorithm; Updating the network weights and the thresholds through the adaptive GA-PSO algorithm to obtain a target recurrent neural network includes: Initializing a population through an adaptive genetic algorithm, where the population includes a preset number of individuals, and each individual is a set of parameters of network weights and thresholds; Calculating the individual fitness value of each individual; Obtaining the maximum value of the individual fitness values and the average value of the individual fitness; Setting a crossover base probability, a mutation base probability, a crossover constant, and a mutation constant, where the crossover base probability is less than the crossover constant, and the mutation base probability is less than the mutation constant; Performing crossover and mutation on the individuals through the individual fitness value, the maximum value, the average value, the crossover base probability, the crossover constant, the mutation base probability, and the mutation constant; Obtaining a first individual fitness value, a second individual fitness value, and the number of consecutive decreases in the individual fitness value according to the individual fitness value; Calculating an update probability through the first individual fitness value, the second individual fitness value, and the number of consecutive decreases in the individual fitness value; Judging whether an individual performs crossover and mutation according to the update probability; When the update probability is equal to a preset probability value, outputting the current population; Using an adaptive particle swarm optimization algorithm to calculate the current population to update the network weights and the thresholds to obtain a target recurrent neural network.
4. The soil heavy metal content prediction method according to claim 3, wherein The using an adaptive particle swarm optimization algorithm to calculate the current population to update the network weights and the thresholds to obtain a target recurrent neural network includes: Initializing the current population through an adaptive particle swarm optimization algorithm, and setting the initial position and the initial velocity of the current population; Calculating the individual fitness values in the current population to obtain individual extreme values and a global extreme value; Setting a maximum number of iterations and obtaining the current number of iterations; Setting a position parameter according to the maximum number of iterations; Setting a first inertia weight, a second inertia weight, a first learning factor, a second learning factor, and a random value, where the first inertia weight is less than the second inertia weight, and the first learning factor is less than the second learning factor; Calculating a current inertia weight through the first inertia weight, the second inertia weight, the position parameter, and the current number of iterations; Calculating a current learning factor according to the first learning factor, the second learning factor, the maximum number of iterations, and the current number of iterations; Calculating a compression factor through the current learning factor; Calculating a particle update velocity according to the compression factor, the current inertia weight, the current learning factor, the random value, the individual extreme value, the global extreme value, the initial velocity, and the initial position; Performing local optimization and global optimization training according to the particle update velocity, and determining updated network weights and an updated threshold when the training requirements are met; Optimizing the network error of the recurrent neural network according to the updated network weights and the updated threshold to obtain a target recurrent neural network.
5. The soil heavy metal content prediction method according to claim 1, wherein Clustering the initial soil data set to obtain target clustering clusters includes: Select target points from the initial soil dataset, and use the target points as the initial clustering centers; Calculate the first distances from each point in the initial soil dataset to the initial clustering centers; Screen the initial clustering centers through the first distances to obtain target clustering centers; Divide the initial soil dataset based on the target clustering centers to obtain target clustering clusters.
6. The method for predicting soil heavy metal content according to any one of claims 1-5, characterized in that, The obtaining of the initial soil dataset includes: Obtain the original soil dataset; Perform data cleaning on the original soil dataset to obtain abnormal data in the original soil dataset; Remove the abnormal data from the original soil dataset to obtain the initial soil dataset.
7. A device for predicting the heavy metal content in soil, characterized in that, The soil heavy metal content prediction device includes: An acquisition module for acquiring the initial soil dataset; A clustering module for clustering the initial soil dataset to obtain target clustering clusters; The acquisition module is further configured to obtain a reference soil dataset through the target clustering clusters; A filtering module for filtering the reference soil dataset through an entropy weight distance detection strategy to obtain a target soil dataset; An optimization module for optimizing a recursive neural network through the target soil dataset to obtain a target recursive neural network; A training module for training the target recursive neural network through the target soil dataset to obtain a soil heavy metal content prediction model; A prediction module for predicting the soil heavy metal content through the soil heavy metal content prediction model; The optimization module is further configured to obtain the quantity and actual values of the target soil dataset; Obtain the predicted values of the recursive neural network; Calculate a mean square error function through the quantity, the actual values, and the predicted values; Set the weights in the target soil dataset; Calculate a weight decay term according to the weights; Obtain the Jacobian matrix and Hessian matrix of the target soil dataset; Calculate the number of effective weights through the Jacobian matrix, the Hessian matrix, and the quantity; Calculate a regularization parameter through the number of effective weights, the weight decay term, the quantity, and the mean square error function; Calculate an objective function through the regularization parameter, the weight decay term, and the mean square error function; Optimize the structure of the recursive neural network according to the objective function and the regularization parameter to obtain the structure of the target recursive neural network; Obtain the network weights and thresholds of the structure of the target recursive neural network; Update the network weights and the thresholds through an adaptive GA-PSO algorithm to optimize the network error of the recursive neural network to obtain a target recursive neural network, where the formulas for calculating individual crossover and mutation in the adaptive GA-PSO algorithm are: , P ci is the crossover probability, P mi is the mutation probability, c 1 and c 2 are the basic crossover probabilities, c 3 is the crossover constant, 0 < c 1 < c 2 < c 3 < 1, c 4 and c 5 are the basic mutation probabilities, c 6 is the mutation constant, 0 < c 4 < c 5 < c 6 < 1, f i is the individual fitness value, f max is the maximum value of the individual fitness value, f avg is the average value of the individual fitness; among them, the formula for calculating the current inertia weight in the adaptive GA-PSO algorithm is: , ω is the current inertia weight, ω min is the first inertia weight, ω max is the second inertia weight, t is the position parameter, and k is the current iteration number.
8. A soil heavy metal content prediction device, characterized in that, The soil heavy metal content prediction device includes: a memory, a processor, and a soil heavy metal content prediction program stored on the memory and executable on the processor, and the soil heavy metal content prediction program is configured to implement the soil heavy metal content prediction method according to any one of claims 1 to 6.
9. A storage medium, characterized in that, A soil heavy metal content prediction program is stored on the storage medium. When the soil heavy metal content prediction program is executed by a processor, the soil heavy metal content prediction method according to any one of claims 1 to 6 is implemented.