Soil pollution analysis and treatment method based on big data analysis
Through the pre-treatment of soil data and the construction of multi-dimensional data models, combined with the joint optimization of data encoder and loss function, the problems of insufficient utilization of data multi-dimensional information and difficulty in noise processing in soil pollution analysis in the existing technology are solved, and more accurate prediction of soil pollution level and formulation of treatment plans are achieved.
Patent Information
- Application Number
- CN202510128345.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-23
AI Technical Summary
Existing soil pollution analysis methods cannot make full use of the multi-dimensional information of big data, and it is difficult to effectively process noise and complex relationships in the data, resulting in inaccurate division of pollution levels, which in turn affects the scientificity and effectiveness of the governance plan.
Through soil data acquisition and pretreatment, a multi-dimensional data model of soil pollution is constructed, a low-dimensional representation of features is obtained using a data encoder, the correlation matrix is reconstructed, and a joint optimization is carried out based on local similarity loss and global classification loss to predict soil pollution levels.
Effectively reduce data dimensions, reduce noise interference, improve the accuracy and efficiency of soil pollution analysis, ensure the accuracy of pollution level prediction, and provide a reliable basis for formulating scientific and reasonable governance plans.
Smart Images

Figure CN120030304A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis and soil pollution control, and more specifically, to a soil pollution analysis and control method based on big data analysis. Background Art
[0002] With the rapid development of industrialization and urbanization, soil pollution has become increasingly prominent and has become a global environmental problem. Traditional soil pollution analysis methods mainly rely on laboratory testing and on-site sampling and analysis. These methods are not only time-consuming and laborious, but also difficult to fully reflect the complexity and dynamic changes of soil pollution. In recent years, big data technology has gradually been applied in the field of environmental science. By collecting and analyzing a large amount of soil data, it can more efficiently identify pollution sources, assess the degree of pollution, and formulate treatment plans. However, when processing soil pollution big data, existing technologies often face problems such as high data dimensionality, high noise, and complex correlation, which limits the accuracy and reliability of the analysis results.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: the existing soil pollution analysis method cannot fully utilize the multi-dimensional information of big data, and it is difficult to effectively deal with the noise and complex correlations in the data, resulting in inaccurate classification of pollution levels, which in turn affects the scientificity and effectiveness of the control plan. Summary of the invention
[0004] The present invention provides a soil pollution analysis and control method based on big data analysis, comprising: Soil data collection and preprocessing; Construct a multi-dimensional data model of soil pollution based on the pre-processed soil data; Obtaining a low-dimensional representation of the features of the multi-dimensional data model through a data encoder; Reconstructing an association matrix of the multidimensional data model based on the feature low-dimensional representation, and calculating a local similarity loss based on the reconstructed association matrix and an original association matrix of the multidimensional data model; Using a clustering-based and highly robust analysis method, the low-dimensional representation of the features is input to classify the soil samples and generate pollution level labels; A fully connected layer is set after the data encoder, and the low-dimensional representation of the feature is input into the fully connected layer to obtain an intermediate analysis result; Calculating a global classification loss based on the pollution level labels and the intermediate analysis results; Performing joint optimization based on the local similarity loss and the global classification loss; The soil pollution level is predicted based on the result of the joint optimization to obtain a treatment plan.
[0005] Furthermore, the soil data collection and preprocessing specifically include: for each soil sample, synthesizing the geographical location, pollution type and pollution source information of the sample into a set of data, and then preprocessing the set of data; All the features of the preprocessed data are sent to the trained feature vector model to obtain the feature vector of each feature and the feature vectors of all the features are averaged as the comprehensive feature vector of the soil sample; The common pollution sources, common geographical locations and common pollution types of different soil samples were preprocessed to obtain the common pollution sources, common geographical locations and common pollution types with the same expression.
[0006] Furthermore, the preprocessing of the group of data specifically includes: standardizing the geographic location coordinates, removing fuzzy expressions in the pollution type, removing repeated pollution source information, segmenting the data with specific separators, and removing irrelevant features and features with a length less than a certain threshold; The preprocessing of common pollution sources, common geographical locations and common pollution types of different soil samples specifically includes: for common pollution sources, standardization processing, unification of name writing order and normalization of pollution source identification; for common geographical locations, coordinate standardization, removal of redundant spaces, segmentation of data with specific separators, removal of irrelevant features and features with a length less than a certain threshold; for common pollution types, standardization processing, removal of ambiguous expressions, segmentation of data with specific separators, removal of irrelevant features and features with a length less than a certain threshold.
[0007] Furthermore, constructing a soil pollution multidimensional data model based on the preprocessed soil data specifically includes: taking each soil sample as a node of the multidimensional data model; Using the comprehensive feature vector of each soil sample as the node feature of the sample in the multi-dimensional data model; Calculate the similarity between the common pollution source, common geographical location and common pollution type of the two soil samples and set the corresponding similarity threshold. Among the three attributes of the common pollution source, common geographical location and common pollution type of the two soil samples, if the similarity of one attribute exceeds the corresponding similarity threshold, establish an edge of the attribute between the two soil sample nodes.
[0008] Furthermore, for common pollution sources and common pollution types, feature overlap is used to calculate similarity, and for common geographical locations, the inverse of the Euclidean distance is used as the metric for similarity.
[0009] Furthermore, when obtaining the characteristic low-dimensional representation of the multidimensional data model through the data encoder, a two-layer long short-term memory network (LSTM) is used as the data encoder, the input of each layer of the long short-term memory network is the characteristic low-dimensional representation of the previous layer, and the output is the characteristic low-dimensional representation of this layer, and the input of the first layer is the comprehensive feature vector of the soil sample.
[0010] Furthermore, the local similarity loss is calculated based on the reconstructed association matrix and the original association matrix of the multidimensional data model as follows: the objective function of the local similarity loss is designed to minimize the mean square error loss between the reconstructed association matrix and the original association matrix. :
[0011] In the formula, is the reconstructed correlation matrix; is an element in the reconstructed association matrix, indicating the probability of predicting the association between node i and node j, with a value range of [0,1]; A is the original association matrix of the multidimensional data model; is an element in the original association matrix of the multidimensional data model, and its value is 0 or 1; N is the number of nodes in the multidimensional data model.
[0012] Furthermore, the global classification loss is calculated based on the pollution level label and the intermediate analysis result as follows: The cross entropy loss function between the pollution level label and the intermediate analysis result is defined as the global classification loss :
[0013] Where, C is the intermediate analysis result; represents the probability that node i belongs to category c, and its value range is [0,1]; Y is the pollution level label, represents the true category label of node i, which takes a value of 0 or 1; N is the number of nodes in the multidimensional data model; C is the total number of pollution level categories.
[0014] Furthermore, the joint optimization based on the local similarity loss and the global classification loss is specifically as follows: Using global classification loss and local similarity loss to achieve a balance between them, that is,
[0015] In the formula, is the weighted loss; is a hyperparameter set empirically; In getting the weighted loss Then, using the stochastic gradient descent algorithm, based on the weighted loss The parameters of the data encoder and the fully connected layer are trained for multiple rounds, and the parameters of the data encoder and the fully connected layer are jointly optimized through training.
[0016] Furthermore, the soil pollution level is predicted based on the result of the joint optimization to obtain a treatment plan, specifically: the pollution level label generated by the last round of training is taken as the basis for the final treatment plan.
[0017] According to the above-mentioned embodiments of the present invention, at least the following beneficial effects are achieved: First, by constructing a multi-dimensional data model of soil pollution and using a data encoder to represent features in a low-dimensional manner, the data dimension can be effectively reduced, noise interference can be reduced, and key feature information of soil pollution data can be retained, thereby improving the accuracy and efficiency of soil pollution analysis. Secondly, the joint optimization method based on local similarity loss and global classification loss can better balance the local correlation between soil samples and the global classification accuracy, making the prediction of soil pollution levels more accurate, thereby providing a reliable basis for formulating scientific and reasonable soil pollution control plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which: Figure 1 A schematic diagram of a process for analyzing and treating soil pollution based on big data analysis provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0019] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0020] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, apparatus, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0021] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0022] Reference below Figure 1 , Figure 1 A schematic diagram of a process flow of a soil pollution analysis and treatment method based on big data analysis provided by an embodiment of the present invention. Figure 1 As shown, a soil pollution analysis and control method 100 based on big data analysis includes: Step 101, soil data collection and preprocessing; Step 102, constructing a soil pollution multi-dimensional data model based on the pre-processed soil data; Step 103, obtaining a low-dimensional representation of the features of the multi-dimensional data model through a data encoder; Step 104, reconstructing the association matrix of the multidimensional data model based on the feature low-dimensional representation, and calculating the local similarity loss based on the reconstructed association matrix and the original association matrix of the multidimensional data model; Step 105, using a clustering-based and highly robust analysis method, inputting the feature low-dimensional representation to classify the soil sample and generate a pollution level label; Step 106, setting a fully connected layer after the data encoder, and inputting the feature low-dimensional representation into the fully connected layer to obtain an intermediate analysis result; Step 107, calculating a global classification loss based on the pollution level label and the intermediate analysis result; Step 108, performing joint optimization based on the local similarity loss and the global classification loss; Step 109: predicting the soil pollution level based on the result of the joint optimization to obtain a treatment plan.
[0023] It should be noted that the present invention proposes a soil pollution analysis and control method based on big data analysis, the core of which is to achieve accurate prediction of soil pollution levels and formulation of control plans through a series of data processing and analysis steps. Among them, soil data collection and preprocessing refers to the collection and preliminary processing of information such as the geographical location, pollution type and pollution source of soil samples to remove noise and redundant information and ensure the quality and availability of data. The multidimensional data model refers to the integration of multiple characteristics of soil samples (such as geographical location, pollution type, etc.) into a model for analyzing the correlation between samples. Through technical means such as low-dimensional representation of features and local similarity loss, the performance of the model can be further optimized to improve the accuracy and efficiency of the analysis.
[0024] Specifically, soil data collection and preprocessing involve standardizing the geographic location coordinates of soil samples, removing fuzzy expressions in pollution types, and removing duplicate pollution source information. The purpose of these operations is to convert raw data into standardized data that can be used for subsequent analysis. For example, geographic location coordinates can be standardized by normalizing longitude and latitude; pollution types can improve data accuracy by removing fuzzy expressions (such as changing possible heavy metals to heavy metals). In addition, when constructing a multidimensional data model, each soil sample will be taken as a node, and its comprehensive feature vector will be used as a node feature. When calculating similarity, for common pollution sources and pollution types, feature overlap can be used to measure; for geographic locations, the inverse of the Euclidean distance is used as a measure of similarity. The selection of these parameters and concepts is to better reflect the correlation between soil samples.
[0025] Preferably, the operation steps can be further refined in the soil data collection and preprocessing stage. For example, for the standardization of geographic location coordinates, a specific geographic coordinate system (such as WGS-84) can be used for unified processing; for the standardization of pollution types, a detailed pollution type dictionary can be established to uniformly classify fuzzy expressions into standard terms in the dictionary. In the construction of a multidimensional data model, a similarity threshold can be set. For example, when the overlap of the common pollution source characteristics of two soil samples exceeds 0.8, it is considered that there is an association between them; for the similarity of geographic location, the association can be established when the inverse of the Euclidean distance is greater than 0.5. In addition, in the low-dimensional representation stage of features, a two-layer long short-term memory network (LSTM) can be used as a data encoder. The input of the first layer is the comprehensive feature vector of the soil sample. By extracting features layer by layer, a low-dimensional representation is finally obtained for subsequent analysis and optimization.
[0026] In some embodiments, the soil data collection and preprocessing specifically include: for each soil sample, synthesizing the geographical location, pollution type and pollution source information of the sample into a set of data, and then preprocessing the set of data; All the features of the preprocessed data are sent to the trained feature vector model to obtain the feature vector of each feature and the feature vectors of all the features are averaged as the comprehensive feature vector of the soil sample; The common pollution sources, common geographical locations and common pollution types of different soil samples were preprocessed to obtain the common pollution sources, common geographical locations and common pollution types with the same expression.
[0027] It should be noted that the soil data collection and preprocessing mentioned in the present invention is the basic link of the entire soil pollution analysis and control method. Its main purpose is to integrate the geographical location, pollution type and pollution source of the soil sample into a set of data, and remove the noise and redundant information in the data through preprocessing operations, and extract representative feature vectors. Among them, the comprehensive feature vector refers to the vector obtained by weighted averaging of multiple features of the soil sample, which can effectively characterize the overall characteristics of the soil sample and provide a data basis for subsequent model construction and analysis.
[0028] Specifically, soil data collection and preprocessing include the following key steps: First, for each soil sample, its geographical location, pollution type and pollution source information are integrated into a set of data. Among them, geographical location information can be obtained through GPS positioning and expressed in the form of longitude and latitude; pollution types can include heavy metal pollution, organic pollution, etc.; pollution source information can be factory emissions, agricultural fertilization, etc. Then, this set of data is preprocessed, including standardization of geographical location coordinates, removal of fuzzy expressions in pollution types, and removal of repeated pollution source information. For example, the standardization of geographical location coordinates can be normalized to map longitude and latitude values to the [0,1] interval; for fuzzy expressions in pollution types, such as possible heavy metals, it can be clearly stated as containing heavy metals or not containing heavy metals. In addition, for common pollution sources, common geographical locations and common pollution types of different soil samples, unified preprocessing is required to ensure their consistency in expression. For example, for common pollution sources, standardization can be performed to unify the name writing order and normalize the pollution source identifier; for common geographical locations, redundant spaces can be removed and data can be split with specific separators.
[0029] Preferably, the operation steps can be further refined during the soil data collection and preprocessing stage. For example, for the standardization of geographic location coordinates, a specific geographic coordinate system (such as WGS-84) can be used for unified processing and converted into a unified projection coordinate system for subsequent calculations. For the standardization of pollution types, a detailed pollution type dictionary can be established to uniformly classify fuzzy expressions into standard terms in the dictionary. In the process of generating feature vectors, different weights can be assigned to each feature. For example, a higher weight can be assigned to geographic location information because it usually plays a more important role in soil pollution analysis. In addition, for the preprocessing of common pollution sources and common pollution types, more complex text processing techniques can be used, such as word embedding technology in natural language processing, to convert text information into numerical vectors for better similarity calculation and analysis.
[0030] In some embodiments, the preprocessing of the group of data specifically includes: standardizing geographic location coordinates, removing ambiguous expressions in pollution types, removing repeated pollution source information, segmenting data with specific separators, and removing irrelevant features and features with lengths less than a certain threshold; The preprocessing of common pollution sources, common geographical locations and common pollution types of different soil samples specifically includes: for common pollution sources, standardization processing, unification of name writing order and normalization of pollution source identification; for common geographical locations, coordinate standardization, removal of redundant spaces, segmentation of data with specific separators, removal of irrelevant features and features with a length less than a certain threshold; for common pollution types, standardization processing, removal of ambiguous expressions, segmentation of data with specific separators, removal of irrelevant features and features with a length less than a certain threshold.
[0031] It should be noted that the steps of soil data collection and preprocessing are further refined in the present invention, aiming to improve the quality and consistency of data through a series of specific operations, and provide more accurate input data for subsequent soil pollution analysis. Among them, preprocessing refers to the process of cleaning, standardizing and normalizing the collected soil data, with the purpose of removing noise, duplicate information and irrelevant features in the data, while ensuring the integrity and consistency of the data. For example, geographic location coordinate standardization refers to the uniform conversion of geographic location information in different formats into a standard latitude and longitude format; removing fuzzy expressions in pollution types refers to converting ambiguous descriptions of pollution types into clear classifications; removing duplicate pollution source information refers to deleting duplicate pollution source records to avoid data redundancy.
[0032] Specifically, during the preprocessing process, the geographic location coordinates can be converted into a unified coordinate system through standardization, such as using the WGS-84 coordinate system for unification. For fuzzy expressions in pollution types, a standardized pollution type dictionary can be established to map all possible fuzzy expressions to clear classifications. For example, the possible presence of heavy metals is uniformly mapped as heavy metal pollution. For repeated pollution source information, it can be cleaned up through a data deduplication algorithm. In addition, the preprocessing of common pollution sources, common geographic locations, and common pollution types can be further refined as follows: for common pollution sources, the names are standardized, such as unifying a chemical plant as a chemical enterprise; for common geographic locations, extra spaces are removed, and the data is separated by a unified separator (such as a comma); for common pollution types, standardization is performed to remove fuzzy expressions and separate the data with a specific separator. The purpose of these operations is to ensure the consistency and comparability of data expression.
[0033] Preferably, in the preprocessing stage, the operation steps can be further refined. For example, for the standardization of geographic location coordinates, more accurate geographic information system (GIS) technology can be used to convert coordinate data into a unified geographic projection format for more accurate spatial analysis. For the standardization of pollution types, machine learning algorithms can be introduced to automatically identify and correct ambiguous statements. For the processing of common pollution sources, industry standards and specifications can be combined to classify and encode pollution sources for better statistics and analysis. In addition, for the preprocessing of common geographic locations and common pollution types, natural language processing technology can be used to perform word segmentation, part-of-speech tagging and entity recognition on text data to improve the accuracy and consistency of the data.
[0034] In some embodiments, constructing a soil pollution multidimensional data model based on the preprocessed soil data specifically includes: taking each soil sample as a node of the multidimensional data model; Using the comprehensive feature vector of each soil sample as the node feature of the sample in the multi-dimensional data model; Calculate the similarity between the common pollution source, common geographical location and common pollution type of the two soil samples and set the corresponding similarity threshold. Among the three attributes of the common pollution source, common geographical location and common pollution type of the two soil samples, if the similarity of one attribute exceeds the corresponding similarity threshold, establish an edge of the attribute between the two soil sample nodes.
[0035] It should be noted that the present invention describes in detail the construction of a multidimensional data model for soil pollution, which is the core link in realizing the soil pollution analysis and treatment method. The so-called multidimensional data model refers to integrating multiple characteristics of soil samples (such as geographical location, pollution type, pollution source, etc.) into a model to reflect the complex correlation between soil samples. Each soil sample is used as a node, and its comprehensive feature vector is used as a node feature. By calculating the similarity between samples and setting a similarity threshold, edges can be established between sample nodes that meet the conditions, thereby constructing a network model reflecting the soil pollution relationship.
[0036] Specifically, when constructing a multidimensional data model, each soil sample is first taken as a node of the model, and its comprehensive feature vector is taken as the feature of the node. The comprehensive feature vector is obtained by weighted averaging multiple features of the soil sample (such as geographic location coordinates, pollution type, pollution source, etc.). Next, the similarity between the two soil samples is calculated, including the similarity between common pollution sources, common geographic locations, and common pollution types. For the calculation of similarity, a specific similarity threshold can be set. For example, when the overlap of the common pollution source features of two soil samples exceeds 0.7, it is considered that there is an association between them; for common geographic locations, the reciprocal of the Euclidean distance can be used as a measure of similarity. When the reciprocal of the Euclidean distance is greater than 0.6, it is considered that the two samples are associated in terms of geographical location. In this way, edges can be established between sample nodes that meet the similarity conditions, thereby constructing a complete multidimensional data model.
[0037] Preferably, when constructing a multidimensional data model, the operation steps can be further refined. For example, the setting of the similarity threshold can be adjusted according to the actual application scenario. For the similarity calculation of common pollution sources, in addition to feature overlap, other similarity measurement methods, such as the Jaccard similarity coefficient, can be introduced to more comprehensively evaluate the correlation between samples. For the similarity calculation of common geographical locations, in addition to the Euclidean distance, other geographical distance measurement methods, such as the Manhattan distance or the Mahalanobis distance, can also be considered to adapt to different geographical data distributions. In addition, during the model construction process, a weight mechanism can be introduced to assign different weights to different types of similarities to better reflect the comprehensive correlation between soil samples.
[0038] In some embodiments, for common pollution sources and common pollution types, feature overlap is used to calculate similarity, and for common geographic locations, the inverse of the Euclidean distance is used as a measure of similarity.
[0039] It should be noted that the present invention further clarifies the specific method of similarity calculation when constructing a multidimensional data model of soil pollution. For common pollution sources and common pollution types, feature overlap is used to calculate similarity; and for common geographical locations, the reciprocal of the Euclidean distance is used as the measure of similarity. The feature overlap here refers to the ratio of the number of common features of two samples on a certain feature to the total number of features, reflecting the degree of similarity of the samples on the feature; the reciprocal of the Euclidean distance is a similarity measurement method based on spatial distance, and the closer the distance, the higher the similarity. Through these two different similarity calculation methods, the correlation between soil samples can be more accurately measured, providing a more reliable data basis for subsequent pollution analysis and governance.
[0040] Specifically, for the similarity calculation of common pollution sources, the feature overlap can be calculated by the following formula: Assuming that two soil samples A and B have sets SA and SB in pollution source characteristics, their feature overlap is ,in represents the size of the intersection of two sets, represents the size of the union of two sets. For common pollution types, the feature overlap calculation method is also used, and the pollution type is regarded as a feature set for similarity evaluation. For common geographical locations, it is assumed that the geographical coordinates of the two soil samples are and , then the Euclidean distance between them is , whose reciprocal is the similarity. By setting the similarity threshold, such as feature overlap greater than 0.6 or the reciprocal of the Euclidean distance greater than 0.5, edges can be established between sample nodes that meet the conditions, thereby constructing a network model that reflects the soil pollution relationship.
[0041] Preferably, in practical applications, the similarity calculation method can be adjusted according to the specific characteristics of the soil samples and the application scenarios. For example, for the similarity calculation of common pollution sources, in addition to feature overlap, weighted feature overlap can also be introduced to assign different weights according to the importance of different pollution sources. For the similarity calculation of common geographical locations, in addition to the reciprocal of the Euclidean distance, other geographical distance measurement methods can also be considered, such as the reciprocal of the Manhattan distance or the cosine similarity based on geographical coordinates, to adapt to different geographical data distribution situations. In addition, when setting the similarity threshold, it can be dynamically adjusted according to the distribution density and pollution degree of the soil samples to more accurately reflect the correlation between samples.
[0042] In some embodiments, when obtaining the characteristic low-dimensional representation of the multidimensional data model through a data encoder, a two-layer long short-term memory network (LSTM) is used as a data encoder, the input of each layer of the long short-term memory network is the characteristic low-dimensional representation of the previous layer, and the output is the characteristic low-dimensional representation of this layer, and the input of the first layer is the comprehensive feature vector of the soil sample.
[0043] It should be noted that when obtaining the characteristic low-dimensional representation of the multi-dimensional data model of soil pollution, the present invention uses a two-layer long short-term memory network (LSTM) as a data encoder. LSTM is a special recursive neural network (RNN) that can effectively process and predict long-term dependencies in time series data. Here, the characteristic low-dimensional representation refers to compressing the high-dimensional soil sample feature vector into a representation in a low-dimensional space through a neural network while retaining the key information of the original data. Through the LSTM network, the characteristics of the soil sample can be extracted and encoded layer by layer, thereby providing a more efficient data representation for subsequent analysis and optimization.
[0044] Specifically, each layer of the LSTM network processes the input feature vector and extracts more representative features. In the present invention, the input of the first layer of LSTM is the comprehensive feature vector of the soil sample, which contains information such as the geographical location, pollution type and pollution source of the soil sample. The output of each layer of LSTM is the low-dimensional representation of the features extracted by this layer, and these low-dimensional representations will serve as the input of the next layer. Through the processing of two layers of LSTM, the low-dimensional representation of the features finally obtained can better capture the complex relationships and patterns between soil samples. For example, the number of hidden units of the LSTM network can be set according to the complexity of the actual data set, usually between tens and hundreds. In addition, the training process of the LSTM network can be optimized by the back propagation algorithm to minimize the reconstruction error or classification error.
[0045] Preferably, when using LSTM as a data encoder, the operation steps can be further refined. For example, the number of hidden units of the LSTM network can be adjusted according to the scale and complexity of the soil sample data. If the data set is large and has many features, the number of hidden units can be appropriately increased, for example, set to 128 or 256. In addition, in order to improve the training efficiency and stability of the model, regularization techniques such as Dropout can be introduced to prevent overfitting. During the training process, the Adam optimizer can be used, and its learning rate can be dynamically adjusted according to the convergence during the training process. In addition, the introduction of a bidirectional LSTM (Bi-LSTM) can also be considered, which can simultaneously process the forward and reverse information of the time series data, so as to more comprehensively capture the characteristics of the soil sample.
[0046] In some embodiments, the local similarity loss is calculated based on the reconstructed association matrix and the original association matrix of the multidimensional data model as follows: the objective function of the local similarity loss is designed to minimize the mean square error loss between the reconstructed association matrix and the original association matrix. :
[0047] In the formula, is the reconstructed correlation matrix; is an element in the reconstructed association matrix, indicating the probability of predicting the association between node i and node j, with a value range of [0,1]; A is the original association matrix of the multidimensional data model; is an element in the original association matrix of the multidimensional data model, and its value is 0 or 1; N is the number of nodes in the multidimensional data model.
[0048] It should be noted that, in the present invention, the calculation of local similarity loss is achieved by comparing the difference between the reconstructed association matrix and the original association matrix. The local similarity loss here refers to the optimization of the reconstructed association matrix in the multidimensional data model to make it as close as possible to the original association matrix, thereby retaining the local similarity between soil samples. Specifically, the elements in the reconstructed association matrix represent the probability of association between the predicted nodes, while the elements in the original association matrix represent the actual association relationship. By minimizing the mean square error loss between the two, the local similarity of the model can be effectively optimized and the accuracy of soil pollution analysis can be improved.
[0049] Specifically, the calculation formula for local similarity loss is:
[0050] in, is the reconstructed incidence matrix, whose elements Represents a prediction node and nodes The probability of an association between them is in the range of [0,1]; is the original incidence matrix, whose elements Indicates the actual association relationship, with a value of 0 or 1; N is the number of nodes in the multidimensional data model. In practical applications, the loss function can be minimized by adjusting the parameters of the model to optimize the reconstructed association matrix to make it closer to the original association matrix. For example, during training, stochastic gradient descent (SGD) or other optimization algorithms can be used to update the model parameters to gradually reduce the local similarity loss.
[0051] Preferably, when calculating the local similarity loss, the operation steps can be further refined. For example, for the reconstructed association matrix, a regularization term, such as L2 regularization, can be introduced to prevent the model from overfitting. In addition, during the optimization process, the learning rate can be dynamically adjusted, such as using a learning rate decay strategy, to speed up the training and improve the convergence performance of the model. For the selection of the loss function, in addition to the mean square error loss, other loss functions, such as the cross entropy loss, can also be considered to better measure the difference between the predicted association probability and the actual association. In addition, during the model training process, an early stopping mechanism can be introduced to stop training in advance when the loss on the validation set no longer decreases significantly to avoid overfitting.
[0052] In some embodiments, the global classification loss is calculated based on the pollution level label and the intermediate analysis result as follows: The cross entropy loss function between the pollution level label and the intermediate analysis result is defined as the global classification loss :
[0053] Where, C is the intermediate analysis result; represents the probability that node i belongs to category c, and its value range is [0,1]; Y is the pollution level label, represents the true category label of node i, which takes a value of 0 or 1; N is the number of nodes in the multidimensional data model; C is the total number of pollution level categories.
[0054] It should be noted that the global classification loss mentioned in the present invention is calculated by comparing the difference between the intermediate analysis results and the pollution level label. The global classification loss here refers to the optimization of the output results of the model in the entire multi-dimensional data model of soil pollution, so as to make it as close as possible to the real pollution level label, thereby improving the accuracy of soil pollution level classification. Specifically, the intermediate analysis results are obtained by processing the low-dimensional representation of the features through the fully connected layer, and the pollution level label is the pre-defined soil pollution level. By calculating the cross entropy loss between the two, the classification performance of the model can be effectively measured and the optimization direction of the model can be guided.
[0055] Specifically, the calculation formula for the global classification loss is:
[0056] Among them, C is the intermediate analysis result, which represents the probability of each node belonging to different pollution levels predicted by the model. represents the probability that node i belongs to category c, and its value range is [0,1]; Y is the real pollution level label, represents the true category label of node i, which takes a value of 0 or 1; N is the number of nodes in the multidimensional data model; C is the total number of categories of pollution levels. In practical applications, the classification performance of the model can be optimized by adjusting the parameters of the model to minimize this loss function. For example, during training, stochastic gradient descent (SGD) or other optimization algorithms can be used to update the model parameters to gradually reduce the global classification loss.
[0057] Preferably, when calculating the global classification loss, the operation steps can be further refined. For example, for the generation of intermediate analysis results, a multi-layer perceptron (MLP) can be introduced as an alternative to the fully connected layer to enhance the nonlinear fitting ability of the model. In addition, during the optimization process, the learning rate can be dynamically adjusted, such as using a learning rate decay strategy, to speed up the training and improve the convergence performance of the model. For the selection of loss functions, in addition to the cross entropy loss, other loss functions such as weighted cross entropy loss can also be considered to better deal with the problem of imbalanced categories. In addition, during the model training process, an early stopping mechanism can be introduced to stop training in advance when the loss on the validation set no longer decreases significantly to avoid overfitting.
[0058] In some embodiments, the joint optimization based on the local similarity loss and the global classification loss is specifically: Using global classification loss and local similarity loss to achieve a balance between them, that is,
[0059] In the formula, is the weighted loss; is a hyperparameter set empirically; In getting the weighted loss Then, using the stochastic gradient descent algorithm, based on the weighted loss The parameters of the data encoder and the fully connected layer are trained for multiple rounds, and the parameters of the data encoder and the fully connected layer are jointly optimized through training.
[0060] It should be noted that in the present invention, the joint optimization of local similarity loss and global classification loss is achieved by weighted sum. This joint optimization method aims to balance the local correlation and global classification accuracy in soil pollution analysis, thereby improving the overall performance of the model. Among them, the local similarity loss refers to the loss measured by the difference between the reconstructed association matrix and the original association matrix, reflecting the local similarity between soil samples; the global classification loss refers to the loss measured by comparing the difference between the pollution level predicted by the model and the true label, reflecting the global classification performance of the model. By combining these two losses, the performance of the model in terms of local correlation and global classification accuracy can be optimized simultaneously.
[0061] Specifically, the objective function of joint optimization is defined as: , where MSE is the local similarity loss and CE is the global classification loss. is a hyperparameter that controls the balance between the two losses. The value range of is [0,1] and can be adjusted according to the actual application scenario. When it is close to 1, the model pays more attention to local similarity; when When it is close to 0, the model pays more attention to global classification accuracy. During the training process, stochastic gradient descent (SGD) or other optimization algorithms can be used to perform multiple rounds of training on the parameters of the data encoder and the fully connected layer based on the weighted loss to achieve joint optimization. For example, the learning rate can be set to 0.001 and dynamically adjusted during training to increase the convergence speed of the model.
[0062] Preferably, in the joint optimization process, the operation steps can be further refined. For example, for the hyperparameters In the cross-validation process, the data set can be divided into multiple subsets and different The model performance under the value is finally selected to minimize the loss of the validation set. value. In addition, to improve the robustness of the model, regularization terms such as L2 regularization can be introduced to prevent the model from overfitting. In terms of the selection of optimization algorithms, in addition to stochastic gradient descent (SGD), the Adam optimizer can also be considered, whose adaptive learning rate feature can further improve the training efficiency. In addition, an early stopping mechanism can be introduced to stop training in advance when the loss on the validation set no longer decreases significantly to avoid overfitting.
[0063] In some embodiments, predicting the soil pollution level based on the result of the joint optimization to obtain a treatment plan is specifically as follows: taking the pollution level label generated by the last round of training as the basis for the final treatment plan.
[0064] It should be noted that in the present invention, predicting the soil pollution level based on the results of joint optimization is the ultimate goal of the entire method, and its purpose is to output accurate pollution level labels through the optimized model, thereby providing a scientific basis for the formulation of soil pollution control plans. The joint optimization here refers to optimizing the model parameters by comprehensively considering the local similarity loss and the global classification loss to achieve a balance between local correlation and global classification accuracy. The final pollution level prediction result is obtained by evaluating and analyzing the optimized model, and its accuracy directly affects the effectiveness of the subsequent control plan.
[0065] Specifically, when predicting the soil pollution level, it is first necessary to generate the final pollution level label based on the results of joint optimization, that is, the model parameters obtained by minimizing the weighted loss function. In the final stage of model training, a set of optimized model parameters will be obtained, which can enable the model to achieve the best balance between local similarity and global classification accuracy. By applying these parameters to the fully connected layer and the data encoder, the model can output the pollution level prediction results for each soil sample. These prediction results are usually expressed in the form of probabilities, that is, the probability that each sample belongs to different pollution levels. Finally, the pollution level label of each sample can be determined based on these probability values. For example, a threshold can be set, and when the probability that a sample belongs to a certain pollution level exceeds the threshold, it is classified as that pollution level.
[0066] Preferably, when predicting the soil pollution level, the operation steps can be further refined. For example, when determining the pollution level label, a multi-classification strategy can be introduced, such as using the Softmax function to normalize the probability of the model output to ensure that the sum of the pollution level probabilities of each sample is 1. In addition, in order to improve the reliability of the prediction results, an ensemble learning method can be used to perform weighted averaging of multiple optimized model results to obtain a more stable pollution level prediction. In practical applications, the pollution level can also be subdivided according to the specific situation of soil pollution, such as dividing the pollution level into slight pollution, moderate pollution and severe pollution, and formulating corresponding control measures according to different pollution levels.
[0067] The above-mentioned embodiments of the present invention have the following beneficial effects: Through soil data collection and preprocessing, the present invention can convert complex soil sample information into standardized feature vectors, providing a clear data basis for subsequent analysis. Constructing a multidimensional data model and using low-dimensional representation of features can effectively reduce the complexity of data processing, while retaining key information and improving analysis efficiency. The joint optimization method of local similarity loss and global classification loss can balance the correlation and classification accuracy between soil samples, making the prediction of pollution level more accurate. In addition, through specific preprocessing steps and similarity calculation methods, the accuracy and consistency of the data can be further improved, providing a more reliable basis for soil pollution control. Finally, predicting the soil pollution level based on the results of joint optimization can provide a scientific basis for the formulation of control plans and improve the pertinence and effectiveness of soil pollution control.
[0068] In addition, the present invention also covers detailed optimization of the data processing and analysis process. For example, using a long short-term memory network (LSTM) as a data encoder can effectively capture the time series characteristics of soil data and further improve the accuracy of feature extraction. Through standardized processing and specific similarity calculation methods, information such as the geographical location, pollution source, and pollution type of soil samples can be better processed to ensure the consistency and comparability of the data. These optimization measures not only improve the accuracy of soil pollution analysis, but also enhance the robustness and adaptability of the method, enabling it to better cope with complex and changeable soil pollution scenarios.
[0069] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0070] The above descriptions are only some preferred embodiments of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A soil pollution analysis and control method based on big data analysis, characterized in that: The following steps are involved: Soil data collection and preprocessing; Construct a multi-dimensional data model of soil pollution based on the pre-processed soil data; Obtaining a low-dimensional representation of the features of the multi-dimensional data model through a data encoder; Reconstructing an association matrix of the multidimensional data model based on the feature low-dimensional representation, and calculating a local similarity loss based on the reconstructed association matrix and an original association matrix of the multidimensional data model; Using a clustering-based and highly robust analysis method, the low-dimensional representation of the features is input to classify the soil samples and generate pollution level labels; A fully connected layer is set after the data encoder, and the low-dimensional representation of the feature is input into the fully connected layer to obtain an intermediate analysis result; Calculating a global classification loss based on the pollution level labels and the intermediate analysis results; Performing joint optimization based on the local similarity loss and the global classification loss; The soil pollution level is predicted based on the result of the joint optimization to obtain a treatment plan.
2. According to the soil pollution analysis and control method based on big data analysis in claim 1, it is characterized in that: The soil data collection and preprocessing specifically include: for each soil sample, synthesizing the geographical location, pollution type and pollution source information of the sample into a set of data, and then preprocessing the set of data; All the features of the preprocessed data are sent to the trained feature vector model to obtain the feature vector of each feature and the feature vectors of all the features are averaged as the comprehensive feature vector of the soil sample; The common pollution sources, common geographical locations and common pollution types of different soil samples were preprocessed to obtain the common pollution sources, common geographical locations and common pollution types with the same expression.
3. The soil pollution analysis and control method based on big data analysis according to claim 2 is characterized in that: The preprocessing of the group of data specifically includes: standardizing the geographic location coordinates, removing fuzzy expressions in the pollution type, removing repeated pollution source information, segmenting the data with specific separators, and removing irrelevant features and features with a length less than a certain threshold; The preprocessing of common pollution sources, common geographical locations and common pollution types of different soil samples specifically includes: for common pollution sources, standardization processing, unification of name writing order and normalization of pollution source identification; for common geographical locations, coordinate standardization, removal of redundant spaces, segmentation of data with specific separators, removal of irrelevant features and features with a length less than a certain threshold; for common pollution types, standardization processing, removal of ambiguous expressions, segmentation of data with specific separators, removal of irrelevant features and features with a length less than a certain threshold.
4. The soil pollution analysis and control method based on big data analysis according to claim 3 is characterized in that: Constructing a soil pollution multidimensional data model based on the preprocessed soil data specifically includes: taking each soil sample as a node of the multidimensional data model; Using the comprehensive feature vector of each soil sample as the node feature of the sample in the multi-dimensional data model; Calculate the similarity between the common pollution source, common geographical location and common pollution type of the two soil samples and set the corresponding similarity threshold. Among the three attributes of the common pollution source, common geographical location and common pollution type of the two soil samples, if the similarity of one attribute exceeds the corresponding similarity threshold, establish an edge of the attribute between the two soil sample nodes.
5. The soil pollution analysis and control method based on big data analysis according to claim 4 is characterized in that: For common pollution sources and common pollution types, feature overlap is used to calculate similarity, and for common geographical locations, the inverse of the Euclidean distance is used as the similarity metric.
6. The soil pollution analysis and control method based on big data analysis according to any one of claims 1 to 5, characterized in that: When obtaining the characteristic low-dimensional representation of the multidimensional data model through the data encoder, a two-layer long short-term memory network (LSTM) is used as the data encoder, the input of each layer of the long short-term memory network is the characteristic low-dimensional representation of the previous layer, and the output is the characteristic low-dimensional representation of the current layer, and the input of the first layer is the comprehensive feature vector of the soil sample.
7. The soil pollution analysis and control method based on big data analysis according to claim 6 is characterized in that: The local similarity loss is calculated based on the reconstructed association matrix and the original association matrix of the multidimensional data model. Specifically, the objective function of the local similarity loss is designed to minimize the mean square error loss between the reconstructed association matrix and the original association matrix. : In the formula, is the reconstructed correlation matrix; is an element in the reconstructed association matrix, indicating the probability of predicting the association between node i and node j, with a value range of [0,1]; A is the original association matrix of the multidimensional data model; is an element in the original association matrix of the multidimensional data model, and its value is 0 or 1; N is the number of nodes in the multidimensional data model.
8. The soil pollution analysis and treatment method based on big data analysis according to claim 7 is characterized in that: The global classification loss is calculated based on the pollution level label and the intermediate analysis result as follows: The cross entropy loss function between the pollution level label and the intermediate analysis result is defined as the global classification loss : Where, C is the intermediate analysis result; represents the probability that node i belongs to category c, and its value range is [0,1]; Y is the pollution level label, Represents the true category label of node i, which takes a value of 0 or 1; N is the number of nodes in the multidimensional data model; C is the total number of pollution level categories.
9. The soil pollution analysis and treatment method based on big data analysis according to claim 8 is characterized in that: The joint optimization based on the local similarity loss and the global classification loss is specifically as follows: Using global classification loss and local similarity loss to achieve a balance between them, that is, In the formula, is the weighted loss; is a hyperparameter set empirically; In getting the weighted loss Then, using the stochastic gradient descent algorithm, based on the weighted loss The parameters of the data encoder and the fully connected layer are trained for multiple rounds, and the parameters of the data encoder and the fully connected layer are jointly optimized through training.
10. The soil pollution analysis and control method based on big data analysis according to claim 9 is characterized in that: Predicting the soil pollution level based on the result of the joint optimization to obtain a treatment plan is specifically as follows: taking the pollution level label generated by the last round of training as the basis for the final treatment plan.