A soil pollution data automatic fusion method and system based on multi-modal alignment and attribute supplement
By using multimodal alignment and attribute supplementation methods, soil pollution data is processed automatically, which solves the problems of low preprocessing efficiency and insufficient spatiotemporal consistency of multi-source heterogeneous data, and generates a high-quality fused dataset that is suitable for pollution source identification and risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-04
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies are inefficient in preprocessing multi-source heterogeneous soil pollution data, lack key attributes, and suffer from insufficient spatiotemporal consistency, making it difficult to achieve automated, high-quality data fusion.
A multimodal alignment and attribute supplementation method is adopted. Multimodal data is processed through an automated cleaning and structural transformation module, missing key attributes are supplemented by a multidimensional information complementarity model, and precise temporal and spatial alignment is achieved through a spatiotemporal heterogeneous data alignment engine, and finally data fusion is performed.
It significantly improves the preprocessing efficiency of multi-source soil pollution data, generates a high-quality fusion dataset with complete attributes and spatiotemporal consistency, solves the problems of low efficiency and unstable quality in traditional methods, and provides a more reliable data foundation.
Smart Images

Figure CN121435165B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental data processing technology, and in particular to an automated fusion method and system for soil pollution data based on multimodal alignment and attribute supplementation. Background Technology
[0002] In practical work on the fusion of multi-source heterogeneous soil pollution data, three core technical bottlenecks are commonly encountered. First, the data preprocessing stage is inefficient. Traditional methods heavily rely on manual intervention to identify outliers, remove duplicates, and standardize formats from census, monitoring, and survey data. This is not only time-consuming but also results in inconsistent processing quality due to variations in human experience and operational methods. Especially when dealing with hundreds of thousands or even more multimodal data points, manual cleaning and structuring often become bottlenecks in the overall process. Second, differences in acquisition frequency and spatial reference between multi-source data lead to insufficient spatiotemporal consistency. For example, remote sensing data is often based on monthly or weekly grids, while ground monitoring data is mostly from annual or even lower-frequency stations. When the time steps and spatial coordinate systems of the two are mismatched, simple overlay will produce significant temporal misalignment and spatial bias, making it difficult to accurately reflect the spatiotemporal evolution of pollution. Third, the lack of key attributes is common. This is manifested in the fact that key information such as the process of pollution source section, the status of pollution source existence, and the coordinate system in daily soil monitoring are incomplete in historical data and existing survey records. Traditional methods often use manual filling, mean filling, or direct discarding of missing data, which is not only inefficient but also easy to introduce large biases, thus significantly reducing the usability of fused data.
[0003] To address the aforementioned issues, the industry has proposed alternative solutions including manual-assisted fusion, single-modality-first fusion, and traditional statistical fusion. For example, in manual-assisted fusion, technical personnel and business experts collaborate using tools such as Excel, Python scripts, and ArcGIS for data cleaning, attribute completion, and spatial stitching. While feasible for small-scale, small-sample data scenarios, it suffers from significant drawbacks in large-scale, multi-source heterogeneous data scenarios, including long processing cycles, high subjectivity, and unstable spatiotemporal calibration accuracy. Similarly, single-modality-first solutions often focus on detailed soil survey data, simply matching a few key fields from pollution source survey data into the detailed survey dataset. Remote sensing data is only used for trend reference, and missing attributes are handled using simple methods such as mean filling or similar replacement. This hinders the realization of the complementary value of multimodal data, resulting in insufficient precision and reliability of the fusion results. Traditional statistical fusion methods can achieve data aggregation and rough correlation to some extent, but they generally lack fine alignment for spatiotemporally heterogeneous data and intelligent attribute supplementation mechanisms based on domain knowledge and historical big data, making it difficult to simultaneously meet the efficiency, completeness, and accuracy requirements of data fusion. Summary of the Invention
[0004] To overcome the shortcomings of the prior art, the purpose of this invention is to provide an automated fusion method and system for soil pollution data based on multimodal alignment and attribute supplementation. This invention solves the problem that existing technologies struggle to automate and achieve high-quality fusion of multi-source heterogeneous soil pollution data in terms of preprocessing efficiency, key attribute completion accuracy, and spatiotemporal multimodal alignment.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] An automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation includes:
[0007] Collect raw data on multimodal soil pollution;
[0008] Outlier identification and processing, duplicate data identification and removal, missing value labeling, and data format parsing were performed on the raw multimodal soil pollution data to obtain a standardized structured dataset;
[0009] Using a multidimensional information complementarity model, attribute verification is performed on the standardized structured dataset, and missing key attributes are automatically predicted and supplemented based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes.
[0010] Time interpolation and alignment are performed on remote sensing monitoring data with different monitoring frequencies and time steps in a complete soil pollution standardization dataset, and time-aligned data under a unified time reference is formed according to a preset target time step.
[0011] Spatial coordinate calibration and spatial interpolation fusion are performed on remote sensing grid data with different spatial resolutions and spatial coordinate systems in a soil pollution standardization dataset with complete attributes, and ground monitoring station data are used to obtain spatially aligned data.
[0012] The spatially aligned data and the temporally aligned data are then fused to obtain the final fused data.
[0013] An automated fusion system for soil pollution data based on multimodal alignment and attribute supplementation includes:
[0014] A multimodal data access module is used to collect raw data on multimodal soil pollution.
[0015] The automated cleaning and structural transformation module is used to identify and process outliers, identify and remove duplicate data, mark missing values, and parse data formats from raw multimodal soil pollution data to obtain a standardized structured dataset.
[0016] The multidimensional information complementarity model module is used to perform attribute verification on standardized structured datasets using the multidimensional information complementarity model, and to automatically predict and supplement missing key attributes based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes.
[0017] The time alignment module is used to perform time interpolation and alignment processing on remote sensing monitoring data and ground monitoring station data with different monitoring frequencies and time steps in a complete soil pollution standardization dataset, and form time-aligned data under a unified time reference according to a preset target time step.
[0018] The spatial alignment module is used to perform spatial coordinate calibration and spatial interpolation fusion on remote sensing grid data with different spatial resolutions and spatial coordinate systems and ground monitoring station data in a soil pollution standardization dataset with complete attributes, so as to obtain spatially aligned data.
[0019] The fusion module is used to fuse the spatially aligned data and the temporally aligned data to obtain the final fused data.
[0020] The present invention discloses the following technical effects:
[0021] This invention provides an automated fusion method and system for soil pollution data based on multimodal alignment and attribute supplementation. The method includes: collecting raw multimodal soil pollution data; performing outlier identification and processing, duplicate data identification and removal, missing value marking, and data format parsing on the raw multimodal soil pollution data to obtain a standardized structured dataset; using a multidimensional information complementarity model to perform attribute verification on the standardized structured dataset, and automatically predicting and supplementing missing key attributes based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes; performing time interpolation and alignment processing on remote sensing monitoring data with different monitoring frequencies and different time steps and ground monitoring station data in the standardized soil pollution dataset with complete attributes, and forming time-aligned data under a unified time reference according to a preset target time step; performing spatial coordinate calibration and spatial interpolation fusion on remote sensing grid data with different spatial resolutions and different spatial coordinate systems and ground monitoring station data in the standardized soil pollution dataset with complete attributes, to obtain spatially aligned data; and fusing the spatially aligned data and the time-aligned data to obtain the final fused data. This method offers significant advantages over existing technologies in multi-source soil pollution data fusion: First, in the preprocessing stage, the core design of the automated cleaning and structural transformation module, with its built-in optimized cleaning logic and standardized structural transformation rules, automates the entire process of outlier handling, duplicate removal, and format unification for multi-source data. This eliminates the time-consuming and unstable quality issues associated with manual methods, achieving over 90% higher preprocessing efficiency compared to traditional manual methods in scenarios with massive multimodal data. Furthermore, the fixed logic ensures consistent cleaning quality. Second, regarding data integrity, addressing the challenge of missing key attributes such as process flow and pollution source status, a multi-dimensional information complementarity model based on a soil pollution knowledge graph and historical large datasets is introduced. This model automatically identifies and supplements missing key attributes, and a built-in verification mechanism ensures the authenticity and rationality of the supplemented results. This fundamentally solves the problem of traditional methods struggling to balance supplementation efficiency and accuracy, generating a standardized dataset with complete attributes. Furthermore, in terms of spatiotemporal fusion, a spatiotemporal heterogeneous data alignment engine was constructed. By utilizing the temporal analysis capabilities of large models combined with interpolation and alignment algorithms, time synchronization of data from different monitoring frequencies was achieved. Through spatial coordinate calibration and spatial correlation analysis, accurate spatial matching of remote sensing grid data and ground station data was completed, reducing the spatiotemporal consistency error of the final fused dataset to approximately 3% or less. This provides a more reliable and accurate data foundation for subsequent work such as pollution source identification and risk assessment. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart of an automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation provided in an embodiment of the present invention;
[0024] Figure 2 A comparison chart of the accuracy of supplementary data attributes for soil pollution in industrial areas provided in this embodiment of the invention;
[0025] Figure 3 A comparison chart of data preprocessing efficiency provided for embodiments of the present invention;
[0026] Figure 4 A comparison chart of spatiotemporal consistency errors of soil pollution data in mining areas provided in an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] like Figure 1 As shown, this invention provides an automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation, comprising:
[0030] Step 100: Collect raw data on multimodal soil pollution;
[0031] Step 200: Perform outlier identification and processing, duplicate data identification and removal, missing value labeling, and data format parsing on the original multimodal soil pollution data to obtain a standardized structured dataset;
[0032] Step 300: Using a multidimensional information complementarity model, perform attribute verification on the standardized structured dataset, and automatically predict and supplement missing key attributes based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes.
[0033] Step 400: Perform time interpolation and alignment on remote sensing monitoring data and ground monitoring station data with different monitoring frequencies and time steps in the soil pollution standardization dataset with complete attributes, and form time-aligned data under a unified time reference according to the preset target time step.
[0034] Step 500: Perform spatial coordinate calibration and spatial interpolation fusion on remote sensing grid data with different spatial resolutions and spatial coordinate systems and ground monitoring station data in the soil pollution standardization dataset with complete attributes to obtain spatially aligned data;
[0035] Step 600: Merge the spatially aligned data and the temporally aligned data to obtain the final merged data.
[0036] Specifically, the implementation process is as follows:
[0037] First, the collection and preprocessing of multimodal raw data are completed. Users first integrate various types of data, including pollution source survey data, detailed soil survey data, and routine soil monitoring data, to ensure that all types of data fully cover pollution-related information in the target study area. Then, the system's built-in automated cleaning and structure transformation module is activated. This module automatically identifies and handles outliers, duplicates, and formatting errors in the data. Simultaneously, through built-in structured transformation logic, it converts unstructured and semi-structured data of different formats into a unified structured data format, laying a standardized data foundation for subsequent fusion processes. This process requires minimal manual intervention, significantly improving the efficiency of data cleaning and standardization.
[0038] After data preprocessing, the system proceeds to the automatic supplementation of key missing attributes. The system automatically invokes a multidimensional information complementarity model; users only need to import relevant soil pollution knowledge graphs and historical datasets for the region as model support data. Based on domain knowledge rules and big data association analysis algorithms, the model performs attribute verification on the preprocessed structured data, accurately identifying missing key attributes such as process flow and pollution source status. It then automatically supplements and verifies missing attributes based on the inherent relationships between data and historical experience data, generating a standardized dataset with complete attributes. This effectively solves the data quality problems caused by missing attributes in traditional data fusion.
[0039] After attribute supplementation is completed, the spatiotemporal heterogeneous data alignment engine is activated to perform spatiotemporal fusion of the data. Users need to set the target time step parameter in the system according to the actual situation of the monitoring data. The engine will leverage the time-series analysis capabilities of a large model to perform time-dimensional interpolation and alignment processing on multi-source data from different monitoring frequencies, resolving the issue of inconsistent time steps among multi-frequency monitoring data. Simultaneously, the system will automatically invoke a spatial fusion algorithm to calibrate the spatial coordinates of remote sensing grid data and ground monitoring station data. Accurate fusion is achieved through spatial correlation analysis of the two types of data using a large model. After fusion, the system will output a high-quality, spatiotemporally consistent soil pollution fusion dataset, which can be directly used for subsequent research and applications such as pollution source identification and risk assessment.
[0040] Furthermore, the process of identifying and processing outliers, identifying and removing duplicate data, marking missing values, and parsing data formats from the original multimodal soil pollution data to obtain a standardized structured dataset includes:
[0041] The pollution source survey data, detailed soil survey data, and routine soil monitoring data in the multimodal soil pollution raw data are uniformly integrated to obtain the raw data set;
[0042] Based on the improved Z-score anomaly detection algorithm, outliers are identified in the numerical fields of the original dataset, and the identified outliers are marked to obtain the original dataset with anomaly labels.
[0043] Outliers in the original dataset with anomaly markers are automatically repaired using interpolation methods. The outliers are replaced with interpolated estimates, resulting in the original dataset with complete outlier repair.
[0044] The original dataset after outlier repair is processed by identifying and removing duplicate data to obtain a deduplicated dataset.
[0045] Missing fields in the deduplicated dataset are marked with missing values, and the missing values are explicitly represented in the data records in a uniform labeling format, resulting in a multimodal dataset with missing value labels.
[0046] By calling the XML parsing interface and the JSON parsing interface, the unstructured text data and semi-structured tabular data in the multimodal dataset with missing markers are automatically parsed and fields are extracted. The parsed content is then uniformly converted into structured fields to obtain a structured multi-source dataset.
[0047] By standardizing the data format and fields of a structured multi-source dataset, a standardized structured dataset is obtained.
[0048] Specifically, we first collect multimodal data such as pollution source survey data D1, detailed soil survey data D2, and routine soil monitoring data D3 to construct an original dataset D={D1, D2, D3}.
[0049] This solution introduces a data time-series weight ω. t (t represents the data collection time, with more recent data having higher weight), the optimization formula is:
[0050] ;
[0051] in, , ;
[0052] Improve the accuracy of anomaly detection in time-series data. After anomalies are confirmed, interpolation is used for repair. Simultaneously, an automated structure transformation system calls a bidirectional XML and JSON parsing interface to uniformly convert unstructured text data and semi-structured tabular data into a structured data format D'={d ij |i=1..n,j=1..m}, where i and j represent the indexes of the data samples, used to distinguish different independent data samples. n is the number of data samples, and m is the initial attribute dimension. After preprocessing, the data cleaning efficiency is improved by more than 80% compared with traditional manual methods.
[0053] Furthermore, using a multidimensional information complementarity model, attribute verification is performed on the standardized structured dataset, and missing key attributes are automatically predicted and supplemented based on the multidimensional information complementarity model, resulting in a standardized soil pollution dataset with complete attributes, including:
[0054] Based on knowledge graphs in the field of soil pollution and historical big data sets of soil pollution, we conduct statistical analysis on entities and semantic relationships involving pollution source types, process stages, pollutant types, land use types and spatial locations in standardized structured datasets, and construct attribute association matrices.
[0055] Identify key missing attributes from standardized structured datasets;
[0056] For each record with a key missing attribute, the complete attribute fields in the record are extracted as input feature vectors, and the input feature vectors are mapped to the attribute association matrix to construct an attribute prediction training sample set.
[0057] Based on the attribute prediction training sample set, a gradient boosting tree attribute prediction model is constructed and trained. A loss function with prediction error as the target is adopted, and the parameters of the gradient boosting tree attribute prediction model are optimized through iterative method. At the same time, a domain knowledge constraint term is introduced during the model training and prediction process. When the prediction result conflicts with the existing entity relationship in the soil pollution domain knowledge graph, the conflict prediction result is penalized and adjusted to obtain a gradient boosting tree attribute prediction model that satisfies the domain knowledge constraint.
[0058] Records with key missing attributes in the standardized structured dataset and their corresponding input feature vectors are input into a gradient boosting tree attribute prediction model that satisfies domain knowledge constraints. The key missing attributes of each record are automatically predicted and supplemented to obtain a standardized soil pollution dataset with complete attributes.
[0059] Specifically, firstly, based on the knowledge graph G=(V, E) in the field of soil pollution (V is the entity node, including pollution source type, process section, etc.; E is the association edge, representing the semantic relationship between entities), and combined with the historical big data dataset H, an attribute association matrix M∈R is constructed. k Where R represents the real number space, indicating that the elements of the attribute association matrix M are real numbers; k is the "potential association attribute dimension," referring to the number of newly added attributes that semantically relate to existing attributes, mined from the domain knowledge graph and historical data. For missing attributes y∈Y (Y being a set of key missing attributes such as process segment, current status, etc.), an attribute prediction model based on Gradient Boosting Tree (GBRT) is adopted. The model input is the existing complete attribute feature vector X=(x1, x2, ..., x...). m The output is the predicted value of the missing attribute, ŷ. The model loss function uses mean squared error:
[0060] ;
[0061] Update parameters using gradient descent.
[0062] ;
[0063] in, The set of parameters to be optimized for the GBRT (Gradient Attribute Prediction Model) is the core object to be adjusted by the gradient descent method. The learning rate is a hyperparameter that controls the step size of parameter updates and is used to balance the convergence speed and stability of the model. It is a loss function For parameters The gradient reflects the "direction and rate of change" of the loss function with respect to the parameters. Gradient descent updates the loss function along the opposite direction of the gradient. This achieves the minimization of the loss function. This is the model's loss function (here, mean squared error), used to measure the predicted value. The magnitude of the error between the actual value y and the true value y is the objective function for parameter optimization.
[0064] Simultaneously, a domain knowledge constraint term C( When the predicted value When a conflict arises with entity associations in the knowledge graph, a penalty mechanism is triggered to adjust the prediction result. The final output is a complete dataset D''={d... ij The method, with attributes of type |i=1..n, j=1..m+k}, achieved an accuracy rate of over 92% in attribute completion testing, thus solving the problem of low data availability caused by missing attributes in traditional methods.
[0065] Furthermore, the process of performing time interpolation and alignment on remote sensing monitoring data with different monitoring frequencies and time steps in the complete soil pollution standardized dataset, and forming time-aligned data under a unified time reference according to a preset target time step, includes:
[0066] Time series of remote sensing monitoring data and time series of ground monitoring station data are extracted from a complete soil pollution standardization dataset. The monitoring frequency and time step information corresponding to each time series are identified. The time series are then paired according to spatial location and pollutant type to obtain a set of multi-frequency monitoring time series to be aligned.
[0067] The set of multi-frequency monitoring time series to be aligned is input into a time-series alignment model built based on a large model to obtain a set of time-series feature representations for time alignment.
[0068] Based on the set of temporal feature representations used for time alignment, a dynamic time warping algorithm is used to calculate the temporal distance matrix between each pair of time series;
[0069] Based on the time distance matrix, the optimal time alignment path that minimizes the overall matching cost is searched, and remote sensing monitoring data with different time steps and ground monitoring station data are aligned and mapped under the optimal time alignment path to obtain an aligned time series data set on the same reference time axis.
[0070] Based on the aligned time series data set on the same reference time axis, the time axis is resampled and normalized according to the preset target time step to obtain time-aligned data under a unified time base and a unified time step.
[0071] Furthermore, the spatial coordinate calibration and spatial interpolation fusion of remote sensing grid data with different spatial resolutions and spatial coordinate systems and ground monitoring station data in the complete soil pollution standardized dataset are performed to obtain spatially aligned data, including:
[0072] Extract remote sensing grid data and ground monitoring station data from a complete, attribute-standardized soil pollution dataset;
[0073] The spatial resolution of remote sensing grid data, the spatial coordinates of the center points of each grid, the spatial coordinates of ground monitoring station data, and the corresponding monitored pollution concentration values are obtained. Coordinate system identification and transformation are performed to obtain a set of grid and station mapping data with coordinate calibration completed.
[0074] Based on the grid and station mapping data set with coordinate calibration completed, the Kriging interpolation method is used to spatially interpolate the pollution concentration data of the ground monitoring stations, and the grid pollution concentration data set calculated by spatial interpolation is obtained.
[0075] The grid pollution concentration data set calculated by spatial interpolation is fused with the original remote sensing grid data under a unified spatial coordinate system. For grid data with different spatial resolutions, resampling or grid aggregation is used to unify the resolution. For grid cells with spatial overlap or blank areas, interpolation is performed to complete and boundary consistency is processed to obtain spatially aligned data under a unified spatial coordinate system and unified spatial resolution.
[0076] Specifically, regarding time alignment, to address the issue of inconsistent time steps in multi-frequency monitoring data, a Transformer-based time series alignment model is adopted. Let the time series of data from different frequencies be T1={t...} 11 , t 12 , ..., t 1a} (Step size τ1), T2={t 21 , t 22 , ..., t 2b (Step size τ2). T1 and T2 represent the time series sets corresponding to monitoring data at two different frequencies, and are the input objects of the time series alignment model; where T1 corresponds to a time series with a step size τ1, and T2 corresponds to a time series with a step size τ2; a and b are the number of time points contained in time series T1 and T2, respectively, that is, T1 has a time points, and T2 has b time points. In order to solve the problem of inconsistent time steps of multi-frequency monitoring data, the time step size τ1 is limited.
[0077] Temporal features F=Transformer(T1, T2) are extracted using a model encoder, and then the temporal distance matrix is calculated using the Dynamic Time Warping (DTW) algorithm.
[0078] ;
[0079] Finding the optimal alignment path:
[0080] ;
[0081] This achieves temporal unification of data of different time lengths. F is the temporal feature extracted from time series T1 and T2 by the Transformer encoder, representing an abstract representation of the original time series features (including information such as the time series' changing patterns and trends). Transformer(T1, T2) represents the encoding operation process of the Transformer-based temporal alignment model; that is, inputting two time series T1 and T2 of different time lengths into the model and outputting the corresponding temporal feature F. It is a single element in the time distance matrix, representing the temporal feature corresponding to the i-th time point of time series T1. The temporal features corresponding to the j-th time point of time series T1 The absolute difference between the two points in time is used to measure the "degree of difference" of characteristics at the two points in time. It is the temporal feature corresponding to the i-th time point of time series T1; It is the temporal feature corresponding to the j-th time point of time series T2. P is the optimal alignment path, which is the time point matching method that minimizes the sum of feature differences between the two time series. It is the operational logic for finding the optimal path. "" represents the sum of the characteristic differences at all points in time along a certain path.
[0082] In terms of spatial alignment, for remote sensing grid data S (resolution r×r) and ground monitoring station data (coordinates (x, y)... i y j Using a combination of spatial interpolation and coordinate calibration, the mapping relationship between the grid center point coordinates and the station coordinates is first established, and the pollution concentration values within the grid are calculated using the Kriging interpolation formula.
[0083] ;
[0084] Where λᵢ is the interpolation weight, satisfying Σλᵢ=1, and passing through a semi-variogram function;
[0085] ;
[0086] Spatial correlations are calculated to optimize weights. Finally, a large model fusion layer is used to fuse the features of the spatiotemporally aligned data, outputting a fused dataset. Validation shows that the spatiotemporal consistency error of the fused data is less than 5%, significantly outperforming traditional methods.
[0087] This embodiment also provides an automated fusion system for soil pollution data based on multimodal alignment and attribute supplementation, including:
[0088] A multimodal data access module is used to collect raw data on multimodal soil pollution.
[0089] The automated cleaning and structural transformation module is used to identify and process outliers, identify and remove duplicate data, mark missing values, and parse data formats from raw multimodal soil pollution data to obtain a standardized structured dataset.
[0090] The multidimensional information complementarity model module is used to perform attribute verification on standardized structured datasets using the multidimensional information complementarity model, and to automatically predict and supplement missing key attributes based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes.
[0091] The time alignment module is used to perform time interpolation and alignment processing on remote sensing monitoring data and ground monitoring station data with different monitoring frequencies and time steps in a complete soil pollution standardization dataset, and form time-aligned data under a unified time reference according to a preset target time step.
[0092] The spatial alignment module is used to perform spatial coordinate calibration and spatial interpolation fusion on remote sensing grid data with different spatial resolutions and spatial coordinate systems and ground monitoring station data in a soil pollution standardization dataset with complete attributes, so as to obtain spatially aligned data.
[0093] The fusion module is used to fuse the spatially aligned data and the temporally aligned data to obtain the final fused data.
[0094] Specifically, this patent's core architecture revolves around a three-tiered fusion structure: preprocessing, attribute supplementation, and spatiotemporal alignment. The multimodal data access module serves as the input layer, supporting batch import of various data types, including pollution source survey data, detailed soil survey data, and daily soil monitoring data. It is compatible with multiple formats, such as structured tables (Excel, CSV), semi-structured documents (XML, JSON, JSONLines, Parquet), and unstructured text (reports, logs). The automated cleaning and structure transformation module handles the core data preprocessing function, incorporating an improved Z-score anomaly detection algorithm and multi-format parsing interfaces to remove data noise and standardize formats. The multidimensional information complementarity model module, supported by a knowledge graph and historical large datasets in the soil pollution field, integrates a gradient boosting tree (GBRT) attribute prediction model to automatically identify and accurately supplement missing attributes. The spatiotemporal heterogeneous data alignment engine serves as the core processing layer, including a time alignment module, a spatial alignment module, and a fusion module, relying on a large model to complete the spatiotemporal fusion of cross-source data. The data storage and output module is responsible for storing and exporting the standardized dataset, the attribute-complete dataset, and the final fused dataset, supporting Excel, CSV, and JSON formats. Bidirectional export of formats such as Lines and Parquet is supported, allowing for compatibility with subsequent applications.
[0095] The system adopts a "layered and progressive + module linkage" logical architecture. Each module independently undertakes a specific function, while also achieving full-process automated connection through data interfaces: The input layer (multimodal data access module) breaks down data format and type barriers to achieve "one-stop" import of multi-source data; the preprocessing layer (automatic cleaning and structure transformation module) converts heterogeneous data into a unified structured format through standardized processing logic, laying the foundation for subsequent processes; the core processing layer (multidimensional information complementarity model module + spatiotemporal heterogeneous data alignment engine) first completes attribute integrity repair, and then achieves accurate spatiotemporal dimension alignment to form high-quality fused data; the output layer (data storage and output module) provides flexible data storage and export solutions to meet the application needs of different scenarios.
[0096] Regarding the core parameters of the algorithm, the temporal weight ω of the improved Z-score anomaly detection algorithm is... t The feasible range is 0.6~1.0, the feasible range for outlier detection threshold is 2.5~3.5, and outlier repair supports Lagrange interpolation and cubic spline interpolation; the feasible range for the learning rate η of the GBRT attribute prediction model is 0.1~0.3, the feasible range for the number of decision trees is 150~300, and the feasible range for the attribute association matrix dimension is R. 50x25 ~R 80x40 The time alignment module allows for an encoder layer count of 4-12, a hidden layer dimension of 256-1024, and a target time step of 0.5 hours to 30 days. The spatial alignment module allows for a remote sensing grid resolution r of 10m-100m, supports exponential, Gaussian, and spherical models for Kriging interpolation semivariograms, and is compatible with WGS spatial coordinates. t -84, UTM, National Geodetic Coordinate System 2000; In terms of data processing efficiency parameters, data cleaning efficiency is improved by more than 80% compared with traditional manual methods, the processing time for a single batch of data can be implemented in the range of 5~20 hours, there is no theoretical upper limit to the number of original data samples supported, it can adapt to the parallel processing of n (n≥10000) multimodal data, the initial attribute dimension m can be implemented in the range of 50~80 items, and the supplemented attribute dimension m+k can be implemented in the range of 80~120 items; In terms of fusion accuracy parameters, the attribute supplementation accuracy can be implemented in the range of The accuracy is 92%~98% (optimal value: 95%~96%), the spatiotemporal consistency error is ≤5%, the spatial fusion deviation is ≤5m, and the time alignment deviation is ≤2 hours. In terms of compatibility parameters, it supports multi-frequency monitoring data alignment with time steps τ1 to τ2 (monthly level), and supports importing and exporting formats such as Excel, CSV, XML, JSON, Lines, and Parquet, adapting to the fusion processing needs of different types of soil pollution data such as industrial areas, farmland, and mining areas.
[0097] Furthermore, this invention also provides a data fusion experiment on soil pollution in urban industrial areas:
[0098] Data collection:
[0099] An industrial zone in a certain city was selected as the target area. Data from the 2017 pollution source census (D1, including pollution emission records of 320 enterprises, totaling 86,000 data points), detailed soil survey data (D2, data on heavy metal and organic pollutant content at 120 monitoring points, totaling 38,000 data points), and routine soil monitoring data (D3, data on land use type and soil pollutant monitoring, totaling 21,000 data points) were collected to construct the original dataset D.
[0100] Preprocessing parameters:
[0101] An improved Z-score anomaly detection algorithm is adopted, with time-series weights ω. t Value range: 0.6~1.0 (data from the last 3 months ω) t =1.0, 3~6 months ω t =0.8, 6 months and above ω t =0.6), outlier threshold |Z i Version 3.0 uses linear interpolation to repair outliers; it also uses a bidirectional XML and JSON parsing interface to convert unstructured monitoring reports and semi-structured table data into structured data D' in a unified CSV format.
[0102] Additional attribute parameters:
[0103] Import the knowledge graph G (containing 120 entity nodes and 350 associated edges) in the field of soil pollution and the historical dataset H (150,000 records) of the region from 2018 to 2022, and construct the attribute association matrix M∈R. 60x30 (Initial attribute dimension m=60, potential associated attribute dimension k=30); The learning rate η of the Gradient Boosting Tree (GBRT) model is 0.1~0.3 (optimal η=0.2), and the number of decision trees is 100~300 (optimal 200).
[0104] Spatiotemporal alignment parameters:
[0105] The time step was set to 1 day (target time resolution). Time alignment was achieved using the Transformer temporal alignment model (6 encoder layers, 512 hidden layer dimensions) and the Dynamic Time Warping (DTW) algorithm. Spatial alignment was achieved using Kriging interpolation. The remote sensing grid data resolution was r = 30m × 30m. A spherical model was selected for the semi-variogram, and the interpolation weight λᵢ was optimized through spatial correlation analysis.
[0106] Integration process:
[0107] The process is executed in the order of "preprocessing - attribute supplementation - spatiotemporal alignment", and is fully automated without human intervention.
[0108] Experimental results:
[0109] Product structure identification data:
[0110] The system outputs a fusion dataset F containing 145,000 samples and 90 attribute dimensions (including 30 supplementary attributes such as process flow and pollution source existence status). The data format is uniformly CSV, the spatial coordinates adopt the WGS-84 coordinate system, and the timestamps are accurate to the hour.
[0111] Results data:
[0112] like Figure 2 As shown, data preprocessing took only 8 hours, which is 93.3% more efficient than traditional manual methods (about 120 hours); the accuracy rate of attribute supplementation was 94.7%, of which the accuracy rate of process attribute supplementation was 95.2% and the accuracy rate of pollution source existence status supplementation was 96.1%; the spatiotemporal consistency error was 3.2%, the spatial fusion deviation between remote sensing data and ground station data was ≤2.8m, and the time alignment deviation of multi-frequency data was ≤0.5 hours.
[0113] Furthermore, this invention also provides a data fusion experiment on heavy metal pollution in farmland soil:
[0114] Data collection:
[0115] A plain farmland area (80 km²) was selected, and pollution source survey data D1 (62,000 data points of agricultural non-point source pollution and industrial point source pollution), detailed soil survey data D2 (53,000 data points of heavy metal (Cd, Pb, As, etc.) content at 180 monitoring points), and routine soil monitoring data D3 (35,000 data points of planting type and soil environment monitoring) were collected to form the original data D.
[0116] Preprocessing parameters:
[0117] Improved Z-score algorithm time-series weights ω t Value range: 0.7~1.0 (data from the past month ω) t =1.0, 1~3 months ω t =0.9, 3 months or more ω t =0.7), outlier detection threshold |Zᵢ|>2.8 (optimal value), outliers are repaired using cubic spline interpolation; the target format for structured transformation is Parquet.
[0118] Additional attribute parameters:
[0119] The domain knowledge graph G contains 90 entity nodes and 280 associated edges. The historical dataset H consists of 120,000 data points on farmland soil environmental quality in the region from 2019 to 2023. The GBRT model has a learning rate η of 0.15–0.25 (optimal 0.2), a number of decision trees of 150–250 (optimal 200), and an attribute association matrix M ∈ R. 50x25 .
[0120] Spatiotemporal alignment parameters:
[0121] The target time step is 7 days (weekly), the Transformer model has 4 encoder layers and 256 hidden layer dimensions; the remote sensing grid resolution is r=50m×50m, and the exponential model is used for the Kriging interpolation semivariogram function.
[0122] Integration process:
[0123] To initiate the automated fusion process, only manual setting of the target time step and spatial resolution parameters is required.
[0124] 2. Test Results:
[0125] Product structure identification data:
[0126] The fusion dataset F contains 150,000 samples with 75 attribute dimensions (25 supplementary attributes); the data storage format is Parquet with a compression ratio of 3:1; the spatial coordinates adopt the National Geodetic Coordinate System 2000; and the time series is continuous without any breaks.
[0127] Results data:
[0128] like Figure 3 As shown, the preprocessing time was 10 hours, which is 93.3% more efficient than the traditional manual method (about 150 hours); the attribute supplementation accuracy was 93.5%, of which the accuracy of supplementing the source attributes of agricultural non-point source pollution was 94.1%; the spatiotemporal consistency error was 3.8%, the spatial fusion deviation was ≤3.5m, and the time alignment deviation was ≤1.2 hours.
[0129] Furthermore, this invention also provides a data fusion experiment for monitoring soil pollution and remediation in mining areas:
[0130] Data collection:
[0131] Select an abandoned mining area and its surrounding area, and collect pollution source survey data D1 (78,000 data points on historical mining pollution and emissions from surrounding enterprises), detailed soil survey data D2 (65,000 data points on heavy metal content and pH value from 200 monitoring points), and soil monitoring data D3 (28,000 data points on mining area topography, vegetation cover, and soil pollutant content) to construct the original data D.
[0132] Preprocessing parameters:
[0133] Improved Z-score algorithm time-series weights ω t =0.65~1.0 (data from the last two months ω) t =1.0, 2~5 months ω t =0.85, 5 months or more ω t =0.65), outlier detection threshold |Z i |>3.2, outliers are repaired using Lagrange interpolation; the target format for structured conversion is JSON Lines.
[0134] Additional attribute parameters:
[0135] The domain knowledge graph G contains 150 entity nodes and 420 associated edges. The historical dataset H consists of 180,000 records of mining area pollution and remediation data from 2017 to 2022. The GBRT model has a learning rate η of 0.12 to 0.28 (optimal 0.18), 200 to 300 decision trees (optimal 250), and an attribute association matrix M ∈ R. 70x35 .
[0136] Spatiotemporal alignment parameters:
[0137] The target time step is 12 hours, the Transformer model has 8 encoder layers and 768 hidden layer dimensions; the remote sensing grid resolution is r=20m×20m, and the Gaussian model is used for the Kriging interpolation semivariogram function.
[0138] Integration process:
[0139] The entire process is automated, and it synchronously outputs real-time fusion results and historical trend comparison data.
[0140] Experimental results:
[0141] Product structure identification data: The fusion dataset F contains 163,000 samples with 105 attribute dimensions (35 supplementary attributes); the data format supports bidirectional export of JSON Lines and CSV, the spatial coordinates are compatible with WGS-84 and UTM coordinate systems, and the time resolution reaches 12 hours.
[0142] Results data: such as Figure 4 As shown, the preprocessing time was 11 hours, which is 91.5% more efficient than the traditional manual method (about 130 hours); the attribute supplementation accuracy was 95.3%, of which the process supplementation accuracy of mining section was 96.4% and the attribute supplementation accuracy of pollution control measures was 95.8%; the spatiotemporal consistency error was 2.9%, the spatial fusion deviation was ≤2.5m, and the time alignment deviation was ≤0.3 hours.
[0143] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0144] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation, characterized in that, include: Collect raw data on multimodal soil pollution; Outlier identification and processing, duplicate data identification and removal, missing value labeling, and data format parsing were performed on the raw multimodal soil pollution data to obtain a standardized structured dataset; Using a multidimensional information complementarity model, attribute verification is performed on the standardized structured dataset, and missing key attributes are automatically predicted and supplemented based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes. Time interpolation and alignment are performed on remote sensing monitoring data with different monitoring frequencies and time steps in a complete soil pollution standardization dataset, and time-aligned data under a unified time reference is formed according to a preset target time step. Spatial coordinate calibration and spatial interpolation fusion are performed on remote sensing grid data with different spatial resolutions and spatial coordinate systems in a soil pollution standardization dataset with complete attributes, and ground monitoring station data are used to obtain spatially aligned data. The spatially aligned data and the temporally aligned data are then fused to obtain the final fused data. Using a multidimensional information complementarity model, attribute verification is performed on a standardized structured dataset. Based on this model, missing key attributes are automatically predicted and supplemented, resulting in a standardized soil pollution dataset with complete attributes, including: Based on knowledge graphs in the field of soil pollution and historical big data sets of soil pollution, we conduct statistical analysis on entities and semantic relationships involving pollution source types, process stages, pollutant types, land use types and spatial locations in standardized structured datasets, and construct attribute association matrices. Identify key missing attributes from standardized structured datasets; For each record with a key missing attribute, the complete attribute fields in the record are extracted as input feature vectors, and the input feature vectors are mapped to the attribute association matrix to construct an attribute prediction training sample set. Based on the attribute prediction training sample set, a gradient boosting tree attribute prediction model is constructed and trained. A loss function with prediction error as the target is adopted, and the parameters of the gradient boosting tree attribute prediction model are optimized through iterative method. At the same time, a domain knowledge constraint term is introduced during the model training and prediction process. When the prediction result conflicts with the existing entity relationship in the soil pollution domain knowledge graph, the conflict prediction result is penalized and adjusted to obtain a gradient boosting tree attribute prediction model that satisfies the domain knowledge constraint. Records with key missing attributes in the standardized structured dataset and their corresponding input feature vectors are input into a gradient boosting tree attribute prediction model that satisfies domain knowledge constraints. This model automatically predicts and supplements the key missing attributes of each record, resulting in a standardized soil pollution dataset with complete attributes.
2. The automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation according to claim 1, characterized in that, Raw data on multimodal soil pollution include: Pollution source survey data, detailed soil survey data, and routine soil monitoring data.
3. The automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation according to claim 2, characterized in that, The process involves outlier identification and processing, duplicate data identification and removal, missing value labeling, and data format parsing of the original multimodal soil pollution data to obtain a standardized structured dataset, including: The pollution source survey data, detailed soil survey data, and routine soil monitoring data in the multimodal soil pollution raw data are uniformly integrated to obtain the raw data set; Based on the improved Z-score anomaly detection algorithm, outliers are identified in the numerical fields of the original dataset, and the identified outliers are marked to obtain the original dataset with anomaly labels. Outliers in the original dataset with anomaly markers are automatically repaired using interpolation methods. The outliers are replaced with interpolated estimates, resulting in the original dataset with complete outlier repair. The original dataset after outlier repair is processed by identifying and removing duplicate data to obtain a deduplicated dataset. Missing fields in the deduplicated dataset are marked with missing values to obtain a multimodal dataset with missing value tags; By calling the XML parsing interface and the JSON parsing interface, the unstructured text data and semi-structured tabular data in the multimodal dataset with missing markers are automatically parsed and fields are extracted. The parsed content is then uniformly converted into structured fields to obtain a structured multi-source dataset. By standardizing the data format and fields of a structured multi-source dataset, a standardized structured dataset is obtained.
4. The automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation according to claim 3, characterized in that, The expression for the improved Z-score anomaly detection algorithm is as follows: ; in, , ; in, This represents the anomaly detection value of the i-th data point at time t. This represents the original monitoring value of the i-th data point at time t. This represents the weighted average of the data series corresponding to time t. This represents the weighted standard deviation of the data series corresponding to time t. This indicates the weight of the data collection time.
5. The automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation according to claim 3, characterized in that, The process involves interpolating and aligning remote sensing monitoring data with different monitoring frequencies and time steps from a standardized soil pollution dataset with complete attributes, along with ground monitoring station data, to form time-aligned data under a unified time reference according to a preset target time step. This includes: Time series of remote sensing monitoring data and time series of ground monitoring station data are extracted from a complete soil pollution standardization dataset. The monitoring frequency and time step information corresponding to each time series are identified. The time series are then paired according to spatial location and pollutant type to obtain a set of multi-frequency monitoring time series to be aligned. The set of multi-frequency monitoring time series to be aligned is input into a time-series alignment model built based on a large model to obtain a set of time-series feature representations for time alignment. Based on the set of temporal feature representations used for time alignment, a dynamic time warping algorithm is used to calculate the temporal distance matrix between each pair of time series; Based on the time distance matrix, the optimal time alignment path that minimizes the overall matching cost is searched, and remote sensing monitoring data with different time steps and ground monitoring station data are aligned and mapped under the optimal time alignment path to obtain an aligned time series data set on the same reference time axis. Based on the aligned time series data set on the same reference time axis, the time axis is resampled and normalized according to the preset target time step to obtain time-aligned data under a unified time base and a unified time step.
6. The automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation according to claim 3, characterized in that, The spatial coordinate calibration and spatial interpolation fusion of remote sensing grid data with different spatial resolutions and spatial coordinate systems in the complete soil pollution standardized dataset, along with ground monitoring station data, yields spatially aligned data, including: Extract remote sensing grid data and ground monitoring station data from a complete, attribute-standardized soil pollution dataset; The spatial resolution of remote sensing grid data, the spatial coordinates of the center points of each grid, the spatial coordinates of ground monitoring station data, and the corresponding monitored pollution concentration values are obtained. Coordinate system identification and transformation are performed to obtain a set of grid and station mapping data with coordinate calibration completed. Based on the grid and station mapping data set with coordinate calibration completed, the Kriging interpolation method is used to spatially interpolate the pollution concentration data of the ground monitoring stations, and the grid pollution concentration data set calculated by spatial interpolation is obtained. The grid pollution concentration data set calculated by spatial interpolation is fused with the original remote sensing grid data under a unified spatial coordinate system to obtain spatially aligned data under a unified spatial coordinate system and unified spatial resolution.
7. An automated fusion system for soil pollution data based on multimodal alignment and attribute supplementation, characterized in that, The system is used to implement an automated fusion method for soil pollution data based on multimodal alignment and attribute supplementation as described in any one of claims 1 to 6, the system comprising: A multimodal data access module is used to collect raw data on multimodal soil pollution. The automated cleaning and structural transformation module is used to identify and process outliers, identify and remove duplicate data, mark missing values, and parse data formats from raw multimodal soil pollution data to obtain a standardized structured dataset. The multidimensional information complementarity model module is used to perform attribute verification on standardized structured datasets using the multidimensional information complementarity model, and to automatically predict and supplement missing key attributes based on the multidimensional information complementarity model to obtain a standardized soil pollution dataset with complete attributes. The time alignment module is used to perform time interpolation and alignment processing on remote sensing monitoring data and ground monitoring station data with different monitoring frequencies and time steps in a complete soil pollution standardization dataset, and form time-aligned data under a unified time reference according to a preset target time step. The spatial alignment module is used to perform spatial coordinate calibration and spatial interpolation fusion on remote sensing grid data with different spatial resolutions and spatial coordinate systems and ground monitoring station data in a soil pollution standardization dataset with complete attributes, so as to obtain spatially aligned data. The fusion module is used to fuse the spatially aligned data and the temporally aligned data to obtain the final fused data.
Citation Information
Patent Citations
Environment detection data acquisition and processing method and system based on big data
CN120850234A
Soil remediation real-time monitoring method utilizing coupling of multispectrum of unmanned aerial vehicle and sensing of internet of things
CN120948762A