A method and system for sediment remediation based on multi-modal machine learning
By combining multimodal machine learning with multi-objective optimization algorithms, high-precision real-time diagnosis and dynamic optimization remediation of sediment pollution in complex waters have been achieved, improving governance efficiency and eco-friendliness, and solving the problems of single data, static decision-making and insufficient equipment accuracy in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUANJIAN ECOLOGICAL RESTORATION (BEIJING) CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-29
Smart Images

Figure CN122114473A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental remediation technology, and more specifically to a method and system for sediment remediation based on multimodal machine learning. Background Technology
[0002] Currently, sediment pollution is a significant contributing factor to the degradation of aquatic ecosystems, and its remediation effectiveness directly impacts water quality safety and ecological restoration. Traditional sediment remediation technologies primarily rely on physical dredging, chemical solidification, and bioremediation.
[0003] However, multiple technical bottlenecks exist in complex aquatic environments: data acquisition and analysis are limited in scope, and pollution diagnosis accuracy is insufficient. Existing technologies typically rely on discrete manual sampling and laboratory testing (such as atomic absorption spectrometry for heavy metal determination), resulting in low sampling density (usually ≤5 points / km²) and long cycles (24-72 hours per test), making it difficult to capture the spatiotemporal dynamics of pollutants. While existing remote sensing technologies can acquire large-scale spectral data, they lack effective integration with ground sensors and laboratory data, making it difficult to accurately model the actual correlation between indicators such as NDVI (Normalized Difference Vegetation Index) and sediment organic matter. Remediation decisions rely on static empirical models, exhibiting poor dynamic adaptability. Mainstream remediation schemes are often based on expert experience databases or statistical regression models, with rigid decision-making logic that cannot respond in real time to sudden changes in environmental parameters (such as pH and dissolved oxygen) or changes in pollution diffusion trends. Existing research attempts to introduce single machine learning algorithms (such as SVM for classifying pollution types), but has not solved the coupling problem between multi-objective optimization (efficiency, cost, and ecological impact) and dynamic decision-making. Engineering implementation relies on heavy equipment, resulting in insufficient precision and eco-friendliness. While existing precision remediation technologies (such as in-situ microbial remediation) can reduce ecological disturbance, the application of chemicals largely relies on manual operation, leading to problems such as uneven coverage (error > 15%) and response delays. Furthermore, remediation effectiveness evaluation still primarily relies on post-remediation laboratory testing, lacking real-time feedback and closed-loop control during implementation. The technology chain is fragmented, lacking an integrated "sensing-decision-execution" system. Currently, the monitoring, diagnosis, decision-making, and execution stages of sediment remediation are mostly completed by independent systems, resulting in data flow disruptions and decision-making delays. For example, monitoring data needs to be manually imported into the decision model, and remediation instructions are then issued to equipment through an independent control system, with an overall response time exceeding 48 hours, failing to meet the real-time governance needs of dynamic pollution scenarios. Existing automated remediation equipment (such as orbital UAVs) can only execute preset path tasks, lacking autonomous decision-making capabilities based on environmental feedback.
[0004] Therefore, how to achieve dynamic pollution diagnosis, real-time optimization of remediation plans, and automated and precise execution, thereby improving the efficiency of sediment treatment in complex waters, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for bottom sediment remediation based on multimodal machine learning, which realizes dynamic pollution diagnosis, real-time optimization of remediation schemes and automated and precise execution, thereby improving the efficiency of bottom sediment treatment in complex waters.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A sediment remediation method based on multimodal machine learning includes: Multimodal data is obtained by acquiring satellite remote sensing data, in-situ sensor data, and laboratory data of the target area; Based on the multimodal data, preprocessing and feature selection are performed sequentially to obtain a key feature set; Based on the aforementioned key feature set, combined with historical pollution data and location information of ecologically sensitive areas, risk-cost trade-off recommendations are generated. An optimized repair scheme is generated based on the aforementioned risk-cost trade-off recommendation and multi-objective optimization algorithm; Based on the optimized repair scheme, a repair task is performed in the target area, and the status data of the target area is obtained during the execution process; The optimized repair scheme is adjusted based on the status data.
[0008] Preferably, the preprocessing specifically includes: Based on the satellite remote sensing data, sensor noise and cloud interference are eliminated, and pollution-related spectral features are extracted to obtain preprocessed remote sensing data; Preprocessed sensor data is obtained by reducing sensor noise and filling in missing data based on the in-situ sensor data. Based on the laboratory data, discrete points are converted into a continuous spatial distribution to eliminate format differences and obtain preprocessed experimental data. Based on the preprocessed remote sensing data, the preprocessed sensor data, and the preprocessed experimental data, a feature alignment algorithm is used to eliminate the spatiotemporal resolution differences of the multi-source data, unify the temporal resolution, and obtain a preprocessed multimodal dataset.
[0009] Preferably, the feature selection specifically includes: The preprocessed multimodal dataset is input into the random forest feature selection layer; The random forest feature selection layer evaluates the feature weights and ranks the importance of the data in the preprocessed multimodal dataset based on the Gini index, and selects the Top-30 data as key data to form the key feature set.
[0010] Preferably, risk-cost trade-off recommendations are generated, specifically including: Based on the key feature set, the spatiotemporal prediction layer of the Long Short-Term Memory network is input. The first layer of 128 LSTM units captures long-distance temporal dependence, and the second layer of 64 units extracts high-level features and outputs a pollution parameter prediction sequence for a future preset duration, thus obtaining a pollution migration trend heatmap. The pollution migration trend heatmap, the historical pollution data, and the location information of the ecologically sensitive area are input into the XGBoost risk modeling layer to generate the risk-cost trade-off recommendation.
[0011] Preferably, the risk-cost trade-off recommendation specifically includes: Based on the aforementioned pollution migration trend heatmap, the pollutant concentration and redox potential are obtained; Risk assessment is conducted based on the pollutant concentration and the redox potential, and risk areas are divided into different risk levels. When the distance between the risk area and the ecologically sensitive area is less than a set value, the risk level is increased by a preset value, and a pollution risk heat map of the target area is generated. Based on the aforementioned risk level, corresponding remediation priority recommendations are provided; The risk-cost trade-off recommendation is generated based on the pollution risk heat map and the remediation priority recommendations.
[0012] Preferably, an optimized repair solution is generated, specifically including: The state space is defined based on a deep Q-network, consisting of pollution feature vectors, environmental parameters, and equipment status; the action space is defined as a combination of remediation processes. Based on the aforementioned risk-cost trade-off, the corresponding combination of repair processes in the action space is selected in the state space; Based on the multi-objective optimization algorithm combined with the comprehensive objective function, the process parameters of the repair process combination are optimized. Pareto optimal solution set is generated by non-dominated sorting and crowding degree calculation. The optimized repair scheme is generated by matching specific repair process parameters through a decision knowledge base.
[0013] Preferably, the comprehensive objective function F Specifically: ; ; ; ; ; ; ; in, ω 1. ω 2 and ω The numbers 3 represent weights, reflecting the importance of restoration efficiency, cost, and ecological disturbance, respectively. , , For the original sub-target, , , To correspond to the normalized index, This represents the objective function for maximizing repair efficiency. This represents the average total amount of pollution removed per unit time; a higher value indicates higher remediation efficiency. A i Indicates the first i Pollution removal volume in each treated area T Indicates the repair time. This represents the overall cost objective function. This represents the sum of total energy and material costs incurred during the repair process; a smaller value indicates lower costs. H Indicates equipment energy consumption. g Indicates electricity price, S Indicates the dosage of microbial agent. j Indicates the unit price of the inoculant. This represents the objective function for ecological disturbance. The overall quantitative assessment measures the degree of disturbance to aquatic ecosystems caused by restoration activities; a smaller value indicates higher eco-friendliness. α and β All represent weighting coefficients. L Indicates the mortality rate of benthic organisms. V This indicates the rate of change in water turbidity.
[0014] Preferably, the repair task is performed, specifically including: Based on the optimized repair scheme, a structured instruction set containing spatial location, operation type, and parameter thresholds is generated; The structured instruction set is used to control the amount of pesticides dispensed by the drone swarm in the target area; The underwater robot collects and provides feedback on the status data of the target area in real time.
[0015] Preferably, adjusting the optimization and repair scheme based on the status data specifically includes: The status data includes the redox potential value and pollutant concentration of the sediment in the target area; Threshold judgment is made based on the redox potential value and the pollutant concentration. When the change of the redox potential value exceeds the first threshold or the pollutant concentration rebounds to the second threshold within a preset time period, the optimized remediation scheme is updated and adjusted based on the state data.
[0016] A sediment remediation system based on multimodal machine learning includes: a data acquisition module, a data processing module, a suggestion generation module, a scheme generation module, a status data acquisition module, and a scheme adjustment module; The data acquisition module is used to acquire multimodal data composed of satellite remote sensing data, in-situ sensor data and laboratory data of the target area; The data processing module is used to perform preprocessing and feature selection on the multimodal data in sequence to obtain a key feature set; The suggestion generation module is used to generate risk-cost trade-off suggestions based on the key feature set combined with historical pollution data and location information of ecologically sensitive areas. The solution generation module is used to generate an optimized repair solution based on the risk-cost trade-off suggestion and the multi-objective optimization algorithm; The status data acquisition module is used to perform a repair task in the target area based on the optimized repair scheme, and to acquire the status data of the target area during the execution process; The scheme adjustment module is used to adjust the optimization and repair scheme based on the status data.
[0017] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a method and system for sediment remediation based on multimodal machine learning, which has the following beneficial effects: 1. By integrating a hybrid algorithm model of Random Forest and Long Short-Term Memory (LSTM) network, satellite remote sensing spectral data, in-situ sensor monitoring parameters (such as pH, redox potential, and pollutant concentration) and laboratory test results are dynamically analyzed to achieve high-precision real-time diagnosis of sediment pollution type, spatial distribution and migration trend.
[0018] 2. Combining the multi-objective optimization algorithm (NSGA-II), based on pollution characteristics and remediation cost constraints, the system autonomously matches the optimal remediation scheme (such as in-situ microbial induced mineralization remediation or ex-situ dredging technology), and drives drone swarms and underwater biomimetic robots to perform precise tasks (such as targeted drug delivery and pollution hotspot sampling) through end-to-end control protocols.
[0019] 3. Deep integration of machine learning algorithms with remediation engineering: The XGBoost algorithm is used to build a pollution risk prediction model, and LSTM time series analysis is used to predict the pollution diffusion pattern. At the same time, the decision path is dynamically optimized to form a closed-loop system of "perception-decision-execution-feedback". This improves efficiency by more than 30% compared with the traditional manual intervention mode. It is especially suitable for the precise treatment of bottom sediments in complex waters such as large lakes and estuaries, and has dynamic adaptability and potential for large-scale application.
[0020] 4. By combining multimodal data fusion with machine learning algorithms, dynamic pollution diagnosis, real-time optimization of remediation plans, and automated and precise execution can be achieved, significantly improving the efficiency and eco-friendliness of sediment treatment in complex waters. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0022] Figure 1 The present invention provides a flowchart of a sediment remediation method based on multimodal machine learning.
[0023] Figure 2 This is a schematic diagram of a drone-underwater robot collaborative operation provided by the present invention.
[0024] Figure 3 This is a schematic diagram of a sediment remediation system based on multimodal machine learning, provided by the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Example 1 like Figure 1 As shown, this embodiment of the invention discloses a method for sediment remediation based on multimodal machine learning, comprising: Multimodal data is obtained by acquiring satellite remote sensing data, in-situ sensor data, and laboratory data of the target area; Based on the multimodal data, preprocessing and feature selection are performed sequentially to obtain the key feature set; Risk-cost trade-off recommendations are generated based on key feature sets combined with historical pollution data and location information of ecologically sensitive areas. An optimized repair scheme is generated based on risk-cost trade-off recommendations and multi-objective optimization algorithms; Based on the optimized repair scheme, a repair task is performed in the target area, and the status data of the target area is obtained during the execution process; Adjust and optimize the repair plan based on status data.
[0027] Example 2 This invention discloses a method for sediment remediation based on multimodal machine learning, comprising: Multimodal data is obtained by acquiring satellite remote sensing data, in-situ sensor data, and laboratory data of the target area.
[0028] Preferably, in this embodiment, by accessing Level-2A surface reflectance data from the Sentinel-2 MSI satellite platform, which has a spatial resolution of 10 meters and a revisit period of 5 days, a high-frequency data source is provided for dynamic monitoring of the target area. The NDVI of band 8A and the chlorophyll-a concentration of band 3 are extracted from it. That is, the satellite remote sensing data includes: normalized vegetation index and chlorophyll-a concentration.
[0029] Preferably, NDVI (Normalized Difference Vegetation Index) is calculated using band 8A (near-infrared, 842nm) B8A and band 4 (red, 665nm) B4: NDVI = (B8A - B4) / (B8A + B4); Normalized Difference Vegetation Index (NDVI) is used to indirectly indicate the organic matter content in sediment. NDVI < 0.1 indicates bare sediment areas, 0.1 ≤ NDVI ≤ 0.3 indicates transitional areas (sparse vegetation or sediment-vegetation mixed areas), and > 0.3 indicates vegetation-covered areas.
[0030] Preferably, the chlorophyll-a concentration can be estimated using a reflectance ratio model based on the red-edge band. As an example, the calculation formula can be expressed as:
[0031] in, The coefficient B, representing the water reflectance, needs to be calibrated using measured data from the target water area.
[0032] Preferably, in this embodiment, a Sea-Bird SBE 37-SMP CTD multi-parameter probe is used, which can measure pH (accuracy ±0.01, range 0-14), OPR (oxidation-reduction potential, accuracy ±5mV, range 2000-2000mV), and conductivity (accuracy ±0.05mS / cm, range 0-100mS / cm). The probe is deployed in a 500m×500m grid in the target area, with the grid density increased to 100m×100m in the core pollution area to improve monitoring accuracy. For cadmium ion detection, a Metrohm 6.0502.100 ion-selective electrode is used, with a detection limit of 0.01ppm and a linear range of 0.05-10ppm. Three-point calibration (0.1ppm, 1ppm, and 5ppm Cd²) is automatically performed daily. + (Standard solutions) are used to ensure data reliability. pH, redox potential, conductivity, and heavy metal concentration are obtained as in-situ sensor data.
[0033] In this embodiment, the heavy metal concentration is: cadmium ion concentration.
[0034] The IoT network is constructed using the LoRaWAN protocol, with a node spacing of ≤1km and a data packet transmission interval of 1 minute. Outliers, such as abnormal data with a redox potential mutation rate >100mV / min, are removed using the Z-score algorithm (threshold |Z|>3). Dual probes are deployed at key nodes, and manual verification is triggered when the data difference is >10%, thus achieving redundancy backup.
[0035] Preferably, sediment columnar samples (depth 0-30cm) are collected at 100m intervals in turbidity anomaly areas (>50NTU) identified by satellite imagery and in areas where Cd >0.5ppm is monitored by sensors. Samples were freeze-dried, passed through a 100-mesh sieve, and then digested with aqua regia (HNO3:HCl=3:1) using a CEM Mars6 microwave digester. Arsenic (As) (m / z=75), lead (Pb) (m / z=208), and cadmium (Cd) (m / z=111) were detected using an Agilent 7900 ICP-MS. Quality control indicators were a spiked recovery rate of 85-115% and a repeat sample RSD <5%. The obtained arsenic, lead, and cadmium concentrations were used as laboratory data.
[0036] The key feature set is obtained by sequentially preprocessing and feature selection based on multimodal data.
[0037] Preferably, the preprocessing specifically includes: Preprocessed remote sensing data is obtained by eliminating sensor noise and cloud interference based on satellite remote sensing data and extracting pollution-related spectral features. Preprocessed sensor data is obtained by reducing sensor noise and filling in missing data based on in-situ sensor data. Based on laboratory data, discrete points are converted into a continuous spatial distribution to eliminate format differences and obtain preprocessed experimental data; Based on preprocessed remote sensing data, preprocessed sensor data, and preprocessed experimental data, a feature alignment algorithm is used to eliminate the spatiotemporal resolution differences of multi-source data, unify the temporal resolution, and obtain a preprocessed multimodal dataset.
[0038] Preferably, in terms of data format, a feature matrix is constructed by integrating multimodal data (number of samples = 10,000, number of features = 28), with each row corresponding to a spatiotemporal feature vector of a 10m×10m grid cell, to achieve spatial alignment and dimensional uniformity of preprocessed remote sensing data, preprocessed sensor data, and preprocessed experimental data.
[0039] Preferably, the spatiotemporal KNN algorithm (k=5, spatiotemporal radius 50m × 1 hour) is used to fill missing values in all three types of data. For heterogeneous data characteristics, such as NDVI range [-1,1] and Cd concentration at the ppm level, the Z-score normalization method is used, with the following formula: X_normalized=(X-μ) / σ Where μ is the feature mean and σ is the standard deviation, ensuring data dimensionality consistency.
[0040] Preferably, feature selection specifically includes: The preprocessed multimodal dataset is input into the random forest feature selection layer; The random forest feature selection layer evaluates the feature weights and ranks the importance of the data in the preprocessed multimodal dataset based on the Gini index, and selects the top-30 data as key data to form a key feature set.
[0041] Preferably, a random forest algorithm is used, with n_estimators=500 to balance computational efficiency and the stability of feature importance evaluation. The depth of a single tree is limited by max_depth=10 to prevent overfitting. The criterion='gini' is selected to evaluate feature splitting quality based on the Gini index. Feature importance is calculated based on the average Gini index decrease value. The Top-30 key features with a cumulative contribution rate ≥95% are selected from the preprocessed multimodal dataset, for example: NDVI (weight 0.15) and Band 8A (near-infrared, weight 0.12) derived from satellite remote sensing data reflect sediment vegetation cover and spectral anomalies; The rate of change of redox potential (Δredox potential / Δt, weight 0.18) and Cd concentration (weight 0.22) in the sensor time-series characteristics characterize pollution dynamics and environmental stability; The laboratory-interpolated As concentration (weight 0.09) is used as the true ground value for heavy metal pollution; The dimensionality reduction effect verification shows that the AUC of the logistic regression model trained on the filtered feature set increased from 0.82 to 0.89, indicating that redundant features were effectively removed and the discriminative ability of the feature space was significantly enhanced.
[0042] Risk-cost trade-off recommendations are generated based on key feature sets combined with historical pollution data and location information of ecologically sensitive areas.
[0043] Preferably, risk-cost trade-off recommendations are generated, specifically including: Based on the key feature set input into the spatiotemporal prediction layer of the Long Short-Term Memory network, the first layer of 128 LSTM units captures long-distance temporal dependence, the second layer of 64 units extracts high-level features, and outputs a pollution parameter prediction sequence for a future preset duration, thus obtaining a pollution migration trend heatmap. Based on the pollution migration trend heat map, historical pollution data, and location information of ecologically sensitive areas, the XGBoost risk modeling layer is used to generate risk-cost trade-off recommendations.
[0044] Preferably, the spatiotemporal prediction layer of the Long Short-Term Memory Network fills missing values using the spatiotemporal KNN algorithm (k=5, spatiotemporal radius 50m×1 hour), and uses the Savitzky-Golay filter (window length 15, polynomial order 2) to filter noise in the key feature set to form the original time series sequence before standardization. The input layer receives standardized time-series data, and the hidden layer extracts features through two layers of bidirectional LSTM units—the first layer captures the long-range dependencies of the time-series data (such as the impact of day-night cycles on OPR) by outputting the complete sequence. The second layer extracts the core trends of time series characteristics through the final state output (such as whether the Cd concentration is on an upward trend). The output layer maps high-level features to a spatial dimension, enabling cross-modal prediction of "time series → spatial distribution," and providing spatiotemporally coupled pollution prediction data for the XGBoost risk modeling layer. The output layer maps temporal features to a 10m×10m spatial grid through a fully connected layer (5 neurons), enabling prediction of the spatial distribution of Cd concentration over the next 6 hours, and outputting a pollution migration trend heatmap with a spatial resolution of 10m×10m.
[0045] Preferably, ecologically sensitive areas include: mangrove distribution areas (marked with a 50-meter buffer zone) and seasonal fish spawning areas, with the presence or absence of pollution events indicated by binary labels: 0 / 1.
[0046] Preferably, the latitude and longitude coordinates of the mangrove distribution area are obtained through protected area boundary data, remote sensing extraction (such as Sentinel-2+ random forest), or ground RTK measurement.
[0047] Preferably, the historical pollution data includes: a set of 1000+ historical pollution events (including pollution type and remediation effect).
[0048] Preferred risk-cost trade-off recommendations are as follows: Based on the heat map of pollution migration trends, pollutant concentrations and redox potentials were obtained; Risk assessments are conducted based on pollutant concentrations and redox potentials, and risk zones are divided into different risk levels. When the distance between the risk area and the ecological sensitive area is less than the set value, the risk level is increased by a preset value, and a pollution risk heat map of the target area is generated; Obtain corresponding repair priority suggestions based on the risk level; Generate risk-cost trade-off suggestions based on the pollution risk heat map and the repair priority suggestions.
[0049] Preferably, risk assessment is carried out based on pollutant concentration and redox potential, and the risk probability (0-1) of each 10m×10m grid cell in the target area is output. The judgment conditions for risk areas of different risk levels are specifically as follows: Low-risk area (probability < 0.3): Cd ≤ 1 ppm and redox potential ≥ -100 mV; Medium-risk area (probability 0.3 - 0.7): 1 ppm < Cd ≤ 2 ppm or -200 mV ≤ redox potential < -100 mV; High-risk area (probability > 0.7): Cd > 2 ppm and redox potential < -200 mV; Through the above risk levels and corresponding standards, the model realizes the quantitative spatial distribution prediction and level division of pollution risks, providing data-driven decision support for environmental management.
[0050] Preferably, when the distance between the risk area and the ecological sensitive area is less than 50 meters, the risk level is increased by 1 level, and the risk levels of each grid cell in the target area are obtained, and then the pollution risk heat map of the target area is obtained: High-risk scenario: Cd = 2.1 ppm, OPR = -210 mV, 50 meters away from the mangrove forest, determined as high risk.
[0051] Medium-risk scenario: Cd = 1.8 ppm, OPR = -150 mV, 150 meters away from the mangrove forest, determined as medium risk.
[0052] Low-risk scenario: Cd = 0.5 ppm, OPR = -50 mV, non-sensitive area, determined as low risk, only quarterly sampling and monitoring are required.
[0053] Preferably, the random forest feature selection layer, the long short-term memory network spatio-temporal prediction layer and the XGBoost risk modeling layer jointly constitute a hybrid machine learning model.
[0054] Generate an optimized repair plan based on the risk-cost trade-off suggestions and the multi-objective optimization algorithm.
[0055] Preferably, generating an optimized repair plan specifically includes: Define the state space based on the deep Q network: pollution feature vector, environmental parameters and equipment status; define the action space as the repair process combination; Based on the risk-cost trade-off, the corresponding repair process combination in the action space is selected in the state space; Based on a multi-objective optimization algorithm combined with a comprehensive objective function, the process parameters of the repair process combination are optimized. Pareto optimal solution set is generated by non-dominated sorting and crowding degree calculation. Then, the specific repair process parameters are matched by a decision knowledge base to generate an optimized repair scheme.
[0056] Preferably, the state space is defined as a five-dimensional vector, designed to comprehensively and accurately represent the dynamic environment and device states of the system. Specifically: Pollution type coding: A classification coding method is used (e.g., heavy metals = 1, organic matter = 2, eutrophication = 3). In this case, the specific code is 1, which clearly indicates heavy metal pollution. Pollution feature vector: Cd average concentration: The 6-hour average concentration value predicted by the LSTM model, which is 1.8 ppm in the current example. When this concentration exceeds the set threshold of 1 ppm, corresponding remediation measures will be triggered. Environmental parameters: Taking water temperature as an example, the current value is 25℃. Water temperature has a significant impact on the efficiency of microbial remediation, and its optimal range is 20-30℃. Equipment status: This includes the drone's battery level, which is currently at 80%. When the battery level drops below 20%, the return-to-home mechanism will be automatically triggered. It also includes the robot's positioning status, which is currently displayed as normal. Abnormal statuses include sensor malfunctions, communication interruptions, and other situations.
[0057] Preferably, the action space includes six types of remediation process combinations, each of which is a multi-technology collaborative strategy. Taking "microbial remediation (Pseudomonas putida) + local dredging" as an example action: Microbial remediation: The selected inoculant was Pseudomonas putida, which is effective against Cd²⁺. + Its adsorption efficiency is ≥85%, and its mechanism of action is mainly to fix heavy metals through cell wall complexation and biomineralization processes.
[0058] Local dredging: The target area is set as the core pollution zone with Cd concentration >2ppm. The dredging depth is dynamically adjusted based on the NSGA-II optimization results, and is 0.8m in this case.
[0059] The reward function quantifies the overall benefits of a decision, and the weighting is based on expert experience and validated with historical data. The formula for calculating the reward function is: R = 0.6 × pollution removal rate + 0.3 × (1 - cost weight) - 0.1 × ecological disturbance index.
[0060] Preferred, comprehensive objective function F Specifically: ; ; ; ; ; ; ; in, ω 1. ω 2 and ω The numbers 3 represent weights, reflecting the importance of restoration efficiency, cost, and ecological disturbance, respectively. , , For the original sub-target, , , To correspond to the normalized index, This represents the objective function for maximizing repair efficiency. This represents the average total amount of pollution removed per unit time; a higher value indicates higher remediation efficiency. A i Indicates the first i Pollution removal volume in each treated area T Indicates the repair time. This represents the overall cost objective function. This represents the sum of total energy and material costs incurred during the repair process; a smaller value indicates lower costs. H Indicates equipment energy consumption. g Indicates electricity price, S Indicates the dosage of microbial agent. j Indicates the unit price of the inoculant. This represents the objective function for ecological disturbance. The overall quantitative assessment measures the degree of disturbance to aquatic ecosystems caused by restoration activities; a smaller value indicates higher eco-friendliness. α and β All represent weighting coefficients. L Indicates the mortality rate of benthic organisms. V This indicates the rate of change in water turbidity.
[0061] Preferably, in this embodiment, α=0.6 and β=0.4, which are obtained based on ecological expert scores.
[0062] Preferably, the constraints include: the repair time is limited to ≤48 hours due to the tidal cycle; the total budget (equipment rental, chemicals, labor, etc.) does not exceed 500,000 yuan; and the dredging depth is limited to the range of [0.3m, 1.2m] to avoid damage to the bottom sediment structure and to ensure the feasibility of the project and environmental safety.
[0063] The preferred method for generating Pareto optimal solutions is as follows: Population initialization: 100 initial solutions are randomly generated, with variables including the amount of microbial agent added, dredging depth, and priority of the remediation area, to construct a diversified decision-making starting point.
[0064] Non-dominated sorting: The solution set is stratified by combining the objective function values, eliminating "dominated solutions" that are comprehensively superior to other solutions, and retaining Pareto front candidate solutions.
[0065] Pareto Front: Generate 20 sets of non-dominated solutions, forming a three-dimensional surface of restoration efficiency, cost, and ecological disturbance, and visualize the multi-objective conflict and trade-off relationship.
[0066] Optimal solution selection: Based on the decision knowledge base, match the actual scenario parameters. For example, for the current Cd concentration of 1.8 ppm, determine the inoculant dosage to be 120 g / m³ using the linear interpolation formula (dosage = 50 + 40 × Cd). 3 The dredging depth was selected as 0.8m to meet the dual constraints of repair efficiency ≥80% and cost ≤450,000 yuan, thus achieving the optimal balance between project objectives and resource constraints.
[0067] This method combines multi-objective mathematical modeling with intelligent optimization algorithms to provide an optimized remediation scheme for heavy metal pollution that balances efficiency, cost, and ecological protection.
[0068] Based on the optimized repair scheme, a repair task is performed in the target area, and the status data of the target area is obtained during the execution process.
[0069] Preferably, the repair task is performed, specifically including: A structured instruction set containing spatial location, operation type, and parameter thresholds is generated based on the optimized repair scheme; Controlling the amount of pesticides dispensed by a drone swarm in a target area based on structured instruction sets; The underwater robot collects and reports the status data of the target area in real time.
[0070] Preferred, such as Figure 2 As shown, the drone swarm is equipped with a hyperspectral imager and a pesticide spraying device. It controls the amount of pesticide added based on a structured instruction set and a Gaussian distribution model, with an addition error of less than 2%. The underwater robot is equipped with a microelectrode sensor array to provide real-time feedback on the redox potential value and heavy metal concentration of the bottom sediment, with a positioning accuracy of ±0.1 meters.
[0071] Preferably, the drone swarm plans its flight path using a Gaussian distribution model (μ is the coordinate of the core pollution area, σ=50m) to precisely spray Pseudomonas putida inoculant, ensuring that the dosage error is controlled within 1.5%. When the wind speed reaches 10m / s, it automatically switches to low-altitude operation mode (flight altitude reduced to 5m) to cope with abnormal weather conditions. During path generation, the flight sub-areas of each drone are allocated according to the drone swarm size (10 drones) using a Gaussian probability density function. The spraying path uses concentric spiral circles spaced 10m apart, and the flight speed is set at 5m / s to ensure sufficient agent settling time. Nordson EFD is used for agent dosing control. The PicoJet™ piezoelectric microdroplet jetting device features a 50nL droplet volume and a 100Hz jetting frequency. Its control logic dynamically adjusts the jetting volume based on real-time flight position (GPS positioning, accuracy ±0.1m) to meet Gaussian concentration gradient requirements. In the error calibration stage, an onboard hyperspectral imager (wavelength range 400-1000nm) monitors the ground agent coverage density in real time and provides feedback to adjust jetting parameters, ensuring an application error of <1.5%. The anomaly response mechanism includes wind speed monitoring, mode switching, and emergency obstacle avoidance. Wind speed monitoring utilizes an FT205 digital anemometer (range 0-30m / s, accuracy ±0.3m / s). When the wind speed ≥10m / s or gust frequency >5Hz, a low-altitude mode is triggered, reducing the flight altitude from the standard 15m to 5m to reduce wind disturbance. It then switches to a terrain-following mode using LiDAR real-time terrain scanning to maintain a constant altitude. Emergency obstacle avoidance relies on 77GHz millimeter-wave radar to detect obstacles and trigger dynamic path replanning (RRT). Algorithm, response time <0.5s).
[0072] Preferably, the underwater robot collects sediment samples at a depth of 30cm using its robotic arm at the coordinate location, displays the Cd concentration in real time, and simultaneously transmits the oxidation-reduction potential via underwater acoustic communication. The hardware configuration of the robotic arm precision sampling system includes a six-degree-of-freedom robotic arm (Blue Robotics ARM5 model) with a load capacity of 2kg and a repeatability of ±0.1m. It uses an inverse kinematics algorithm (IKFast) to calculate joint angles in real time to adapt to complex sediment terrain. It is equipped with a piston-type columnar sediment sampler (5cm inner diameter, sampling depth 0-50cm, hydraulic drive pressure 20kPa), which can automatically trigger sampling based on preset coordinates (x=123.45, y=45.67) or a real-time pollution heat map. The pollution detection and feedback process integrates a microelectrode sensor array, in which Cd²... +The selective electrode has a detection range of 0.01-10ppm and a response time of <30s. The redox potential sensor has a range of -2000 to 2000mV and an accuracy of ±5mV. It integrates multiple parameter detection values (such as Cd=2.1ppm, redox potential=-210mV) every 5 minutes to achieve data fusion. The underwater acoustic communication link adopts FSK frequency shift keying modulation with a bandwidth of 10kHz and a transmission rate of 10kbps. It is equipped with a customized lightweight protocol (the packet header includes CRC-16 checksum) to ensure a packet loss rate of <1%.
[0073] Adjust and optimize the repair plan based on status data.
[0074] Preferred solutions for adjusting and optimizing repair based on status data include: The state data includes the redox potential values and heavy metal concentrations of the sediment in the target area; Threshold judgment is made based on redox potential value and heavy metal concentration. When the change of redox potential value exceeds the first threshold or the pollution concentration rebounds to the second threshold within a preset time period, the remediation plan is updated, adjusted and optimized based on the status data.
[0075] Preferably, when the underwater robot detects a rebound in Cd concentration in the target area exceeding 5% during operation (i.e., the measured concentration increases by more than 5% from the initial value after a decrease before remediation), a second optimization of the remediation plan is immediately triggered. Subsequently, the parameters of the LSTM network are updated using 200 newly collected real-time data points (with a weight of 20%), reducing the prediction error from 7.2% to 6.5%, effectively improving the model's adaptability to dynamic changes in pollution. The remediation parameters are adjusted to increase the bacterial agent dosage to 150 g / m³ (a 25% increase compared to the initial plan) and increase the dredging depth to 1.0 m (breaking the original 0.8 m setting) to address the risk of continuous release of heavy metals from the sediment.
[0076] Preferably, when the Cd concentration exceeds the baseline value after remediation by 5% for three consecutive tests (e.g., the baseline value of 0.27 ppm corresponds to the threshold of 0.284 ppm, and the actual measured value is 0.3 ppm) or the Δ redox potential is >50 mV (e.g., the initial value of -160 mV increases to the current value of -210 mV), the test data is uploaded to the cloud decision engine and marked as an "abnormal event". The state space is recalculated by the reinforcement learning module (Cd=0.3 ppm, redox potential=-210 mV), and new parameters are output by the multi-objective optimization algorithm (dredging depth increases from 0.8 m to 1.0 m, and the amount of microbial agent added is adjusted from 120 g / m³ to 150 g / m³). Finally, the optimized remediation plan is updated and issued for execution.
[0077] Preferably, the end-to-end collaborative verification system achieves time synchronization between the UAV swarm and the underwater robot (error <1ms) through the NTP protocol, and the cloud-based dispatch center dynamically allocates task priorities based on the degree of pollution (e.g., prioritizing high-pollution areas with Cd>2ppm). Its performance indicators cover the positioning accuracy of UAVs ±0.1m (RTK-GPS) and underwater robots ±0.1m (underwater SLAM), the response latency of <10s from data feedback to command update, and the system robustness of the remaining equipment automatically taking over the task and the coverage rate decreasing by <5% when a single node fails. The technical verification results show that the reagent dosing accuracy reaches the Gaussian distribution with a measured error of 1.2% (better than the target <1.5%), the coverage uniformity is improved by 40%, and the Cd rebound rate after dredging depth adjustment is reduced from 6% to 3%, verifying the effectiveness of replanning. In the 10km² area restoration task, the equipment collaborative operation time is shortened by 60% compared with the traditional method, significantly improving the collaborative efficiency.
[0078] Preferably, based on state data (especially high-frequency data after anomaly triggering and new scenario data), the prediction and decision-making model is continuously optimized through two core sub-strategies to ensure its adaptability and accuracy: First, incremental learning (continuous model fine-tuning) is used to update the existing model in small increments using newly collected and filtered valid data (screening criteria: Pearson correlation coefficient > 0.7). A new and old data weight balancing mechanism is introduced (new data weight 20%, old data weight 80%) to prevent model drift. The LSTM uses a loss function of "Loss = 0.8 × Loss_old + 0.2 × Loss_new". The weights of the first two layers are frozen and only the last fully connected layer is fine-tuned (learning rate = 0.0001, batch size = 32, training period = 10). XGBoost uses the partial_fit method to update the model every 100 new data points and dynamically adjusts the Z-score normalization parameter of the new data (μ_new = 0.8μ_old + 0.2μ_new). At the same time, a prediction error threshold of MSE < 8% is set for the test set. If it is exceeded, the entire model is retrained, and cross-validation is performed every 24 hours.
[0079] Secondly, there is transfer learning (cross-scenario knowledge transfer and rapid adaptation). When the system is applied to a new environment or encounters significantly different scenarios, it utilizes a model pre-trained with integrated historical data (covering multiple pollution types, 100,000 samples, and multi-dimensional annotation information) from 10 water scenarios (5 lakes, 3 estuaries, and 2 ports). The model (LSTM inputs standardized multimodal time-series data and outputs pollution concentration prediction, using HuberLoss and specific hyperparameters; XGBoost performs multi-task learning and configures corresponding parameters) and aligns the new scenario data with a 1-hour time window through the high-throughput real-time data pipeline of Apache Kafka. This triggers a fine-tuning mechanism—LSTM freezes the weights of the first two layers and only fine-tunes the fully connected layers, while XGBoost retains the basic tree structure and adds relevant subtrees. The initial fine-tuning learning rate is set to 1 / 10 of the pre-trained rate and dynamically adjusted in conjunction with the ReduceLROnPlateau mechanism.
[0080] Example 3 like Figure 3 As shown, a sediment remediation system based on multimodal machine learning includes: a data acquisition module, a data processing module, a suggestion generation module, a scheme generation module, a status data acquisition module, and a scheme adjustment module; The data acquisition module is used to acquire multimodal data composed of satellite remote sensing data, in-situ sensor data and laboratory data of the target area; The data processing module is used to perform preprocessing and feature selection on the multimodal data in sequence to obtain a key feature set; The suggestion generation module is used to generate risk-cost trade-off suggestions based on the key feature set combined with historical pollution data and location information of ecologically sensitive areas. The solution generation module is used to generate an optimized repair solution based on the risk-cost trade-off suggestion and the multi-objective optimization algorithm; The status data acquisition module is used to perform a repair task in the target area based on the optimized repair scheme, and to acquire the status data of the target area during the execution process; The scheme adjustment module is used to adjust the optimization and repair scheme based on the status data.
[0081] Preferably, in this embodiment, the functional implementation of each functional module corresponds one-to-one with the content of the above method.
[0082] Example 4 Based on the same inventive concept, the present invention also provides a computer device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes a program stored in memory, it is able to implement a sediment remediation method based on multimodal machine learning, as in Embodiment 1 or 2.
[0083] The electronic device may include a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions in the memory to execute a sediment remediation method based on multimodal machine learning as described in Embodiment 1 or 2.
[0084] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a method and system for sediment remediation based on multimodal machine learning, which has the following beneficial effects: 1. By integrating a hybrid algorithm model of Random Forest and Long Short-Term Memory (LSTM) network, satellite remote sensing spectral data, in-situ sensor monitoring parameters (such as pH, redox potential, heavy metal concentration) and laboratory test results are dynamically analyzed to achieve high-precision real-time diagnosis of sediment pollution type, spatial distribution and migration trend.
[0086] 2. Combining the multi-objective optimization algorithm (NSGA-II), based on pollution characteristics and remediation cost constraints, the system autonomously matches the optimal remediation scheme (such as in-situ microbial induced mineralization remediation or ex-situ dredging technology), and drives drone swarms and underwater biomimetic robots to perform precise tasks (such as targeted drug delivery and pollution hotspot sampling) through end-to-end control protocols.
[0087] 3. Deep integration of machine learning algorithms with remediation engineering: The XGBoost algorithm is used to build a pollution risk prediction model, and LSTM time series analysis is used to predict the pollution diffusion pattern. At the same time, the decision path is dynamically optimized to form a closed-loop system of "perception-decision-execution-feedback". This improves efficiency by more than 30% compared with the traditional manual intervention mode. It is especially suitable for the precise treatment of bottom sediments in complex waters such as large lakes and estuaries, and has dynamic adaptability and potential for large-scale application.
[0088] 4. By combining multimodal data fusion with machine learning algorithms, dynamic pollution diagnosis, real-time optimization of remediation plans, and automated and precise execution can be achieved, significantly improving the efficiency and eco-friendliness of sediment treatment in complex waters.
[0089] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0090] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for sediment remediation based on multimodal machine learning, characterized in that, include: Multimodal data is obtained by acquiring satellite remote sensing data, in-situ sensor data, and laboratory data of the target area; Based on the multimodal data, preprocessing and feature selection are performed sequentially to obtain a key feature set; Based on the aforementioned key feature set, combined with historical pollution data and location information of ecologically sensitive areas, risk-cost trade-off recommendations are generated. An optimized repair scheme is generated based on the aforementioned risk-cost trade-off recommendation and multi-objective optimization algorithm; Based on the optimized repair scheme, a repair task is performed in the target area, and the status data of the target area is obtained during the execution process; The optimized repair scheme is adjusted based on the status data.
2. The method for sediment remediation based on multimodal machine learning according to claim 1, characterized in that, The preprocessing specifically includes: Based on the satellite remote sensing data, sensor noise and cloud interference are eliminated, and pollution-related spectral features are extracted to obtain preprocessed remote sensing data; Preprocessed sensor data is obtained by reducing sensor noise and filling in missing data based on the in-situ sensor data. Based on the laboratory data, discrete points are converted into a continuous spatial distribution to eliminate format differences and obtain preprocessed experimental data. Based on the preprocessed remote sensing data, the preprocessed sensor data, and the preprocessed experimental data, a feature alignment algorithm is used to eliminate the spatiotemporal resolution differences of the multi-source data, unify the temporal resolution, and obtain a preprocessed multimodal dataset.
3. The method for sediment remediation based on multimodal machine learning according to claim 2, characterized in that, The feature selection specifically includes: The preprocessed multimodal dataset is input into the random forest feature selection layer; The random forest feature selection layer evaluates the feature weights and ranks the importance of the data in the preprocessed multimodal dataset based on the Gini index, and selects the Top-30 data as key data to form the key feature set.
4. The method for sediment remediation based on multimodal machine learning according to claim 1, characterized in that, Generate risk-cost trade-off recommendations, specifically including: Based on the key feature set input into the spatiotemporal prediction layer of the Long Short-Term Memory network, the first layer of 128 LSTM units captures long-distance temporal dependence, the second layer of 64 units extracts high-level features, and outputs a pollution parameter prediction sequence for a future preset duration to obtain a pollution migration trend heatmap. The pollution migration trend heatmap, the historical pollution data, and the location information of the ecologically sensitive area are input into the XGBoost risk modeling layer to generate the risk-cost trade-off recommendation.
5. The method for sediment remediation based on multimodal machine learning according to claim 4, characterized in that, The specific risk-cost trade-off recommendations are as follows: Based on the aforementioned pollution migration trend heatmap, the pollutant concentration and redox potential are obtained; Risk assessment is conducted based on the pollutant concentration and the redox potential, and risk areas are divided into different risk levels. When the distance between the risk area and the ecologically sensitive area is less than a set value, the risk level is increased by a preset value, and a pollution risk heat map of the target area is generated. Based on the aforementioned risk level, corresponding remediation priority recommendations are provided; The risk-cost trade-off recommendation is generated based on the pollution risk heat map and the remediation priority recommendation.
6. The method for sediment remediation based on multimodal machine learning according to claim 4, characterized in that, Generate an optimized repair plan, specifically including: The state space is defined based on a deep Q-network, consisting of pollution feature vectors, environmental parameters, and equipment status; the action space is defined as a combination of remediation processes. Based on the aforementioned risk-cost trade-off, the corresponding combination of repair processes in the action space is selected in the state space; Based on the multi-objective optimization algorithm combined with the comprehensive objective function, the process parameters of the repair process combination are optimized. Pareto optimal solution set is generated by non-dominated sorting and crowding degree calculation. The optimized repair scheme is generated by matching specific repair process parameters through a decision knowledge base.
7. The method for sediment remediation based on multimodal machine learning according to claim 6, characterized in that, The comprehensive objective function F Specifically: ; ; ; ; ; ; ; in, ω 1. ω 2 and ω The numbers 3 represent weights, reflecting the importance of restoration efficiency, cost, and ecological disturbance, respectively. , , For the original sub-target, , , To correspond to the normalized index, This represents the objective function for maximizing repair efficiency. This represents the average total amount of pollution removed per unit time; a higher value indicates higher remediation efficiency. A i Indicates the first i Pollution removal volume in each treated area T Indicates the repair time. This represents the overall cost objective function. This represents the sum of total energy and material costs incurred during the repair process; a smaller value indicates lower costs. H Indicates equipment energy consumption. g Indicates electricity price, S Indicates the dosage of microbial agent. j Indicates the unit price of the inoculant. This represents the objective function for ecological disturbance. The overall quantitative assessment measures the degree of disturbance to aquatic ecosystems caused by restoration activities; a smaller value indicates higher eco-friendliness. α and β All represent weighting coefficients. L Indicates the mortality rate of benthic organisms. V This indicates the rate of change in water turbidity.
8. The method for sediment remediation based on multimodal machine learning according to claim 1, characterized in that, Perform the repair task, which specifically includes: Based on the optimized repair scheme, a structured instruction set containing spatial location, operation type, and parameter thresholds is generated; The structured instruction set is used to control the amount of pesticides dispensed by the drone swarm in the target area; The underwater robot collects and provides feedback on the status data of the target area in real time.
9. A method for sediment remediation based on multimodal machine learning according to claim 6, characterized in that, Adjusting the optimization and repair scheme based on the aforementioned status data specifically includes: The status data includes the redox potential value and pollutant concentration of the sediment in the target area; Threshold judgment is made based on the redox potential value and the pollutant concentration. When the change of the redox potential value exceeds the first threshold or the pollutant concentration rebounds to the second threshold within a preset time period, the optimized remediation scheme is updated and adjusted based on the state data.
10. A sediment remediation system based on multimodal machine learning, used to execute a sediment remediation method based on multimodal machine learning as described in any one of claims 1-9, characterized in that, include: The module includes a data acquisition module, a data processing module, a suggestion generation module, a solution generation module, a status data acquisition module, and a solution adjustment module. The data acquisition module is used to acquire multimodal data composed of satellite remote sensing data, in-situ sensor data and laboratory data of the target area; The data processing module is used to perform preprocessing and feature selection on the multimodal data in sequence to obtain a key feature set; The suggestion generation module is used to generate risk-cost trade-off suggestions based on the key feature set combined with historical pollution data and location information of ecologically sensitive areas. The solution generation module is used to generate an optimized repair solution based on the risk-cost trade-off suggestion and the multi-objective optimization algorithm; The status data acquisition module is used to perform a repair task in the target area based on the optimized repair scheme, and to acquire the status data of the target area during the execution process; The scheme adjustment module is used to adjust the optimization and repair scheme based on the status data.