An amphibious unmanned aerial vehicle river channel sediment concentration real-time monitoring method based on underwater light field imaging and cross-modal fusion
Patent Information
- Application Number
- CN202510336297.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-03-20
AI Technical Summary
然而,此类方法依赖光线穿透能力,在浑浊水体中有效探测深度受到极大限制,且无法区分悬浮泥沙与藻类、浮游生物的干扰信号
[0082]1)突破了传统监测方法的精度和效率限制,通过空中与水下数据的深度融合,结合先进的深度学习算法,能够实现高分辨率的泥沙浓度反演,具有实时性和高准确度,适应复杂水体环境;
Smart Images

Figure CN120177505B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of river sediment monitoring technology, and more specifically to a method for real-time monitoring of river sediment concentration using an amphibious unmanned aerial vehicle based on underwater light field imaging and cross-modal fusion. Background Technology
[0002] River sediment monitoring is a crucial foundation for water conservancy projects, ecological protection, and flood control, but traditional methods have long faced technical bottlenecks. Current mainstream methods include manual sampling and analysis and single-sensor surveying. Manual sampling relies on boats or manpower to collect water samples at specific locations, and its time cycle is long (usually weekly / monthly), making it difficult to reflect the dynamic process of sediment transport during flood peaks (e.g., the sediment transport rate in the lower reaches of the Yellow River can reach several tons per second during the flood season). While sensors such as sonar and optical turbidimeters can achieve continuous monitoring, they also have significant limitations:
[0003] Traditional sonar retrieves particle concentration by analyzing echo signals, but its accuracy is severely affected by sediment particle size distribution and water turbidity. Studies show that when sediment particle size is large (exceeding 0.1mm-2mm), the nonlinear response of the sonar scattering cross-section introduces significant concentration errors. Furthermore, the multiple scattering effect of suspended particles in highly turbid water leads to signal attenuation, potentially reducing the effective detection depth to less than 1 meter, making it unsuitable for deep-water monitoring.
[0004] Optical equipment such as multispectral cameras and hyperspectral imagers can infer sediment concentration by analyzing water reflectance. However, these methods rely on light penetration capabilities, which greatly limits their effective detection depth in turbid waters, and they cannot distinguish between suspended sediment and interference signals from algae and plankton. Even single forward scattering suppression schemes are insufficient to overcome complex multidirectional scattering noise.
[0005] Economic and reliability issues of multi-device collaboration: To enhance monitoring capabilities, existing technologies require combining drones (aerial remote sensing), underwater robots (sonar detection), and buoy stations (in-situ sensors). However, spatiotemporal registration errors between multiple platforms can lead to data fusion failures. Furthermore, the maintenance costs of multiple systems are high.
[0006] Therefore, proposing a real-time monitoring method for river sediment concentration using amphibious unmanned aerial vehicles based on underwater light field imaging and cross-modal fusion to address the difficulties of existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a method for real-time monitoring of river sediment concentration by amphibious unmanned aerial vehicles based on underwater light field imaging and cross-modal fusion. By collaborating with multi-dimensional data (aerial spectrum + underwater light field), it breaks through the limitations of accuracy and efficiency of traditional methods, realizes real-time and high-resolution inversion of sediment concentration in complex water bodies, and reduces the cost of multi-source equipment collaboration, providing an innovative tool for smart water conservancy.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A real-time monitoring method for river sediment concentration using an amphibious unmanned aerial vehicle (UAV) based on underwater light field imaging and cross-modal fusion includes the following steps:
[0010] S1. Acquire aerial multispectral data and preprocess the aerial multispectral data;
[0011] S2. The integrated polarization light field camera is mounted on an amphibious drone to capture underwater light field data. Through optical models and lightweight deep learning networks, underwater three-dimensional light field information is reconstructed and sediment scattering features are extracted.
[0012] S3. By fusing aerial multispectral data and sediment scattering characteristics through spatiotemporal alignment and attention mechanisms, a sediment concentration prediction model is generated, which then outputs a spatial distribution map of river sediment concentration.
[0013] S4. Based on the spatial distribution map of river sediment concentration, real-time monitoring and early warning of river sediment concentration are carried out through edge computing and cloud collaborative architecture.
[0014] Optionally, in the above method, the specific content of acquiring and preprocessing the aerial multispectral data in S1 is as follows:
[0015] S11, Multispectral data acquisition and radiometric correction;
[0016] S12, Geometric correction;
[0017] S13, Optical parameter inversion;
[0018] S14. Output and Formatting.
[0019] The above method, optionally, includes the following specific details regarding radiation correction in S11:
[0020] Radiometric calibration: An empirical linear method was used to establish the numerical relationship between DN and incident radiance L using the reflectivity values of a ground calibration plate. λ The linear relationship is expressed by the following formula:
[0021] L λ =a·DN λ +b
[0022] Among them, coefficients a and b are calibrated by the least squares method and independently fitted band by band;
[0023] Atmospheric correction: Simulation of atmospheric path radiation L using the MODTRAN model path Combined with solar irradiance E sun Inverting atmospheric transmittance T to retrieve water surface reflectance ρ λ The formula is as follows:
[0024]
[0025] Where cosθ is the angle between solar radiation and the ground normal.
[0026] The above method, optionally, includes the following specific details regarding geometric correction in S12:
[0027] Spatial registration: RTK coordinates are used to address the hovering and shaking of the UAV, and Kalman filtering is employed to fuse GPS data, as shown in the following formula:
[0028]
[0029] Where (X0, Y0, Z0) are the coordinates of the camera projection center, a i b i c i Here are the projection matrix parameters, f is the focal length, and (X, Y, Z) are the three-dimensional coordinates of a point in space to be projected onto the image, in meters.
[0030] An affine transformation matrix is constructed using the coordinates of the matching points to optimize the geometric positioning error of the image, as shown in the following formula:
[0031]
[0032] Where, m ij For rotation and scaling parameters, t x ,t y The translation parameters are solved using the least squares method.
[0033] Optionally, in the above method, the specific content of S2 for reconstructing underwater three-dimensional light field information and extracting sediment scattering features through an optical model and a lightweight deep learning network is as follows:
[0034] S21, Polarization field imaging hardware configuration;
[0035] S22. Underwater light field reconstruction and scattering suppression;
[0036] S23, Multi-dimensional light field feature extraction;
[0037] S24, Cross-modal feature association;
[0038] S25. Output and Formatting.
[0039] The above method, optionally, includes the following specific details regarding underwater light field reconstruction and scattering suppression in S22:
[0040] Light field depth estimation: A depth map is generated using the consistency of the light field focal length, and the scene depth z is solved through integral optimization.
[0041]
[0042] in, Let be the gradient field of the i-th viewpoint, TV(z) be the total variational regularization term to suppress noise interference, λ be the regularization weight coefficient, and N be the number of viewpoints or images.
[0043] Polarization scattering suppression: Based on Mueller matrix decomposition of water scattering components, recover the target reflected light intensity I. object :
[0044] S out =M water ·M object ·S in
[0045] Among them, M water M is the water scattering matrix. object S is the scattering matrix of the target object. in S out These are the Stokes vectors for the incident and output light, respectively, representing the polarization state of the beam;
[0046] Modeling the attenuation coefficient using Beer-Lambert's law:
[0047]
[0048] Where c is the water extinction coefficient, which is fed back in real time by a pre-set turbidity sensor, z is the depth, and β(z) is the polarization coupling coefficient related to the incident / observation angle.
[0049] The above method, optionally, includes the following specific content for dimensional light field feature extraction in S23:
[0050] Basic modeling of sediment scattering: Constructing scattering phase functions for different particle sizes based on Mie scattering theory:
[0051]
[0052] Among them, a n b n Let π be the Mie coefficient. n τ represents the radiation mode or radiation directionality during the scattering process. nThe function is related to the scattering intensity and particle size parameters of the particle, where d is the characteristic size of the scattering particle and n is the order index in the Mie scattering series expansion;
[0053] Feature encoding network: Design a lightweight dual-stream CNN to process the optical field flow and polarization components separately;
[0054] Optical flow: The input is a 9×9 sub-aperture view with a viewing baseline Δ = 0.5°. A 3D convolution kernel is used to extract parallax latent features. Polarization components: The input is a Stokes vector, S0, S1, S2. Anisotropic scattering features are extracted using angle-sensitive convolution. The two flow features are fused through cross-attention gating to output a multi-scale feature vector F∈R. 512 .
[0055] Optionally, in the above method, S3 uses spatiotemporal alignment and attention mechanisms to fuse aerial multispectral data and sediment scattering characteristics to generate a sediment concentration prediction model, thereby outputting the spatial distribution map of river sediment concentration.
[0056] S31, Multimodal Spatiotemporal Registration and Coding;
[0057] S32, Hierarchical cross-modal attention mechanism;
[0058] S33, Concentration inversion under physical constraints;
[0059] S34. Output and Formatting.
[0060] The above method, optionally, includes the following specific details regarding the hierarchical cross-modal attention mechanism in S32:
[0061] Pixel-level modulation: Utilizing the highest resolution features in the air to modulate coarse-grained features underwater, enhancing edge sensitivity, as shown in the following formula:
[0062]
[0063] Where σ is the Sigmoid function, ⊙ is the element-wise multiplication, and G is the modulation weight. l ∈[0,1] H×W , The feature map is extracted from the air branch at the l-th network layer. Conv is a feature map extracted from the underwater branch. 1×1 For a convolution operation with a kernel size of k×1, G l For modulation weights, The improved underwater features are obtained after pixel-level modulation;
[0064] Semantic Cross-Attention: Construct a multi-head attention mechanism, using underwater features as the query and aerial features as the key-value pair, to calculate cross-modal associations, as shown in the following formula:
[0065]
[0066] in, W q W k W v For the learnable weight matrix, d k The feature dimension or number of channels of the Key vector;
[0067] Residual fusion and upsampling: Attention outputs are stacked layer by layer, as shown in the following formula:
[0068]
[0069] Transposed convolution is used to progressively upsample to the original resolution;
[0070] The specific content of the concentration inversion of physical constraints in S33 is as follows:
[0071] Coupled sediment transport equation: The sediment transport continuity equation is introduced as a physical sensing loss and solved using the finite difference method, as shown in the following formula:
[0072]
[0073] Where c is the sediment concentration, D is the diffusion coefficient, and q s υ is the flux, μ is the sediment sink term, representing the rate of sediment settling, deposition, or removal, μ is the sediment source term, representing the intensity of new or replenished sediment, and x, y, z are the three-dimensional spatial coordinates in the river channel or water area.
[0074] Regression Network: Design a lightweight U-Net decoder, with the final fused features as input. Output sediment concentration field C∈R H×W The formula is as follows:
[0075]
[0076] The above method, optionally, includes the following specific content in S4 for real-time monitoring and early warning of river sediment concentration:
[0077] S41. Design the architecture of the early warning system;
[0078] S42. Establish a multi-level intelligent early warning mechanism;
[0079] S43, Early Warning Verification and Decision Support;
[0080] S44, System Outputs and Interfaces.
[0081] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a method for real-time monitoring of river sediment concentration based on underwater light field imaging and cross-modal fusion by amphibious unmanned aerial vehicles, the beneficial effects of which are:
[0082] 1) It breaks through the limitations of accuracy and efficiency of traditional monitoring methods. Through the deep fusion of aerial and underwater data and combined with advanced deep learning algorithms, it can achieve high-resolution sediment concentration inversion, with real-time performance and high accuracy, and is adaptable to complex aquatic environments.
[0083] 2) It can not only improve the accuracy and efficiency of river sediment monitoring, but also provide innovative technical means for fields such as smart water conservancy, flood control and environmental protection, and has broad application prospects and significant social benefits. Attached Figure Description
[0084] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0085] Figure 1 The flowchart illustrates a method for real-time monitoring of river sediment concentration using an amphibious unmanned aerial vehicle (UAV) based on underwater light field imaging and cross-modal fusion, as provided by this invention. Detailed Implementation
[0086] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0087] See Figure 1 As shown, this invention discloses a method for real-time monitoring of river sediment concentration using an amphibious unmanned aerial vehicle (UAV) based on underwater light field imaging and cross-modal fusion, comprising the following steps:
[0088] S1. Acquire aerial multispectral data and preprocess the aerial multispectral data;
[0089] S2. The integrated polarization light field camera is mounted on an amphibious drone to capture underwater light field data. Through optical models and lightweight deep learning networks, underwater three-dimensional light field information is reconstructed and sediment scattering features are extracted.
[0090] S3. By fusing aerial multispectral data and sediment scattering characteristics through spatiotemporal alignment and attention mechanisms, a sediment concentration prediction model is generated, which then outputs a spatial distribution map of river sediment concentration.
[0091] S4. Based on the spatial distribution map of river sediment concentration, real-time monitoring and early warning of river sediment concentration are carried out through edge computing and cloud collaborative architecture.
[0092] Furthermore, the specific content of acquiring and preprocessing aerial multispectral data in S1 is as follows:
[0093] S11, Multispectral data acquisition and radiometric correction;
[0094] S12, Geometric correction;
[0095] S13, Optical parameter inversion;
[0096] Specifically, the normalized turbidity index (NDI) quantifies the concentration of suspended sediment on the surface based on the ratio of the reflectance difference between the red band (650 nm) and the near-infrared band (850 nm). The formula is as follows:
[0097]
[0098] When NDTI > 0.4, it is identified as a high turbidity region (sediment concentration > 200 mg / L).
[0099] Chlorophyll-a concentration (Chl-a): A three-band algorithm was used to eliminate water absorption interference and suppress false positive signals from algae.
[0100]
[0101] The coefficients in the formula were obtained through laboratory calibration of actual river water bodies;
[0102] Denoising and outlier removal: Applying a guided filter to the reflectance image:
[0103]
[0104] Where I is the guide image (near-infrared band), ω k For local windows, minimize |qp| 2 (p is the input image) Preserve edge texture.
[0105] S14. Output and Formatting.
[0106] Specifically, the processed data is stored in a gridded GeoTIFF format, including: surface reflectance of each band (ρ1~ρ5); derived parameters (NDTI, Chl-a); spatially registered RGB true-color imagery; and metadata (timestamp, flight altitude, solar geometry parameters).
[0107] Specifically, aerial multispectral data is collected by a UAV at a constant flight altitude (50m) and speed (5m / s). The multispectral camera is set with five narrow bands (center wavelengths: 450nm, 550nm, 650nm, 750nm, 850nm, bandwidth ±10nm), simultaneously recording the solar zenith angle θ and the sensor's viewing angle φ. The RTK module provides centimeter-level positioning (horizontal ±1cm, vertical ±2cm).
[0108] Furthermore, the specific details of radiation correction in S11 are as follows:
[0109] Radiometric calibration: An empirical linear method was used to establish the relationship between the digital numerical value DN and the incident radiance L using the reflectivity values (20%, 40%, 60%) of the ground calibration plate. λ The linear relationship is expressed by the following formula:
[0110] L λ =a·DN λ +b
[0111] Among them, coefficients a and b are calibrated by the least squares method and independently fitted band by band;
[0112] Atmospheric correction: Simulation of atmospheric path radiation L using the MODTRAN model path Combined with solar irradiance E sun Inverting atmospheric transmittance T to retrieve water surface reflectance ρ λ The formula is as follows:
[0113]
[0114] Where cosθ is the angle between solar radiation and the ground normal.
[0115] Specifically, the sky light reflection component (the proportion of water surface ripples) needs to be deducted in the calculation.
[0116] Furthermore, the specific details of geometric correction in S12 are as follows:
[0117] Spatial registration: RTK coordinates are used to address the hovering and swaying motion of the UAV (amplitude < 3cm), and Kalman filtering is employed to fuse GPS data, as shown in the following formula:
[0118]
[0119] Where (X0, Y0, Z0) are the coordinates of the camera projection center, ai b i c i The projection matrix parameters are provided by the UAV's factory calibration parameters, with the default three-axis attitude being horizontal aircraft state. f is the focal length, and (X, Y, Z) are the three-dimensional coordinates of a point in space to be projected onto the image (such as a point in a river, a ground target, etc.), in meters, consistent with the RTK / GPS coordinate system used by the UAV.
[0120] An affine transformation matrix is constructed using the coordinates of the matching points to optimize the geometric positioning error of the image, as shown in the following formula:
[0121]
[0122] Where, m ij For rotation and scaling parameters, t x ,t y The translation parameters are solved using the least squares method.
[0123] Specifically, to eliminate residual distortion caused by flight trajectory deviation, 15cm×15cm high reflectivity targets (with known sub-meter GPS coordinates) were deployed along both banks of the monitored river section, spaced 50-100 meters apart. SIFT features were extracted from the near-infrared band (850nm) of the multispectral imagery and matched against a target template library, with matches having a confidence level >0.9 selected.
[0124] Furthermore, in S2, the specific content of reconstructing underwater three-dimensional light field information and extracting sediment scattering features through an optical model and a lightweight deep learning network is as follows:
[0125] S21, Polarization field imaging hardware configuration;
[0126] Specifically, the lens assembly uses a multi-aperture polarization decomposition sensor array, which achieves four-dimensional vector acquisition through an optical path folding device. Each sub-aperture corresponds to four polarization directions (0°, 45°, 90°, 135°).
[0127] Sensor: Global shutter CMOS (resolution 4096×2160), dynamic range 14-bit, supports multi-exposure fusion (short exposure 50μs to capture the main scattering peak, long exposure 5ms to extract deep details);
[0128] Light source compensation: Dual-sided 450nm blue LED array (peak power 80W), synchronous pulse triggering (delay <10μs), beam angle 20° to suppress lateral scattering.
[0129] S22. Underwater light field reconstruction and scattering suppression;
[0130] S23, Multi-dimensional light field feature extraction;
[0131] S24, Cross-modal feature association;
[0132] Specifically, air-water feature alignment: aligning the light field feature F... underwater With aerial multispectral data F air We construct a joint embedding space and maximize cross-modal correlations through canonical correlation analysis.
[0133] S25. Output and Formatting.
[0134] Specifically, the processed data is stored in point cloud matrix format, including: 3D light field point cloud (XYZ coordinates + Stokes vector + scattering phase); light field feature tensor (512-dimensional feature vector + attention weight map); scattering parameter matrix (particle size distribution, attenuation coefficient, optical depth); and cross-modal association weights (CCA projection matrix W). u W a ).
[0135] Furthermore, the specific details of underwater light field reconstruction and scattering suppression in S22 are as follows:
[0136] Light field depth estimation: A depth map is generated using the focal stack, and the scene depth z is solved through integral optimization.
[0137]
[0138] in, Let be the gradient field of the i-th viewpoint, TV(z) be the total variational regularization term to suppress noise interference, λ be the regularization weight coefficient, and N be the number of viewpoints or images.
[0139] Polarization scattering suppression: Based on Mueller matrix decomposition of water scattering components, recover the target reflected light intensity I. object :
[0140] S out =M water ·M object ·S in
[0141] Among them, M water M is the water scattering matrix. object S is the scattering matrix of the target object. in S out These are the Stokes vectors for the incident and output light, respectively, representing the polarization state of the beam;
[0142] Modeling the attenuation coefficient using Beer-Lambert's law:
[0143]
[0144] Where c is the water extinction coefficient, fed back in real time by a pre-set turbidity sensor, z is the depth, and β(z) is the polarization coupling coefficient related to the incident / observation angle. The target reflected light intensity I is extracted by inversely solving the Mueller equation. object .
[0145] Furthermore, the specific details of dimensional light field feature extraction in S23 are as follows:
[0146] Basic modeling of sediment scattering: Constructing scattering phase functions for different particle sizes (d = 0.05-2 mm) based on Mie scattering theory:
[0147]
[0148] Among them, a n b n Let π be the Mie coefficient. n τ represents the radiation mode or radiation directionality during the scattering process. n It is a function related to the scattering intensity and particle size parameters (such as particle diameter) of the particle, where d is the characteristic size of the scattering particle, and n is the order index in the Mie scattering series expansion, with a value of d > 1;
[0149] Specifically, π n Typically a function related to the scattering angle, it represents the radiation mode or radiation directionality during the scattering process, particularly relevant to the scattering characteristics of spherical particles; τ n This is a function related to the scattering intensity and particle size parameters (such as particle size), which helps characterize the scattering effect of electromagnetic waves of different wavelengths on particles. The measured polarized light field is matched with the theoretical phase function to generate a particle size probability distribution map.
[0150] Feature encoding network: Design a lightweight dual-stream CNN (LFA-Net) to process the optical field flow and polarization components separately;
[0151] Optical flow: The input is a 9×9 sub-aperture view with a viewing baseline Δ = 0.5°. A 3D convolutional kernel (3×3×3) is used to extract parallax latent features. Polarization components: The input is a Stokes vector, S0, S1, S2. Anisotropic scattering features are extracted using an oriented kernel. The two flow features are fused using a cross-attention gate to output a multi-scale feature vector F∈R. 512 .
[0152] Furthermore, S3 integrates aerial multispectral data and sediment scattering characteristics through spatiotemporal alignment and attention mechanisms to generate a sediment concentration prediction model, which then outputs the spatial distribution map of river sediment concentration.
[0153] S31, Multimodal Spatiotemporal Registration and Coding;
[0154] Specifically, spatiotemporal alignment: Utilizing the Precision Time Protocol (PTP) between the UAV and the underwater camera, timestamps are aligned to the microsecond level (deviation <1ms); based on GPS / INS integrated navigation data, a rigid body transformation matrix T is used... water →air maps underwater point clouds to a multispectral image coordinate system.
[0155] Feature encoder:
[0156] Aerial branch: Input is multispectral reflectance (ρ) 450 , ρ 550 , ρ 650 , ρ 750 , ρ 850 ) and NDTI, multi-scale features F extracted by ResNet-18 l air (Level l = 1, 2, 3, 4), with 64, 128, 256, and 512 output channels.
[0157] Underwater branch: The input is the light field feature tensor F underwater ∈R 512 Using the scattering parameter matrix, voxelized features F are generated using PointNet++. waterl ∈R H×W×C Aligned with the resolution of aerial features.
[0158] S32, Hierarchical cross-modal attention mechanism;
[0159] S33, Concentration inversion under physical constraints;
[0160] S34. Output and Formatting.
[0161] Specifically, the fusion results are stored in multi-layer HDF5 format, including: a three-dimensional sediment concentration field (XYZ grid + concentration values, resolution 0.1m×0.1m×0.05m); an attention weight distribution map (attention heatmaps at each level); a spatiotemporal alignment parameter set (rotation matrix, translation vector, projection error residual); and an inversion uncertainty map (calculated pixel-wise standard deviation σ based on Monte Carlo Dropout).
[0162] Furthermore, the specific details of the hierarchical cross-modal attention mechanism in S32 are as follows:
[0163] Pixel-level modulation: Utilizing the highest resolution features in the air to modulate coarse-grained features underwater, enhancing edge sensitivity, as shown in the following formula:
[0164]
[0165] Where σ is the Sigmoid function, ⊙ is the element-wise multiplication, and G is the modulation weight. l ∈[0,1] H×W , The feature map is extracted from the air branch at the l-th network layer. Conv is a feature map extracted from the underwater branch. 1×1 For a convolution operation with a kernel size of k×1, G l For modulation weights, The improved underwater features are obtained after pixel-level modulation;
[0166] Semantic Cross-Attention: Construct a multi-head attention mechanism (4 heads), using underwater features as the query and aerial features as the key-value pair, to calculate cross-modal associations, as shown in the following formula:
[0167]
[0168] in, W q W k W v For the learnable weight matrix, d k This is the feature dimension or number of channels of the Key vector, and can take a value of 64.
[0169] Residual fusion and upsampling: Attention outputs are stacked layer by layer, as shown in the following formula:
[0170]
[0171] Transposed convolution is used to progressively upsample to the original resolution (0.1m);
[0172] The specific content of the concentration inversion of physical constraints in S33 is as follows:
[0173] Coupled sediment transport equation: The sediment transport continuity equation is introduced as a physical sensing loss and solved using the finite difference method, as shown in the following formula:
[0174]
[0175] Where c is the sediment concentration, D is the diffusion coefficient, and q s υ is the flux, representing the rate of sediment deposition, siltation, or removal; μ is the sediment source, representing the intensity of new or replenished sediment; and x, y, z are the three-dimensional spatial coordinates in the river channel or water area.
[0176] Regression Network: Design a lightweight U-Net decoder, with the final fused features as input. Output sediment concentration field C∈R H×W(Unit: mg / L), the formula is as follows:
[0177]
[0178] Specifically, the maximum value is constrained to 1500 mg / L (corresponding to high sediment content scenarios such as the Yellow River).
[0179] Furthermore, the specific details of real-time monitoring and early warning of river sediment concentration in S4 are as follows:
[0180] S41. Design the architecture of the early warning system;
[0181] Specifically, the hardware deployment includes: Edge computing units: Deploying embedded devices equipped with high-performance GPUs (such as NVIDIA Jetson) to run lightweight AI models directly on drones and ground stations, completing data preprocessing (noise reduction, registration) and preliminary analysis, with a response time controlled within 50 milliseconds; Communication links: 5G dual-redundant network (main link: Sub-6GHz 100Mbps, backup link: satellite UHF band 10kbps), ensuring an online rate of >99.9%; Alarm triggering devices: LoRa IoT audible and visual alarms distributed along the coast (working radius 3km, trigger response time <0.5s).
[0182] Data processing flow: The three-dimensional concentration field data is transmitted after being compressed by Huffman coding (compression ratio 5:1); the edge nodes perform sliding window spatiotemporal filtering (window size 5×5×3, interval 0.1s); the parameters of the digital twin hydrodynamic model are updated synchronously in the cloud (time step Δt=10s).
[0183] S42. Establish a multi-level intelligent early warning mechanism;
[0184] Specifically, dynamic threshold setting:
[0185] Risk classification: Combining historical data and real-time hydrological and meteorological information, dynamic thresholds are divided into three levels: yellow, orange, and red (e.g., yellow alert is a concentration >250mg / L, and red alert is a concentration >1000mg / L with a continuously expanding range).
[0186] Adaptive adjustment: When the upstream rainfall suddenly increases or the water flow velocity fluctuates beyond the preset value, the warning level is automatically raised and the threshold is dynamically optimized using an exponential weighted algorithm.
[0187] Flood peak propagation forecast:
[0188] Based on the hydrodynamic model to simulate the movement path and arrival time of the sediment front, combined with real-time flow velocity data, the risk time window of key downstream areas (such as dams and bridges) is estimated with an error controlled within 15 minutes; if the prediction shows that important facilities will be impacted by high concentrations of sediment within 2 hours, an orange or higher warning will be immediately activated.
[0189] Automated response strategy:
[0190] Primary response: Notify water conservancy management personnel via SMS and APP push, and mark high-risk areas; Intermediate response: Trigger the riverside broadcasting system to play a voice alarm, and the drone automatically flies to the target area to take pictures and verify; Advanced response: Link the water conservancy gate control system to adjust the opening and closing degree according to the pre-programmed program to guide the diversion of sediment.
[0191] S43, Early Warning Verification and Decision Support;
[0192] Specifically, digital twin verification:
[0193] A digital twin model that is a complete mirror image of the physical environment is built in the cloud, and the actual monitoring data and simulation results are compared in real time. If the deviation exceeds 20%, the model parameters are automatically calibrated. Historical flood events are used for inversion testing to ensure the accuracy of the early warning (false alarm rate <5%).
[0194] Event Retrospective and Origin Tracing:
[0195] Based on blockchain technology, the entire chain of data (including original images, processing logs, and early warning instructions) is encrypted and stored, providing tamper-proof evidence traceability; it can retrieve 3D scene snapshots at any time point to analyze the sediment diffusion path and the effectiveness of prevention and control measures.
[0196] S44, System Outputs and Interfaces.
[0197] Specifically, real-time early warning signals are pushed to government emergency platforms and mobile terminals in standard JSON format, including risk level, scope of impact, and recommended measures;
[0198] Visualization terminal: Supports access from multiple platforms including PC-based WebGL 3D sandbox, AR head-mounted holographic interface, and lightweight mobile app;
[0199] Emergency control interface: Open API for calling automatic control systems of water conservancy facilities, such as gate opening adjustment command format in Modbus / TCP protocol.
[0200] Specifically, the technical solution of this invention can be applied to the following scenarios:
[0201] 1. Flood control and disaster relief emergency command
[0202] In flood control and disaster relief emergency command scenarios, the system needs to achieve dynamic tracking of flood peaks and prediction of dam breach risks. Dynamic flood peak tracking, in emergencies such as torrential rains and dam breaches, predicts in real time the diffusion path and arrival time of high-sediment-laden floods, providing a valuable time window for personnel evacuation and material dispatch. Dam breach risk prediction monitors sediment deposition around key dam structures and uses digital twin models to assess structural bearing capacity, preventing chain collapses caused by localized scouring. System applications include pushing 3D risk maps to the command center two hours before the flood front arrives, automatically identifying inundated areas (such as villages and roads); using AR sand tables to simulate the impact range of different flood discharge schemes (such as the water flow trajectory after opening specific gates), supporting rapid decision-making by commanders; and using drone swarms to execute tasks such as breach sealing and material airdrops. A typical case is ice jam control in the lower reaches of the Yellow River, where ice jam monitoring data is combined to predict the location of ice blockages, allowing for advance blasting and diversion to reduce the combined risks of ice impact and sediment deposition.
[0203] 2. Safety Operation and Maintenance of Water Conservancy Projects
[0204] In the context of safe operation and maintenance of water conservancy projects, the system needs to realize reservoir dredging scheduling and hydropower station equipment protection. Reservoir dredging scheduling optimizes dredging cycles and machinery deployment paths by accurately monitoring the sediment settling rate in the reservoir area, thereby reducing operation and maintenance costs. Hydropower station equipment protection extends the fault-free operating time of turbine units by predicting the risk of sediment blockage at the intake. System applications include generating monthly three-dimensional heat maps of reservoir sediment thickness and recommending dredging priorities based on historical data; overlaying real-time status of underwater gates and sediment accumulation warning lines on the AR inspection interface to prompt maintenance personnel to focus on worn parts; and automatically generating optimal path planning for dredging vehicles.
[0205] 3. River ecological protection monitoring
[0206] In the context of river ecological protection and monitoring, the system needs to protect endangered species habitats and monitor illegal sand mining. Endangered species habitat protection involves identifying areas of water quality deterioration caused by abnormal sediment concentrations, thus preventing large-scale deaths of benthic organisms. Illegal sand mining monitoring uses abrupt changes in riverbed topography to pinpoint suspected illegal mining sites, assisting in law enforcement and evidence collection.
[0207] 4. Urban flooding early warning linkage
[0208] In urban flooding early warning and linkage scenarios, the system needs to realize pipeline blockage early warning and traffic emergency management. Pipeline blockage early warning prevents drainage system paralysis by monitoring the amount of sediment deposited at water inlets during heavy rain; traffic emergency management prevents accidents caused by vehicles getting stuck in mud by predicting the sediment content of road flooded areas.
[0209] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0210] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for real-time monitoring of river sediment concentration using an amphibious unmanned aerial vehicle (UAV) based on underwater light field imaging and cross-modal fusion, characterized in that, Includes the following steps: S1. Acquire aerial multispectral data and preprocess the aerial multispectral data; S2. The integrated polarization light field camera is mounted on an amphibious drone to capture underwater light field data. Through optical models and lightweight deep learning networks, underwater three-dimensional light field information is reconstructed and sediment scattering features are extracted. S3. By fusing aerial multispectral data and sediment scattering characteristics through spatiotemporal alignment and attention mechanisms, a sediment concentration prediction model is generated, which then outputs a spatial distribution map of river sediment concentration. S4. Based on the spatial distribution map of river sediment concentration, real-time monitoring and early warning of river sediment concentration are carried out through edge computing and cloud collaborative architecture. The specific content of S2, which reconstructs underwater three-dimensional light field information and extracts sediment scattering features through an optical model and a lightweight deep learning network, is as follows: S21, Polarization field imaging hardware configuration; S22. Underwater light field reconstruction and scattering suppression; S23, Multi-dimensional light field feature extraction; S24, Cross-modal feature association; S25. Output and Formatting; The specific details of underwater light field reconstruction and scattering suppression in S22 are as follows: Light field depth estimation: A depth map is generated using the consistency of the light field focal length, and the scene depth z is solved through integral optimization. in, For the first i Gradient field of viewpoint The total variational regularization term suppresses noise interference. These are the regularization weight coefficients. Number of viewpoints or images; Polarization scattering suppression: Based on Mueller matrix decomposition of water scattering components, recover the intensity of target reflected light. I object : in, The water scattering matrix, The scattering matrix of the target object. These are the Stokes vectors for the incident and output light, respectively, representing the polarization state of the beam; Modeling the attenuation coefficient using Beer-Lambert's law: in, c The extinction coefficient of the water body is fed back in real time by a pre-installed turbidity sensor. For depth, The polarization coupling coefficient is related to the incident / observation angle. The angle between the direction of incident light propagation and the direction of observation is determined by the direction of illumination from the light source and the optical axis of the camera. It is used to describe the influence of the geometric relationship of light propagation on the polarization coupling effect during scattering, and thus affect the value of the coefficient β(z); The specific content of dimensional light field feature extraction in S23 is as follows: Basic modeling of sediment scattering: Constructing scattering phase functions for different particle sizes based on Mie scattering theory: in, , The Mie coefficient, This refers to the radiation mode or radiation directionality during the scattering process. It is a function related to the scattering intensity and particle size parameters of the particle. The characteristic size of the scattering particles, For the order index in the Mie scattering series expansion; Feature encoding network: Design a lightweight dual-stream CNN to process the optical field flow and polarization components separately; Optical flow: The input is a 9×9 sub-aperture view with a viewing baseline Δ=0.5°. A 3D convolution kernel is used to extract parallax latent features. Polarization component: The input is a Stokes vector, i.e. S 0 , S 1 , S 2 Anisotropic scattering features are extracted through angle-sensitive convolution; the two stream features are fused through cross-attention gating to output a multi-scale feature vector F∈R. 512 ; In S3, a sediment concentration prediction model is generated by fusing aerial multispectral data and sediment scattering characteristics through spatiotemporal alignment and attention mechanisms. The specific content of the output spatial distribution map of river sediment concentration is as follows: S31, Multimodal Spatiotemporal Registration and Coding; S32, Hierarchical cross-modal attention mechanism; S33, Concentration inversion under physical constraints; S34. Output and Format; The specific details of the hierarchical cross-modal attention mechanism in S32 are as follows: Pixel-level modulation: Utilizing the highest resolution features in the air to modulate coarse-grained features underwater, enhancing edge sensitivity, as shown in the following formula: in, For the Sigmoid function, Modulate weights for element-wise multiplication. , The feature map is extracted from the air branch at the l-th network layer. The feature map extracted from the underwater branch, For convolution operations with a kernel size of k×1, For modulation weights, The improved underwater features are obtained after pixel-level modulation; Semantic Cross-Attention: Construct a multi-head attention mechanism, using underwater features as the query and aerial features as the key-value pair, to calculate cross-modal associations, as shown in the following formula: in, , , , For learnable weight matrix, The feature dimension or number of channels of the Key vector; Residual fusion and upsampling: Attention outputs are stacked layer by layer, as shown in the following formula: Transposed convolution is used to progressively upsample to the original resolution; The specific content of the concentration inversion of physical constraints in S33 is as follows: Coupled sediment transport equation: The sediment transport continuity equation is introduced as a physical sensing loss and solved using the finite difference method, as shown in the following formula: in, c The concentration of sediment. D Where is the diffusion coefficient. q s For flux, For sediment accumulation, it indicates the rate at which sediment settles, accumulates, or is removed. For sediment source terms, it represents the intensity of newly added or replenished sediment. , , Three-dimensional spatial coordinates within a river or body of water; Regression Network: Design a lightweight U-Net decoder, with the final fused features as input. Output sediment concentration field ∈ × The formula is as follows: 。 2. The method for real-time monitoring of river sediment concentration based on underwater light field imaging and cross-modal fusion by an amphibious unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The specific steps in S1 for acquiring and preprocessing aerial multispectral data are as follows: S11, Multispectral data acquisition and radiometric correction; S12, Geometric correction; S13, Optical parameter inversion; S14. Output and Formatting.
3. The method for real-time monitoring of river sediment concentration based on underwater light field imaging and cross-modal fusion by an amphibious unmanned aerial vehicle (UAV) according to claim 2, characterized in that, The specific details of radiation correction in S11 are as follows: Radiometric calibration: An empirical linear method is used to establish a digital numerical value DN and incident radiance based on the reflectivity values of a ground calibration plate. L λ The linear relationship is expressed by the following formula: Among them, coefficient , Calibration was performed using the least squares method, and independent fitting was performed band by band. Atmospheric correction: Simulation of atmospheric path radiation using the MODTRAN model L path Combined with solar irradiance E sun With atmospheric transmittance T Inverting water surface reflectivity ρ λ The formula is as follows: Among them, cos θ It is the angle between solar radiation and the ground normal.
4. The method for real-time monitoring of river sediment concentration based on underwater light field imaging and cross-modal fusion by an amphibious unmanned aerial vehicle (UAV) according to claim 2, characterized in that, The specific details of geometric correction in S12 are as follows: Spatial registration: RTK coordinates are used to address the hovering and shaking of the UAV, and Kalman filtering is employed to fuse GPS data, as shown in the following formula: in,( , , () represents the coordinates of the camera projection center. a i , b i , c i These are the projection matrix parameters. f For focal length, ( , , () represents the three-dimensional coordinates of a point in space to be projected onto the image, in meters; An affine transformation matrix is constructed using the coordinates of the matching points to optimize the geometric positioning error of the image, as shown in the following formula: in, For rotation and scaling parameters, , The translation parameters are solved using the least squares method.
5. The method for real-time monitoring of river sediment concentration based on underwater light field imaging and cross-modal fusion by an amphibious unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The specific content of S4 regarding real-time monitoring and early warning of river sediment concentration is as follows: S41. Design the architecture of the early warning system; S42. Establish a multi-level intelligent early warning mechanism; S43, Early Warning Verification and Decision Support; S44, System Outputs and Interfaces.
Citation Information
Patent Citations
High concentration underwater polarization imaging method
CN109187364A
Method for judging validity of high-resolution multispectral water depth inversion data based on spectral roughness information
CN114459438A