An airborne laser radar sounding data processing method and system

By employing a joint denoising strategy combining machine learning and signal-to-noise ratio optimization, along with Gaussian mixture models and frequency-domain low-pass filtering, the problems of noise suppression and underwater echo signal preservation in airborne lidar depth sounding data were solved. This enabled high-precision extraction of water depth information and inversion of water optical parameters, filling the measurement gap in extremely shallow waters and generating complete integrated land-sea point cloud data.

CN121114969BActive Publication Date: 2026-01-27GEOPHYSICAL SURVEY TEAM OF SHANDONG COALFIELD GEOLOGY BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511682194.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-01-27
Estimated Expiration
2045-11-17

AI Technical Summary

Technical Problem

In existing airborne lidar depth sounding technology, the raw ALB waveform data is mixed with complex background noise and random noise. Existing denoising algorithms have difficulty accurately modeling the statistical characteristics of background noise, which leads to over-smoothing or weakening of weak but critical bottom echo signals while suppressing noise, especially in deep or turbid water, resulting in the loss or error of water depth information.

Method used

A joint denoising strategy based on machine learning-based background noise modeling and frequency domain low-pass filtering based on signal-to-noise ratio optimization is adopted. The probability distribution of background noise is fitted by Gaussian mixture model, and the optimal cutoff frequency is determined by iterative optimization for low-pass filtering. The infrared laser waveform amplitude and blue-green laser multi-channel waveform features are combined for high-precision classification and decomposition. The bounded Gaussian decomposition model is used to separate the water surface, water body and bottom reflection echo components.

Benefits of technology

It effectively avoids the problem of water depth information loss caused by excessive smoothing in traditional methods, provides input data with high signal-to-noise ratio and high fidelity, improves the accuracy of water depth information extraction and the reliability of water body optical parameter inversion, enhances the ability to distinguish complex land features, especially the ability to identify extremely shallow water areas, and generates more complete and accurate integrated land and sea point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121114969B_ABST
    Figure CN121114969B_ABST
Patent Text Reader

Abstract

The application discloses a kind of airborne laser radar sounding data processing method and system, it is related to marine surveying and mapping and earth observation technical field, including the following steps: collecting the original full waveform data of airborne laser radar sounding system, the original full waveform data includes infrared laser channel data and blue-green laser multichannel data;The original full waveform data is carried out joint denoising processing, obtains the waveform data after denoising, wherein the joint denoising processing includes background noise modeling and removal based on machine learning and random noise filtering based on frequency domain low-pass filtering;Through the fusion based on machine learning background noise modeling and based on signal-to-noise ratio optimization frequency domain low-pass filtering, the cooperative joint denoising strategy is formed.The method can not only accurately estimate and remove complex background noise, but also adaptively filter out random noise, to the greatest extent retain weak water bottom echo signal, effectively avoid the water depth information loss problem caused by excessive smoothing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine surveying and Earth observation technology, and more specifically, to a method and system for processing airborne lidar depth sounding data. Background Technology

[0002] Airborne lidar depth sounding (ALB) is a novel active remote sensing technology that rapidly acquires large-scale, high-precision underwater topographic information by emitting lasers (usually infrared and blue-green lasers) from an airborne platform and receiving echo signals from the sea surface, water column, and seabed. Compared with traditional shipborne acoustic detection, ALB technology has significant advantages such as high efficiency, wide coverage, and non-contact measurement, and has irreplaceable application value in coastal zone resource surveys, marine engineering construction, channel dredging, island and reef mapping, and marine ecological environment protection.

[0003] The core data of the ALB system is the full waveform data recording the change of echo signal intensity over time. This waveform data contains reflection information from various targets, including the water surface, suspended particles within the water body, and the bottom. By precisely processing the waveform data, laser propagation time can be extracted to calculate water depth, three-dimensional point clouds can be generated to construct digital elevation models, and optical parameters of the water body (such as suspended sediment concentration and chlorophyll concentration) can be retrieved. Therefore, the quality of waveform data processing directly determines the final application effectiveness of the ALB system.

[0004] However, existing ALB waveform data processing methods still have significant technical bottlenecks: the raw ALB waveform is mixed with complex background noise (such as solar background radiation) and random noise. Existing denoising algorithms, such as fixed-threshold waveform filtering or empirical mode decomposition, often struggle to accurately model the statistical characteristics of background noise. This leads to over-smoothing or attenuation of weak but crucial underwater echo signals while suppressing noise, especially in deep or turbid water, resulting in the loss or error of water depth information. Summary of the Invention

[0005] To overcome the aforementioned shortcomings of existing technologies, this invention provides an airborne lidar depth sounding data processing method and system. This method integrates machine learning-based background noise modeling with frequency-domain low-pass filtering based on signal-to-noise ratio optimization, forming a synergistic joint denoising strategy. This effectively avoids the loss of depth information caused by excessive smoothing in traditional methods, providing high signal-to-noise ratio and high-fidelity input data for subsequent waveform classification and decomposition.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method and system for processing airborne lidar depth sounding data includes the following steps: acquiring raw full-waveform data from an airborne lidar depth sounding system, the raw full-waveform data including infrared laser channel data and blue-green laser multi-channel data; performing joint denoising processing on the raw full-waveform data to obtain denoised waveform data, wherein the joint denoising processing includes background noise modeling and removal based on machine learning and random noise filtering based on frequency domain low-pass filtering; performing high-precision classification of land and water waveforms based on the denoised waveform data to obtain land and water classification results, wherein the high-precision classification includes coarse classification based on infrared laser waveform amplitude and misclassification correction by fusing pulse spatial location information and blue-green laser multi-channel waveform features; decomposing the classified blue-green laser waveforms based on a bounded constrained Gaussian decomposition model to separate water surface reflection, water backscattering, and bottom reflection echo components; and calculating the laser propagation time in air and water based on the echo components, and generating seabed topographic point clouds and inverting water optical parameters.

[0008] In a preferred embodiment, the joint denoising process on the original full waveform data specifically includes: segmenting the original full waveform data to separate non-echo signal segments and echo signal segments; modeling the non-echo signal segments using a Gaussian mixture model to fit the probability distribution of background noise; subtracting the background noise component from the original full waveform data based on the fitted background noise probability distribution to remove background noise; performing a Fourier transform on the waveform data after background noise removal to convert it to the frequency domain; in the frequency domain, determining the optimal cutoff frequency for low-pass filtering with the goal of maximizing the signal-to-noise ratio of the underwater echo signal through iterative optimization; and performing an inverse Fourier transform on the filtered frequency domain signal to obtain the time-series waveform signal after removing random noise.

[0009] In a preferred embodiment, the high-precision classification of water and land waveforms based on the denoised waveform data specifically involves: extracting the waveform amplitude features of the infrared laser channel from the denoised waveform data; performing initial clustering of the waveform amplitude features using the K-means clustering algorithm to initially classify the waveforms into water and land categories, obtaining a coarse classification result; extracting the three-dimensional spatial location information of each laser pulse and extracting the waveform features of the deep-water and shallow-water channels of the blue-green laser, the waveform features including waveform skewness, kurtosis, and pulse width; based on the three-dimensional spatial location information and the waveform features, using the DBSCAN clustering algorithm to identify and correct misclassified points in the coarse classification result, obtaining an accurate water and land classification result; for extremely shallow water areas, combining the waveform features of the deep-water and shallow-water channels of the blue-green laser, constructing a multi-channel feature vector and using a support vector machine for classification and discrimination, thereby achieving accurate identification of extremely shallow water areas.

[0010] In a preferred embodiment, the bounded-constraint Gaussian decomposition model decomposes the classified blue-green laser waveforms, specifically by: constructing an airborne laser depth sounding waveform function model, wherein the water surface reflection echo and bottom reflection echo components are modeled as single Gaussian functions, and the water backscattering echo components are modeled as a superposition of multiple Gaussian functions; based on the laser refraction law at the air-water interface and the physical range of the water attenuation coefficient, bounded constraints on the parameters of each echo component are derived and set, the parameters including amplitude, pulse width, and position; under the bounded constraints, a nonlinear least squares optimization problem is established with the goal of minimizing the waveform fitting residual; the Levenberg-Marquardt algorithm is used to solve the optimization problem, iteratively estimating the optimal parameters of each Gaussian component to complete the waveform decomposition.

[0011] In a preferred embodiment, after the classified blue-green laser waveform is decomposed using the bounded-constraint Gaussian decomposition model, the method further includes: calculating the laser propagation time in the water body based on the time difference between the surface reflection echo component and the bottom reflection echo component obtained from the decomposition; performing point cloud repositioning calculations using positioning and attitude determination data provided by the airborne POS system to generate a high-precision seabed topographic point cloud; and inverting the suspended sediment concentration or chlorophyll a concentration optical parameters of the water body based on the amplitude and waveform characteristics of the backscattered echo component of the water body, combined with the water body attenuation model.

[0012] An airborne lidar depth sounding data processing system includes: a data acquisition module configured to acquire raw full-waveform data from the airborne lidar depth sounding system; a joint denoising module configured to jointly remove background noise and random noise from the raw full-waveform data; a land-water classification module configured to perform high-precision classification of land-water waveforms based on the denoised waveform data; a waveform decomposition module configured to decompose blue-green laser waveforms based on a bounded Gaussian decomposition model; and an information extraction module configured to calculate laser propagation time, generate seabed topographic point clouds, and invert water optical parameters based on the decomposition results.

[0013] In a preferred embodiment, the joint denoising module includes: a background noise modeling unit configured to model and remove background noise from non-echo signal segments using a Gaussian mixture model; and a random noise filtering unit configured to filter out random noise by iteratively optimizing a frequency domain low-pass filter with a cutoff frequency.

[0014] In a preferred embodiment, the land-water classification module includes: a coarse classification unit configured to perform preliminary land-water classification based on infrared laser waveform amplitude features and K-means clustering; and a fine classification unit configured to fuse pulse spatial location information and blue-green laser multi-channel waveform features, and use the DBSCAN clustering algorithm and support vector machine to correct misclassification and identify extremely shallow water areas.

[0015] The technical effects and advantages of the airborne lidar depth sounding data processing method and system of the present invention are as follows:

[0016] By integrating machine learning-based background noise modeling with frequency-domain low-pass filtering based on signal-to-noise ratio optimization, a synergistic joint denoising strategy is formed. This method can not only accurately estimate and remove complex background noise, but also adaptively filter out random noise, preserving weak underwater echo signals to the greatest extent. It effectively avoids the problem of water depth information loss caused by excessive smoothing in traditional methods, providing high signal-to-noise ratio and high-fidelity input data for subsequent waveform classification and decomposition.

[0017] By employing a two-level classification architecture combining coarse and fine classification, and innovatively integrating waveform amplitude, pulse spatial location information, and multi-channel waveform features of blue-green lasers, the system significantly enhances its ability to distinguish complex features such as intertidal zones and broken wave zones. In particular, its specialized identification mechanism for extremely shallow waters effectively solves the industry problem of these areas being frequently misidentified as land due to weak signal characteristics, filling a gap in nearshore shallow water depth measurement and resulting in more complete and accurate integrated land-sea point cloud data.

[0018] By introducing bounded constraints based on the physical principles of laser bathymetry, waveform decomposition is transformed from a purely mathematical fitting problem into a physics-driven parameter optimization problem. This ensures that the parameters of the decomposed surface, water body, and seabed echo components are all within physically feasible ranges, fundamentally avoiding decomposition results that are "mathematically correct but physically absurd." This significantly improves the accuracy and reliability of laser propagation time calculation, water depth extraction, and water body optical parameter inversion results.

[0019] This approach organically integrates previously relatively independent processing steps such as denoising, classification, and decomposition into a coherent and automated data processing pipeline. This reduces reliance on manual intervention and experience-based parameter tuning, significantly improving the efficiency and standardization of airborne laser bathymetry (ALB) data processing. It provides strong technical support for the rapid generation of standardized marine mapping products and powerfully promotes the application of ALB technology in large-scale marine mapping projects.

[0020] It can provide more accurate and reliable seabed topography and water environment data for applications such as coastal zone resource surveys, port and waterway maintenance, marine ecological environment protection, and safeguarding maritime rights. By improving the overall performance of the ALB system, the cost and risks of traditional ship-based surveys can be reduced, providing key technical support for the sustainable development of the marine economy, and has broad application prospects and significant socio-economic value. Attached Figure Description

[0021] Figure 1This is a block diagram of a method and system for processing airborne lidar depth sounding data.

[0022] Figure 2 This is a schematic diagram of the joint noise reduction module of an airborne lidar depth sounding data processing method and system.

[0023] Figure 3 This is a schematic diagram of the land and water classification module of an airborne lidar depth sounding data processing method and system.

[0024] Figure 4 This is a schematic diagram of the waveform decomposition module of an airborne lidar depth sounding data processing method and system.

[0025] Figure 5 This is a flowchart illustrating the information extraction module of an airborne lidar depth measurement data processing method and system. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0027] Example 1, Figures 1-5 This invention discloses an airborne lidar depth sounding data processing method and system, comprising the following steps:

[0028] The raw full waveform data of the airborne lidar depth sounding system is collected. The raw full waveform data includes infrared laser channel data and blue-green laser multi-channel data.

[0029] In this embodiment, joint denoising processing is performed on the original full waveform data, specifically as follows: the original full waveform data is segmented into non-echo signal segments and echo signal segments; a Gaussian mixture model is used to model the non-echo signal segments and fit the probability distribution of background noise; based on the fitted background noise probability distribution, the background noise component is subtracted from the original full waveform data to complete the background noise removal; the waveform data after background noise removal is subjected to Fourier transform and converted to the frequency domain; in the frequency domain, with the goal of maximizing the signal-to-noise ratio of the underwater echo signal, the optimal cutoff frequency is determined through iterative optimization for low-pass filtering; the filtered frequency domain signal is subjected to inverse Fourier transform to obtain the time-series waveform signal after removing random noise.

[0030] It should be noted that: Background noise mainly originates from ambient light (such as sunlight and cloud reflection), noise from the electronic circuits of the laser sensor (such as thermal noise), and atmospheric scattering interference. It manifests as persistent low-frequency / stationary noise. Traditional methods often use a single Gaussian model or a fixed threshold for modeling, which cannot adapt to the noise distribution in complex environments (such as the significant difference in noise distribution between near-shore strong light environments and offshore weak light environments), resulting in low modeling accuracy. This can lead to either missing noise or mistakenly deleting weak underwater echoes. Random noise mainly originates from small fluctuations in the laser pulse, interference from airflow on laser propagation, and sensor sampling errors. It manifests as high-frequency clutter superimposed on the useful signal. Traditional methods often use low-pass filtering with a fixed cutoff frequency. If the cutoff frequency is too high, noise filtering is incomplete; if the cutoff frequency is too low, it smooths out the details of the useful signal (especially the peak characteristics of underwater echoes), making it impossible to accurately extract water depth information.

[0031] Furthermore, based on the laser emission-reception timing of the ALB system, the original waveform (horizontal axis is time, vertical axis is signal strength) is divided into two intervals: Non-echo signal segment: the laser has not contacted any target (such as the zero signal period before emission, the latitude period before the laser propagates to the water surface / bottom, and the tail segment after all echoes are received). This interval contains only background noise and no useful echoes (no echoes from the water surface, water body, or bottom). Echo signal segment: the signal interval generated after the laser contacts the target (from the appearance of the peak echo on the water surface to the end of the peak echo on the bottom). This interval contains background noise + random noise + useful echo signal.

[0032] The clean samples of background noise only exist in the non-echo signal segment. If the entire waveform is modeled directly, the useful signals in the echo signal segment will be misclassified as noise, leading to modeling errors. Therefore, segmentation is a prerequisite for achieving accurate background noise modeling.

[0033] Traditional methods do not perform segmentation and directly take the mean noise or standard deviation threshold of the entire waveform, which causes useful signal components to be mixed into the background noise model, resulting in low modeling accuracy.

[0034] The noise data of the non-echo signal segment is input into the Gaussian Mixture Model (GMM), and the model parameters (such as the weight, mean, and variance of each sub-Gaussian distribution) are estimated by the EM (Expectation Maximization) algorithm, and finally the multi-peak probability distribution model of the background noise is obtained.

[0035] In real-world ALB scenarios, background noise is not a single Gaussian distribution (as commonly used in traditional methods), but rather a mixed distribution formed by the superposition of multiple noise sources (such as electronic noise, ambient light noise, and atmospheric scattering noise). For example, electronic noise has a low-mean, low-variance Gaussian distribution, while ambient light noise has a high-mean, high-variance Gaussian distribution. GMM can accurately fit this complex mixed distribution through a linear combination of multiple sub-Gaussian distributions, thus more realistically describing the statistical characteristics of background noise.

[0036] Compared to traditional single Gaussian models, GMM can: adapt to background noise in different environments (such as day / night, nearshore / offshore); capture local features of noise (such as noise peaks caused by a sudden increase in ambient light at a certain time); and provide statistical basis for subsequent accurate denoising (rather than simply applying a threshold).

[0037] Based on the obtained GMM probability distribution, the probability that a point belongs to background noise is calculated for each time point in the original full waveform data (including non-echo and echo bands), and the background noise intensity at that point is estimated based on the probability distribution. Finally, the estimated background noise intensity is subtracted from the original signal intensity to obtain the waveform data after removing background noise.

[0038] Background noise is a stationary component superimposed on the useful signal. The Gaussian Mixture Model (GMM) model yields the statistical expectation of the noise. Subtracting this expectation component removes most of the stationary background interference. For example, if the original signal strength at a certain time point is 100, and the GMM-estimated background noise strength at that point is 30, then the signal strength after removing the background noise is 70 (if it's an echo signal segment, 70 includes the useful signal plus random noise; if it's a non-echo signal segment, 70 should be close to 0, verifying the modeling effectiveness).

[0039] Traditional threshold denoising directly classifies signals below the threshold as noise and sets them to zero. If the underwater echo signal is weak (such as in very shallow water or turbid water), it may be mistakenly identified as background noise and deleted. However, probabilistic denoising based on GMM only subtracts the noise component. Even weak underwater echoes (such as signal strength 40 and noise 30) can retain 10 of the useful signal, achieving high fidelity.

[0040] The time-domain waveform data (horizontal axis: time, vertical axis: signal strength) is converted into a frequency-domain signal (horizontal axis: frequency, vertical axis: signal amplitude / power) using Discrete Fourier Transform (DFT) or Fast Fourier Transform (FFT).

[0041] The time domain and frequency domain of a signal are two representations of the same signal: Useful signals: water surface echoes, water backscattering, and bottom echoes are all low-frequency or specific frequency band signals (because the propagation time of the laser pulse is relatively stable, and the rate of change of the echo signal is slow); Random noise: mostly high-frequency signals (such as electronic thermal noise, signal fluctuations caused by airflow interference, and the rate of change is fast).

[0042] Therefore, after converting the signal to the frequency domain, the useful signal and random noise will be separated on the frequency axis (the useful signal is concentrated at the low frequency end, and the random noise is concentrated at the high frequency end). At this time, high-frequency random noise can be specifically filtered out by low-pass filtering without affecting the low-frequency useful signal. In the time domain, random noise and useful signal are superimposed and difficult to separate; in the frequency domain, the frequency band difference between the two is amplified, which makes subsequent precise filtering possible and avoids the problem of smoothing the useful signal by traditional time-domain filtering.

[0043] The core steps of random noise filtering consist of three sub-steps: Initial cutoff frequency setting: Based on the laser wavelength of the ALB system (e.g., 532nm for blue-green lasers) and the sampling frequency, set an initial low-pass filter cutoff frequency (e.g., 10kHz, to ensure that most of the low-frequency useful signals are initially retained); Filtering and signal-to-noise ratio calculation: Perform low-pass filtering using the initial cutoff frequency (retaining signals below this frequency and filtering out signals above this frequency), then temporarily convert the filtered frequency domain signal back to the time domain and calculate the signal-to-noise ratio (SNR) of the underwater echo signal. SNR = peak intensity of underwater echo / intensity of residual random noise after filtering; Iterative optimization of cutoff frequency: Gradually adjust the cutoff frequency (e.g., increase or decrease by 0.5kHz each time), repeating the process of filtering → converting to the time domain → calculating SNR until the cutoff frequency with the largest SNR is found. This cutoff frequency is the optimal cutoff frequency, and the final low-pass filtering is performed using this frequency.

[0044] The essence of finding the optimal cutoff frequency is to find a balance between 'filtering out more random noise' and 'retaining more useful signals': if the cutoff frequency is too high, high-frequency random noise will not be filtered out completely, resulting in a low SNR; if the cutoff frequency is too low, some low-frequency useful signals (such as the detailed features of underwater echoes) will be filtered out, the peak intensity of underwater echoes will decrease, and the SNR will also decrease; therefore, only the cutoff frequency that maximizes the SNR can simultaneously satisfy the requirements of sufficient random noise filtering and complete retention of useful signals.

[0045] Traditional methods use a fixed cutoff frequency (e.g., 8kHz based on experience), which cannot adapt to different scenarios (e.g., weak underwater echoes in deep water require a slightly higher cutoff frequency to preserve the signal; strong underwater echoes in shallow water require a slightly lower cutoff frequency to filter out more noise). Iterative optimization, on the other hand, enables scenario adaptation, ensuring optimal underwater echo SNR for each measurement area.

[0046] The frequency domain signal filtered at the optimal cutoff frequency is converted back to time domain waveform data through inverse discrete Fourier transform (IDFT) or inverse fast Fourier transform (IFFT), which is the final waveform data after joint denoising.

[0047] The frequency domain is the processing domain. The final waveform classification and decomposition steps need to be based on the characteristics of the time domain waveform (such as echo peak position, pulse width, and skewness). Therefore, the time domain signal must be recovered through inverse transformation.

[0048] At this point, the background noise in the time-domain waveform data has been accurately removed through GMM modeling, and random noise has been adaptively filtered out through optimal cutoff frequency filtering, preserving complete useful signals (water surface echo, water backscattered echo, and bottom echo). Moreover, the SNR of the bottom echo has reached its maximum, laying a high-quality data foundation for subsequent water-land classification (requiring clear echo characteristics) and waveform decomposition (requiring accurate echo peak value and pulse width).

[0049] The original full waveform data is subjected to joint denoising processing to obtain denoised waveform data. The joint denoising processing includes background noise modeling and removal based on machine learning and random noise filtering based on frequency domain low-pass filtering.

[0050] In this embodiment, high-precision classification of water and land waveforms is performed based on the denoised waveform data. Specifically, the waveform amplitude features of the infrared laser channel in the denoised waveform data are extracted; the waveform amplitude features are initially clustered based on the K-means clustering algorithm to classify the waveforms into water and land categories, obtaining a coarse classification result; the three-dimensional spatial location information of each laser pulse is extracted, and the waveform features of the deep-water and shallow-water channels of the blue-green laser are extracted, including waveform skewness, kurtosis, and pulse width; based on the three-dimensional spatial location information and waveform features, the DBSCAN clustering algorithm is used to identify and correct misclassified points in the coarse classification result to obtain an accurate water and land classification result; for extremely shallow water areas, the waveform features of the deep-water and shallow-water channels of the blue-green laser are combined, and a multi-channel feature vector is constructed and a support vector machine is used for classification and discrimination to achieve accurate identification of extremely shallow water areas.

[0051] It should be noted that in the processed high-fidelity denoised waveform, the waveform amplitude features of the infrared laser channel (distinct from the blue-green laser channel) are extracted separately, specifically including: peak amplitude (the maximum reflection intensity of the infrared laser on the target surface); mean amplitude (the average reflection intensity of the entire infrared waveform); and peak amplitude ratio (the ratio of peak amplitude to mean amplitude).

[0052] The difference in the medium reflection characteristics between infrared lasers and blue-green lasers is key to selecting the infrared channel for coarse classification. Infrared lasers (wavelengths typically >1000nm) cannot penetrate water, only exhibiting strong reflection at the water surface (high peak infrared amplitude of water waveforms), while there is no reflection at the bottom (no secondary peak in water infrared waveforms). Land (such as beaches and rocks) provides stable reflection of infrared lasers, but the reflection intensity is usually lower than that of water (low peak infrared amplitude of land waveforms, and lacking the double-peak characteristic of water surface-bottom). This difference in infrared amplitude between land and water is highly discriminative, allowing for the rapid screening of most typical land and water waveforms.

[0053] Using the infrared channel as the basis for coarse classification is more efficient than using the blue-green channel. The water and land characteristics of the infrared waveform are more intuitive. No complex calculations are required, and the initial screening can be completed quickly, reducing the amount of calculation required for subsequent fine classification (equivalent to filtering out obvious water and land waveforms and only performing fine processing on the blurry waveforms).

[0054] The extracted infrared laser amplitude features (such as peak amplitude) are used as clustering samples. K=2 is set (target categories: water and land). Unsupervised clustering is performed using the K-means algorithm: two cluster centers are randomly initialized (initial water amplitude center and land amplitude center); the distance between the amplitude feature of each laser pulse and the two centers is calculated, and the pulse is assigned to the category with the closer distance; the mean amplitude feature of each category is recalculated, and the cluster centers are updated; the distance calculation-classification-center update is repeated until the cluster centers no longer change, and the preliminary water and land classification results are finally output (most typical waveforms have been correctly classified, and only a few misclassified points of blurred waveforms remain).

[0055] K-means is a highly efficient and easy-to-implement algorithm for unsupervised clustering, suitable for processing massive waveform data of ALB (a single flight can generate millions of laser pulses). Its core logic is that samples of the same class have high feature similarity, while samples of different classes have large feature differences. The infrared amplitude characteristics of high values ​​for water and low values ​​for land perfectly match the clustering assumptions of K-means.

[0056] Fixed amplitude threshold classification (e.g., setting amplitude > X as water area and < X as land area) requires manual parameter tuning and cannot adapt to different regions (e.g., the amplitude thresholds of nearshore high-reflection water surfaces and offshore low-reflection water surfaces differ greatly); while K-means determines cluster centers through data self-learning, without manual intervention, and has stronger adaptability.

[0057] This step is only a rough classification and may still result in misclassification (e.g., very shallow water has low infrared amplitude and may be classified as land; low-reflectivity land has high infrared amplitude and may be classified as water), which will need to be corrected in subsequent steps.

[0058] To supplement the coarse classification results, two key features are added to provide a basis for fine classification correction: 1) Three-dimensional spatial location information of laser pulses: The three-dimensional coordinates (X, Y, Z) of each laser pulse are obtained from the POS (Positioning and Attitude System) of the ALB system, reflecting the pulse's position in real geographic space and demonstrating the spatial correlation between adjacent pulses (e.g., if the X and Y coordinates of a pulse are surrounded by water pulses, it is more likely to be a water area). 2) Multi-channel waveform features of blue-green lasers: Morphological features of deep-water channels (for water depths > 3 meters, with strong waveform signals and clearly separated peaks) and shallow-water channels (for water depths 1-3 meters, with weak waveform signals and relatively close peaks) are extracted. Specifically, these include: skewness (the degree of waveform asymmetry: water waveforms have a double peak value (surface-bottom double peak), resulting in skewness closer to 0; land waveforms are mostly single peaks, resulting in a large absolute value of skewness); kurtosis (the sharpness of the waveform: water waveforms have low kurtosis due to the double peak value, while land waveforms have high kurtosis due to the single peak value); and pulse width (waveform duration: water waveforms have a longer width because they include surface, water body, and bottom echoes; land waveforms have a shorter width due to only a single echo).

[0059] Coarse classification uses only the single-dimensional feature of infrared amplitude, while fine classification requires the fusion of multi-dimensional features. Spatial location information supplements macroscopic correlation, and blue-green multi-channel morphological features supplement microscopic waveform differences. The combination of the two can effectively distinguish fuzzy samples in coarse classification (e.g., a pulse infrared amplitude is low (coarse classification is land), but the spatial location is in a contiguous area of ​​water, and the blue-green waveform skewness is close to 0 (water feature), then it can be corrected to water).

[0060] This step is a crucial transition from 'single feature classification' to 'multi-feature collaborative classification', providing a dual basis for spatial and morphological judgment for subsequent correction of misclassification points, and solving the shortcomings of traditional methods that lack spatial correlation and have single features.

[0061] Each laser pulse is treated as a multi-feature sample (feature vector = 3D coordinate X + 3D coordinate Y + 3D coordinate Z + skewness of blue-green deep channel + kurtosis of blue-green shallow channel + pulse width of blue-green multi-channel), and input into DBSCAN (density-based spatial clustering algorithm) to correct misclassified points: set the core parameters of DBSCAN (ε: neighborhood radius, MinPts: minimum number of samples in the neighborhood). ε is set according to the laser point cloud density of the ALB system (e.g., if the point cloud density is 10 points / m², then ε is set to 0.3 meters to ensure that adjacent pulses can be included in the same neighborhood). It iterates through all samples, identifying core samples (samples with a neighborhood size ≥ MinPts, representing typical categories of continuous distribution), boundary samples (samples with a neighborhood size < MinPts, but adjacent to core samples, classified as core samples), and noise samples (neither core samples nor adjacent to core samples, i.e., misclassified points in coarse classification). For noise samples, it corrects them according to the category of their nearest core sample (e.g., if the core samples adjacent to a noise sample are all water bodies, it corrects it from land to water). Finally, it outputs accurate land-water classification results without isolated misclassified points.

[0062] DBSCAN's core advantage lies in its density-based identification of continuous regions, which can automatically filter out noise points. It perfectly matches the spatial continuity characteristics of land and water in ALB, where water and land are distributed in contiguous areas (high density) in geographic space, while misclassified points exist in isolation (low density), which corresponds exactly to DBSCAN's noise samples.

[0063] Collaboration logic with K-means: K-means is responsible for fast coarse division to reduce the amount of data, while DBSCAN is responsible for fine correction to improve accuracy. The combination of the two avoids the problem of K-means misclassifying isolated points and solves the defects of DBSCAN in large amount of computation and low efficiency when directly processing massive data, thus achieving a balance between efficiency and accuracy.

[0064] For extremely shallow water areas, a dedicated supplementary approach is used, employing Support Vector Machines (SVM) to accurately identify areas using multi-channel features of blue-green lasers. For potentially missed extremely shallow water areas (depth < 1 meter), a separate classification strategy is designed: Sample screening: Suspected extremely shallow water samples are screened from the denoised waveforms. Blue-green laser deep-water channels exhibit no obvious double peaks (overlapping surface and bottom echoes), shallow-water channels have narrow waveform widths (short propagation distance), and low infrared amplitudes (close to land). Multi-channel feature fusion: A dedicated feature vector is constructed for suspected samples, including the amplitude ratio, peak position difference, and waveform integral value between the blue-green deep-water and shallow-water channels (the amplitude ratio and peak position difference in extremely shallow water areas are significantly smaller than in shallow water areas, and the integral value is even lower). SVM classification training and discrimination: An SVM classifier is trained using labeled real extremely shallow water samples (data verified through GNSS-RTK field measurements) to establish an extremely shallow water-land classification model. Suspected samples are input into the model, and the output is either extremely shallow water or land, supplementing the accurate classification results.

[0065] The waveform characteristics of extremely shallow waters are nonlinear and the sample size is small. Traditional clustering algorithms (such as K-means and DBSCAN) have weak ability to distinguish nonlinear features. However, SVM can map low-dimensional nonlinear features to a high-dimensional linear space through kernel function mapping (such as RBF kernel) and find the optimal classification hyperplane, so as to achieve high-precision classification even with a small sample size.

[0066] This step fills the gap in traditional methods for identifying extremely shallow water areas, which are key areas for coastal ecological protection and intertidal resource surveys (such as mangrove wetlands and mudflats). Accurate identification can generate complete land-sea integrated topographic data, avoiding the blind spot of the last meter in nearshore mapping.

[0067] Based on the denoised waveform data, high-precision classification of land and water waveforms is performed to obtain land and water classification results. The high-precision classification includes coarse classification based on infrared laser waveform amplitude and misclassification correction by fusing pulse spatial location information and blue-green laser multi-channel waveform features.

[0068] The classified blue-green laser waveform is decomposed using a bounded-constraint Gaussian decomposition model. Specifically, an airborne laser bathymetry waveform function model is constructed, where the surface reflection echo and bottom reflection echo components are modeled as single Gaussian functions, and the water backscattered echo components are modeled as a superposition of multiple Gaussian functions. Based on the laser refraction law at the air-water interface and the physical range of the water attenuation coefficient, bounded constraints on the parameters of each echo component are derived and set, including amplitude, pulse width, and position. Under the bounded constraints, a nonlinear least-squares optimization problem is established with the goal of minimizing the waveform fitting residual. The Levenberg-Marquardt algorithm is used to solve the optimization problem, iteratively estimating the optimal parameters of each Gaussian component to complete the waveform decomposition.

[0069] It should be noted that, based on the interaction law between the laser and the medium in the ALB system, the full waveform (measured waveform) of the blue-green laser in the water area is modeled as a superposition of three types of echo components, expressed as:

[0070]

[0071] In the formula, W(t): the measured full waveform of blue-green laser (time-domain signal, t is time), and Ws(t).

[0072] The water surface reflection echo component is modeled as a single Gaussian function: (As is the amplitude, ts is the echo peak position, σs is the pulse width correlation parameter), Ww(t): the backscattered echo component of the water body, modeled as the superposition of multiple Gaussian functions: (n is the number of components, which depends on the turbidity of the water; higher turbidity results in more scattering points, and n is larger.) Wb(t): The underwater reflected echo component, modeled as a single Gaussian function. ε: Fitting residual (ideally should approach 0).

[0073] The Gaussian function modeling of the three types of echo components originates from the physical characteristics of laser pulses. Laser pulses themselves are approximately Gaussian distributed, and after interacting with different media (water surface, water body, and bottom), the time-domain form of the echo signal still maintains Gaussian properties (or superimposed Gaussian properties): the water surface is a smooth reflective surface, the laser is reflected once, the echo is concentrated and there is no scattering, so it is a single Gaussian function; the water body is a scattering medium, the laser is scattered multiple times by water molecules and suspended particles in the water, the echo is dispersed and has a long duration, so it needs to be described by superimposed Gaussian functions; the bottom is a rough reflective surface, the laser is reflected once (although there is weak scattering, the dominant component is reflection), so it is a single Gaussian function.

[0074] This step transforms the complex measured waveform into a decomposable mathematical model, clarifying which components to be decomposed and the mathematical form of each component. This provides a target framework for subsequent decomposition and avoids the problem of arbitrarily setting the number of components in traditional decomposition (such as decomposing too few water scattering components or too many spurious components).

[0075] Based on the physical laws governing laser propagation in air-water-underwater (such as the law of refraction, the law of water attenuation, and ALB system hardware parameters), upper and lower boundaries are set for the parameters of each Gaussian function in step 1 (A: amplitude, t: peak position, σ: pulse width related parameters), forming insurmountable physical constraints. The key constraints are shown in the table below:

[0076]

[0077] The essence of the constraint conditions is to transform physical laws into mathematical inequalities. For example, ts < tb directly stems from the time sequence of the laser propagation path, As being the maximum results from the difference in reflection / scattering rates between the water surface and the water body, and the attenuation of Awi is due to the energy loss caused by the absorption and scattering of the laser by the water body. These constraints are not set subjectively but are objective boundaries that can be derived from the parameters of the ALB system and physical formulas.

[0078] This step fundamentally eliminates physically absurd decomposition results (such as the bottom echo being earlier than the water surface, and the scattering amplitude being greater than the reflection) by setting constraint conditions, enabling each decomposed parameter to have a clear physical meaning and providing a reliable data source for the inversion of water body parameters in the subsequent calculation of laser propagation time.

[0079] Transforming the waveform decomposition into a constrained non - linear least - squares optimization problem, the mathematical expression is:

[0080] .

[0081] s.t. All parameter boundary constraints set in step 2, where: tk (k = 1, 2,..., m): the sampling time points of the measured waveform (m is the number of sampling points); objective function: the sum of the squared residuals between the measured waveform and the decomposed and synthesized waveform is minimized (to ensure the mathematical fitting accuracy); constraint conditions: the parameter boundaries in step 2 (to ensure physical rationality).

[0082] The Gaussian decomposition model of the ALB waveform is a non - linear model (the Gaussian function is a non - linear function, and it becomes more non - linear after the superposition of multiple components). Therefore, the optimization problem belongs to a constrained non - linear least - squares problem. Traditional unconstrained optimization only needs to minimize the residuals, while this step needs to minimize the residuals under the premise of satisfying physical constraints, which is equivalent to finding the optimal fitting solution within the physically feasible domain.

[0083] Select the LM algorithm as the optimization solver, and obtain the optimal parameters of each echo component through iterative calculation, finally completing the waveform decomposition:

[0084] Initializing parameters: Based on the intuitive characteristics of the measured waveform (such as setting the first peak of the waveform as ts, the last peak as tb, and the flat section in the middle as the Ww range), set the initial values of each parameter (which need to satisfy the constraint conditions to avoid entering the infeasible domain in the initial stage of iteration);

[0085] Iterative calculation: Calculate the objective function value (sum of squared residuals) and gradient (partial derivatives of the objective function with respect to each parameter) under the current parameters; combine the damping factor of the LM algorithm (to balance the fast convergence of the Gauss-Newton method and the stability of the steepest descent method), update the parameters (ensuring the new parameters still satisfy the constraints; if they exceed the boundary, adjust them to the boundary values); calculate the updated objective function value; if the residual decrease is less than a set threshold (e.g., 10⁻⁻⁶), then... 6 If the iteration converges, the calculation stops; otherwise, the update is repeated.

[0086] Output results: After iterative convergence, the optimal parameters (As,ts,σs; Awi,twi,σwi; Ab,tb,σb) of each echo component are obtained, and the decomposed surface, water body, and bottom echo waveforms are generated based on these parameters, thus completing the waveform decomposition.

[0087] The LM algorithm is a classic algorithm for solving nonlinear least squares problems. Compared with other algorithms (such as gradient descent and Gauss-Newton method), it has two major advantages: fast convergence: combined with the quadratic convergence characteristic of Gauss-Newton method, the iteration speed is fast when approaching the optimal solution; strong stability: by adjusting the damping factor, it behaves as the steepest descent method in the early stage of iteration (when the parameters are far from the optimal solution), avoiding divergence, and is especially suitable for multi-parameter, strongly nonlinear ALB waveform decomposition.

[0088] This step ensures that constrained optimization problems can be solved efficiently and stably, avoiding the problems of iterative divergence (such as computational crashes caused by parameters exceeding physical boundaries) or slow convergence (low efficiency when processing massive waveforms) of traditional algorithms, and meeting the large-scale, high-precision decomposition requirements of ALB data.

[0089] The classified blue-green laser waveform is decomposed based on a bounded constrained Gaussian decomposition model to separate the water surface reflection, water backscattering and bottom reflection echo components.

[0090] After decomposing the classified blue-green laser waveforms using a bounded-constraint Gaussian decomposition model, the process also includes: calculating the laser propagation time in the water body based on the time difference between the surface reflection echo component and the bottom reflection echo component obtained from the decomposition; performing point cloud repositioning calculations using positioning and attitude determination data provided by the airborne POS system to generate a high-precision seabed topographic point cloud; and inverting the optical parameters of suspended sediment concentration or chlorophyll a concentration in the water body based on the amplitude and waveform characteristics of the backscattered echo component and the water body attenuation model.

[0091] It should be noted that: the peak time of the water surface echo (ts) and the peak time of the bottom echo (tb) are extracted from the decomposition results, and the time difference between the two is calculated (Δt=tb−ts). This time difference is the round-trip propagation time of the laser from the water surface to the bottom and back to the water surface. Then, the round-trip time is divided by 2 to obtain the one-way propagation time of the laser from the water surface to the bottom (twater=Δt / 2).

[0092] Laser light travels slower in water than in air (approximately 300,000 km / s in air and approximately 225,000 km / s in water due to water's refractive index of approximately 1.33), and the propagation time is directly proportional to the water depth (water depth = propagation speed in water × one-way propagation time). Therefore, accurately calculating twater is a prerequisite for subsequent water depth calculations.

[0093] Compared to traditional methods that directly find the peak difference from the original waveform (which is susceptible to noise and waveform overlap), this step uses the echo position after physical constraint decomposition to calculate the time difference. This avoids time difference deviations caused by misjudging noise peaks or water body scattering signals as surface / bottom echoes (for example, in turbid water, traditional methods may treat strong scattering signals as bottom echoes, resulting in a longer calculated propagation time, while this method distinguishes between scattering and bottom echoes through constraints during decomposition, making the time difference more accurate).

[0094] Combined with the data from the airborne POS system, point cloud repositioning calculation is performed to generate a high-precision seabed topographic point cloud. Core operations: (1) Water depth calculation: Multiply the one-way propagation time (twater) obtained in step 1 by the propagation speed of the laser in the water to obtain the relative water depth (i.e., the vertical distance from the water surface to the bottom); (2) Three-dimensional coordinate transformation: Call the three-dimensional coordinates (Xp, Yp, Zp) of the aircraft at the time of laser emission provided by the airborne POS system (positioning and attitude determination system) and the flight attitude (pitch angle, roll angle, heading angle), and combine the laser emission angle to calculate the coordinates of the laser incident point (Xs, Ys, Zs) on the water surface; then, according to the relative water depth and the water surface attitude, convert the bottom position from the vertical distance relative to the water surface to the absolute three-dimensional coordinates (Xb, Yb, Zb); (3) Point cloud generation: Integrate the three-dimensional coordinates of all bottom points, remove abnormal points (such as water depth abnormal values ​​caused by decomposition errors), and form a seabed topographic point cloud.

[0095] The high accuracy of seabed topographic point clouds relies on two core elements: accurate water depth (derived from the propagation time calculation in step 1) and accurate spatial position (derived from the combination of positioning and attitude data from the POS system and laser geometry). For example, changes in the attitude of an aircraft during flight (such as turbulence) can cause a shift in the laser incident angle. If this is not corrected, it will cause a lateral deviation in the coordinates of the seabed point. The attitude data from the POS system can compensate for this shift in real time.

[0096] Traditional point cloud repositioning often ignores the influence of water surface attitude on water depth conversion (assuming the water surface is horizontal). However, this step corrects the water depth projection error caused by tilted water surface by decomposing the obtained water surface echo features (such as waveform width reflecting water surface roughness). For example, in wave regions, tilted water surface will make the actual laser propagation path longer than the vertical distance. This error can be eliminated by attitude correction, thus improving the planar accuracy of the point cloud by 10%-20%.

[0097] Based on the characteristics of backscattering components of water bodies, the core operations for inverting water body optical parameters (suspended sediment concentration or chlorophyll a concentration) are as follows: (1) Extracting scattering characteristics: From the decomposed backscattering echo components of water bodies, calculate two key characteristics: one is the total scattering intensity (the sum of the amplitudes of all scattering components, reflecting the total scattering ability of water bodies to laser); the other is the scattering attenuation rate (the rate at which the amplitude of the scattering components decreases with increasing propagation time, reflecting the rate at which laser energy is lost through absorption and scattering by water bodies); (2) Model inversion: Input the above scattering characteristics into the water body optical parameter inversion model (this model is trained through a large amount of measured data (such as the analysis results of synchronously collected water samples) to establish the mathematical relationship between scattering characteristics and suspended sediment concentration / chlorophyll a concentration), and output the inversion results. For example, the higher the total scattering intensity, the more suspended particles there are in the water body, and the higher the suspended sediment concentration; the faster the scattering attenuation, the stronger the absorption of laser by the water body, and the higher the chlorophyll a concentration may be (chlorophyll absorbs lasers of a specific wavelength strongly).

[0098] Suspended sediment and chlorophyll a in water significantly affect the scattering and absorption characteristics of laser light. Suspended sediment particles (1-100 micrometers in diameter) primarily cause Mie scattering, with the scattering intensity increasing with particle concentration. Chlorophyll a (a pigment found in algae) mainly absorbs blue-green light, leading to rapid attenuation of laser energy. Therefore, the intensity and attenuation pattern of the scattered echo can directly reflect the concentration of these substances.

[0099] Traditional inversion methods directly calculate the total energy of the original waveform, which is easily affected by water surface and bottom reflections (for example, strong water surface reflections can mask the scattering signal). In contrast, this step uses the pure water body scattering component that has been decomposed and separated to perform the inversion, eliminating the signal pollution from water surface / bottom reflections. This reduces the suspended sediment concentration inversion error from the traditional 15%-20% to less than 8%, and improves the chlorophyll a concentration inversion accuracy by more than 30%.

[0100] Example 2, Figure 1 This invention presents an airborne lidar depth sounding data processing system, comprising a data acquisition module, a joint denoising module, a land-water classification module, a waveform decomposition module, and an information extraction module.

[0101] The data acquisition module is configured to acquire raw full waveform data from the airborne lidar depth sounding system.

[0102] The joint denoising module is configured to jointly remove background noise and random noise from the original full waveform data.

[0103] The water and land classification module is configured to perform high-precision classification of water and land waveforms based on denoised waveform data.

[0104] The waveform decomposition module is configured to decompose the blue-green laser waveform based on a bounded-constraint Gaussian decomposition model.

[0105] The information extraction module is configured to calculate laser propagation time, generate seabed topographic point clouds, and invert water optical parameters based on the decomposition results.

[0106] The joint denoising module includes: a background noise modeling unit, configured to model and remove background noise from non-echo signal segments using a Gaussian mixture model; and a random noise filtering unit, configured to filter out random noise through frequency domain low-pass filtering with iteratively optimized cutoff frequency.

[0107] It should be noted that: the system receives raw full waveform data (infrared + blue-green laser channels) from the data acquisition module and automatically identifies the non-echo signal segments (the laser emission / reception timing information must be synchronized with the data acquisition module to ensure accurate segmentation).

[0108] Signal segment selection: Pure background noise samples are extracted from the original data by using a preset laser space distance time threshold (set according to the ranging range of the ALB system; for example, if there is no echo within 100ns after laser emission, it is determined to be a non-echo segment); GMM modeling engine: The built-in Gaussian mixture model algorithm module automatically estimates the number of optimal sub-Gaussian distributions (judged by the BIC information criterion, avoiding manual setting), and iteratively fits the probability distribution of noise through the EM algorithm (outputting the weight, mean, and variance of each sub-distribution); Noise stripping execution: Based on the fitted distribution model, the expected value of the noise component is calculated point by point for the original full waveform data (including echo and non-echo segments), and this value is directly subtracted from the original signal to output the intermediate data after removing background noise.

[0109] This unit's configuration emphasizes automation and adaptability, requiring no manual intervention in threshold settings. It can automatically adjust GMM parameters based on noise distribution in different survey areas (e.g., day / night, nearshore / offshore), solving the problem of poor generalization ability of traditional fixed models. For example, in strong light nearshore environments, it automatically increases the number of sub-Gaussian distributions to fit complex noise; in weak light offshore environments, it simplifies the model to improve efficiency.

[0110] This unit is configured to filter out random noise by frequency domain low-pass filtering through iterative optimization of the cutoff frequency. Its function is to transform the random noise filtering steps into an executable system module. The specific implementation logic is as follows: Data input: Receive the background noise-removed data output by the background noise modeling unit (at this time, the data still contains high-frequency random noise).

[0111] Frequency Domain Conversion Module: Built-in Fast Fourier Transform (FFT) algorithm converts time-domain waveform data into frequency-domain signals (output frequency-amplitude spectrum); Cutoff Frequency Optimization Engine: Presets an initial cutoff frequency range (based on the ALB system sampling frequency, e.g., 0-20kHz); it iterates through candidate frequencies within the range in steps (e.g., 0.5kHz), performing low-pass filtering on each frequency (preserving low-frequency signals); the filtered frequency-domain signal is converted back to the time domain via Inverse Fourier Transform (IFFT), and the underwater echo signal-to-noise ratio (SNR) is calculated (SNR = underwater peak intensity / filtered residual noise intensity); the candidate frequency with the highest SNR is automatically locked as the optimal cutoff frequency; Filtering Execution and Output: The final low-pass filtering is performed using the optimal cutoff frequency, and then converted back to the time domain via IFFT, outputting the jointly denoised clean waveform data (passed to the subsequent land-water classification module).

[0112] The core of this unit's configuration is a dynamic optimization mechanism. Unlike traditional filter modules with fixed cutoff frequencies, its built-in optimization engine can adapt to different signal characteristics in real time (e.g., in deep water areas where the bottom echo is weak, it automatically selects a higher cutoff frequency to preserve the signal; in shallow water areas where the echo is strong, it selects a lower frequency to filter out more noise), ensuring a balance between noise removal and signal preservation.

[0113] The high efficiency of the joint denoising module stems from the timing connection and data closure between the two units, forming a relay processing of stationary noise first, followed by random noise: Pre-processing dependency: The background noise modeling unit must be executed first. If the background noise is not removed, its low-frequency stationary components will be mixed into the frequency domain analysis, causing the random noise filtering unit to mistakenly retain some background noise as useful low-frequency signals, reducing the denoising effect; Data quality transfer: The background-removed data output by the background noise modeling unit provides a high signal-to-noise ratio input for the random noise filtering unit, enabling it to more accurately identify high-frequency random noise (otherwise, strong background noise in the original data will mask the characteristics of random noise, leading to deviations in cutoff frequency optimization); Parameter linkage and adaptation: The core parameters of the two units (such as the number of sub-distributions of the GMM and the cutoff frequency optimization step size) can be dynamically adjusted according to data quality feedback. For example, if there is still obvious residual noise in the data after background noise modeling (judged by a preset threshold), the random noise filtering unit will automatically reduce the cutoff frequency optimization step size (e.g., from 0.5kHz to 0.2kHz) to improve optimization accuracy.

[0114] The structural refinement of the joint denoising module essentially achieves engineering and standardization of the denoising process through modular design. Its value lies in: modular reuse: the two units can be upgraded independently (such as replacing with a more efficient GMM optimization algorithm or introducing wavelet transform-assisted filtering) without modifying other modules of the system, reducing maintenance costs; automation efficiency: the built-in algorithm engines of the units (GMM modeling, cutoff frequency optimization) do not require manual intervention, adapting to the processing needs of massive data (up to TB per flight) of the ALB system, with processing efficiency more than 10 times higher than traditional manual parameter tuning methods; data quality assurance: through the collaborative processing of the two units, the signal-to-noise ratio (SNR) of the output waveform data is improved by an average of 20-30%, especially with significant preservation of weak underwater echoes in extremely shallow water areas and turbid water bodies, laying a high-fidelity data foundation for the accurate classification of the subsequent land-water classification module and the extraction of physical parameters by the waveform decomposition module.

[0115] The water and land classification module includes: a coarse classification unit, configured to perform preliminary water and land classification based on infrared laser waveform amplitude features and K-means clustering; and a fine classification unit, configured to fuse pulse spatial location information and blue-green laser multi-channel waveform features, and use the DBSCAN clustering algorithm and support vector machine to correct misclassification and identify extremely shallow water areas.

[0116] It should be noted that: the system receives high-fidelity denoised waveform data (including infrared laser channel and blue-green laser multi-channel) from the joint denoising module, as well as the three-dimensional spatial coordinates (X, Y, Z) of laser pulses from the airborne POS system.

[0117] Infrared channel feature extraction: Automatically extract the peak amplitude, mean amplitude, and peak percentage of the infrared laser waveform (corresponding to step 1 of claim 3), and output the infrared feature vector (the core basis for coarse classification); Blue-green multi-channel feature extraction: For the deep and shallow channels of blue-green lasers, calculate the skewness, kurtosis, and pulse width (key indicators reflecting waveform morphology), and output the blue-green morphological feature vector (the microscopic basis for fine classification); Spatial feature integration: Convert the three-dimensional coordinates of the POS system into relative spatial position features (such as distance from surrounding pulses and slope), and output the spatial correlation feature vector (used to correct the classification results using spatial continuity).

[0118] The feature extraction engine of this unit has adaptive computing capabilities, dynamically adjusting feature dimensions based on waveform quality (e.g., when the waveform is blurred in extremely shallow water, it automatically adds supplementary features such as the peak position difference integral value of the blue-green channel), avoiding the failure of a single feature in complex scenarios. For example, in the broken wave region, the infrared amplitude feature is unstable, and the skewness feature weight of the blue-green waveform is automatically strengthened.

[0119] The coarse classification unit performs the efficient initial screening function, quickly distinguishing most typical water and land waveforms based on infrared features, reducing the computational load of subsequent fine classification. This corresponds to the engineering implementation of step 2 in claim 3: Data input: Receive the infrared feature vector output by the feature extraction unit.

[0120] K-means clustering engine: Built-in optimized K-means algorithm, automatically sets initial cluster centers (based on historical data of water and land infrared feature distribution), no manual intervention required; Adaptive clustering parameters: Dynamically adjusts the number of iterations according to the amount of input data (parallel computing is enabled to accelerate when the amount of data exceeds 1 million pulses), and evaluates the clustering quality through the silhouette coefficient (ensuring that the discrimination of typical water and land waveforms is >0.8); Coarse classification results output: Divides waveforms into preliminary water class, preliminary land class, and fuzzy class (waveforms with close cluster center distances that are difficult to classify clearly, usually accounting for <10%), among which the fuzzy class will enter the subsequent fine classification stage.

[0121] Compared to traditional fixed threshold classification, this unit's self-learning capability can adapt to different regions (such as high-reflectivity lake surfaces and low-reflectivity sea surfaces). By accumulating survey area data, the cluster center will be automatically updated, avoiding the lag of manual parameter tuning.

[0122] The fine classification correction unit targets the fuzzy classes and potential misclassification points of the coarse classification and achieves accurate correction through multi-feature fusion and spatial correlation analysis. This corresponds to the engineering implementation of steps 3-4 of claim 3: Data input: receiving the fuzzy class waveforms from the coarse classification unit, the blue-green morphological features + spatial correlation features from the feature extraction unit, and the preliminary waveform data of the water / land class (used to construct a spatial reference).

[0123] DBSCAN density clustering engine: Based on three-dimensional spatial coordinates, it calculates the spatial density (e.g., the number of pulses within a 3-meter range) of each fuzzy waveform and surrounding classified waveforms, identifying core water and core land areas (areas with high density and continuous distribution); Multi-feature voting mechanism: For fuzzy waveforms, it combines blue-green morphological features (e.g., a skewness close to 0 indicating a preference for water) and spatial correlation features (e.g., if 80% of the surrounding area is water, it indicates a preference for water) for weighted voting. The weights are dynamically adjusted according to regional characteristics (e.g., higher weight for spatial features in nearshore areas and higher weight for blue-green morphological features in offshore areas); Misclassification point correction: The voting results are compared with the core areas identified by DBSCAN. If the fuzzy waveform is located in the core water area and the vote is biased towards water, it is corrected to water (and vice versa), outputting accurate classification results (including water and land labels).

[0124] The spatial association logic of this unit solves the problem of misclassification of isolated points in traditional classification. For example, a pulse in the intertidal zone is coarsely classified as land due to its low infrared amplitude, but DBSCAN reveals that it is located in the core area of ​​a contiguous water area and the blue-green waveform skewness matches the characteristics of water area. It is finally corrected to water area, improving the classification accuracy by 15%-20%.

[0125] The extremely shallow water identification unit is a supplementary identification module for extremely shallow water (water depth < 1 meter), which solves the problem of misjudging the overlapping waveforms of water surface-bottom echoes in traditional classification. The engineering implementation is as follows: Data input: Receive the waveform of suspected extremely shallow water from the output of the fine classification correction unit (meeting preset conditions: such as no obvious double peaks in the blue-green deep water channel, infrared amplitude close to land, and spatial location near the coastline).

[0126] Sample Filter: Based on conditions such as the peak difference between blue and green shallow water channels being less than the threshold waveform integral value within a specific range, further filters high-probability extremely shallow water waveforms from suspected samples (excluding obvious land waveforms); SVM Classifier: Built-in support vector machine model trained with known measured data from extremely shallow water areas, inputting blue-green multi-channel specific features (such as amplitude ratio, integral decay rate), and outputting the discrimination result of extremely shallow water or land; Result Fusion: Supplementing the SVM discrimination result to the accurate classification result, and finally outputting a complete water-land classification result including extremely shallow water (passed to the subsequent waveform decomposition module).

[0127] The model self-updating mechanism of this unit ensures long-term effectiveness. After each operation, if a misjudged sample is found during on-site verification (such as confirming that a certain area is extremely shallow water through manual sampling), it will be automatically added to the training set, and the SVM model parameters will be updated to adapt to the new geographical environment (such as the waveform difference between silty mudflats and rocky nearshore areas).

[0128] The high efficiency of the land-water classification module stems from the temporal connection and dynamic feedback between units: hierarchical progressive processing: feature extraction unit → coarse classification unit → fine classification correction unit → extremely shallow water area identification unit, forming a processing chain from simple to complex, from general to specialized, avoiding resource waste (e.g., 90% of typical waveforms can be determined in the coarse classification stage, without needing to enter the fine classification stage); data feedback mechanism: the results of the fine classification correction unit will back-optimize the cluster centers of the coarse classification unit (e.g., if extremely shallow water areas are frequently corrected, the infrared feature clustering threshold for that area will be adjusted); misjudged samples from the extremely shallow water area identification unit will be fed back to the feature extraction unit to add targeted features (e.g., specific waveform features for silt beaches); dynamic allocation of computing power: the system automatically allocates computing power according to the amount of data (e.g., the coarse classification unit uses GPU acceleration, and the fine classification unit uses CPU multi-core parallelism due to spatial calculations), ensuring that the processing time for millions of pulses is less than 30 minutes (meeting the real-time requirements of engineering operations).

[0129] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0130] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0131] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0132] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0133] In conclusion, the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for processing airborne lidar depth sounding data, characterized in that, Includes the following steps: The raw full waveform data of the airborne lidar depth sounding system is collected, including infrared laser channel data and blue-green laser multi-channel data. The original full waveform data is subjected to joint denoising processing to obtain denoised waveform data, wherein the joint denoising processing includes background noise modeling and removal based on machine learning and random noise filtering based on frequency domain low-pass filtering. Based on the denoised waveform data, a high-precision classification of land and water waveforms is performed to obtain the land and water classification results. The high-precision classification includes coarse classification based on infrared laser waveform amplitude and misclassification correction by fusing pulse spatial location information and blue-green laser multi-channel waveform features. The classified blue-green laser waveform is decomposed based on a bounded constrained Gaussian decomposition model to separate the water surface reflection, water backscattering and bottom reflection echo components. Based on the echo components, the propagation time of the laser in the air and water is calculated, and seabed topographic point clouds and inverted water optical parameters are generated.

2. The airborne lidar depth sounding data processing method according to claim 1, characterized in that, The original full waveform data is subjected to joint denoising processing, specifically as follows: The original full waveform data is segmented to separate the non-echo signal segment and the echo signal segment; A Gaussian mixture model is used to model the non-echo signal segment and fit the probability distribution of the background noise. Based on the fitted background noise probability distribution, the background noise component is subtracted from the original full waveform data to complete the removal of background noise; Perform a Fourier transform on the waveform data after removing background noise to convert it to the frequency domain; In the frequency domain, with the goal of maximizing the signal-to-noise ratio of the underwater echo signal, the optimal cutoff frequency is determined through iterative optimization for low-pass filtering. Perform an inverse Fourier transform on the filtered frequency domain signal to obtain the time-series waveform signal after removing random noise.

3. The airborne lidar depth sounding data processing method according to claim 2, characterized in that, Based on the denoised waveform data, high-precision classification of land and water waveforms is performed, specifically as follows: Extract the waveform amplitude features of the infrared laser channel from the denoised waveform data; The waveform amplitude features are initially clustered based on the K-means clustering algorithm, and the waveforms are initially classified into water and land categories to obtain coarse classification results. The three-dimensional spatial position information of each laser pulse is extracted, and the waveform features of the deep water channel and shallow water channel of the blue-green laser are extracted. The waveform features include waveform skewness, kurtosis and pulse width. Based on the three-dimensional spatial location information and the waveform features, the DBSCAN clustering algorithm is used to identify and correct misclassified points in the coarse classification results to obtain accurate water and land classification results. For extremely shallow water areas, by combining the waveform characteristics of the blue-green laser deep water and shallow water channels, a multi-channel feature vector is constructed and a support vector machine is used for classification and discrimination to achieve accurate identification of extremely shallow water areas.

4. The airborne lidar depth sounding data processing method according to claim 3, characterized in that, The bounded-constraint Gaussian decomposition model decomposes the classified blue-green laser waveforms as follows: A waveform function model for airborne laser depth sounding is constructed, in which the water surface reflected echo and the bottom reflected echo components are modeled as a single Gaussian function, and the water backscattered echo components are modeled as a superposition of multiple Gaussian functions; Based on the law of laser refraction at the air-water interface and the physical range of water attenuation coefficient, bounded constraints on the parameters of each echo component are derived and set, including amplitude, pulse width and position. Under the bounded constraints, a nonlinear least squares optimization problem is established with the goal of minimizing the waveform fitting residual. The Levenberg-Marquardt algorithm is used to solve the optimization problem, iteratively estimating the optimal parameters of each Gaussian component to complete the waveform decomposition.

5. The airborne lidar depth sounding data processing method according to claim 4, characterized in that, After the bounded-constraint Gaussian decomposition model decomposes the classified blue-green laser waveforms, the method further includes: Based on the time difference between the water surface reflected echo component and the water bottom reflected echo component obtained by decomposition, the propagation time of the laser in the water is calculated. By combining the positioning and attitude determination data provided by the airborne POS system, point cloud repositioning calculations are performed to generate high-precision seabed topographic point clouds. Based on the amplitude and waveform characteristics of the backscattered echo components of the water body, and combined with the water body attenuation model, the optical parameters of suspended sediment concentration or chlorophyll a concentration of the water body are obtained by inversion.

6. An airborne lidar depth sounding data processing system, used to implement the method according to any one of claims 1-5, characterized in that, include: The data acquisition module is configured to acquire raw full waveform data from the airborne lidar depth sounding system. A joint noise reduction module is configured to jointly remove background noise and random noise from the original full waveform data. The water and land classification module is configured to perform high-precision classification of water and land waveforms based on denoised waveform data. The waveform decomposition module is configured to decompose the blue-green laser waveform based on a bounded-constraint Gaussian decomposition model. The information extraction module is configured to calculate laser propagation time, generate seabed topographic point clouds, and invert water optical parameters based on the decomposition results.

7. The system according to claim 6, characterized in that, The joint denoising module includes: Background noise modeling unit, configured to model and remove background noise in non-echo signal segments using a Gaussian mixture model; A random noise filtering unit is configured to filter out random noise by frequency-domain low-pass filtering through iterative optimization of the cutoff frequency.

8. The system according to claim 7, characterized in that, The land and water classification module includes: A coarse classification unit is configured to perform preliminary land and water classification based on infrared laser waveform amplitude features and K-means clustering. The fine classification unit is configured to fuse pulse spatial location information and multi-channel waveform features of blue-green lasers, and uses DBSCAN clustering algorithm and support vector machine to correct misclassification and identify extremely shallow water areas.

Citation Information

Patent Citations

  • Airborne laser sounding signal extraction method and system

    CN110134976A

  • Multi-channel sounding waveform decomposition method and system in combination with morphological prior

    CN118859162A