Method for detecting abnormal value of multi-beam sounding data
By combining CUBE filtering and the isolated forest algorithm, and utilizing local uncertainty estimation and adaptive sliding window, the problem of balancing accuracy and efficiency in multibeam bathymetry data detection in complex seabed topography was solved. This resulted in efficient and accurate outlier detection, improving data quality and analytical reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SECOND INST OF OCEANOGRAPHY MNR
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-28
AI Technical Summary
Existing outlier detection methods for multibeam bathymetry data struggle to balance detection accuracy and computational efficiency when dealing with complex and varied seabed topography. Furthermore, traditional algorithms heavily rely on empirical parameter adjustments, making it difficult for intelligent algorithms to achieve both accuracy and efficiency.
By combining CUBE filtering and the isolated forest algorithm, gross errors are eliminated through local uncertainty estimation, and subtle anomalies are further identified using the isolated forest algorithm. Adaptive sliding window and feature-level fusion methods are employed to improve detection accuracy and robustness.
It significantly improves the detection accuracy of multibeam bathymetry data, reducing the false detection rate and false negative rate by about 30%, and enables efficient and accurate anomaly detection under complex seabed topography conditions, thereby improving data quality and the reliability of subsequent analysis.
Smart Images

Figure CN121934055A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of seabed topography detection technology, specifically relating to a method for detecting outliers in multibeam bathymetry data. Background Technology
[0002] With the rapid development of marine resource development, seabed topography mapping, and marine engineering construction, multibeam bathymetry, as an efficient and high-precision method for seabed topography detection, has been widely applied in the field of marine surveying. Multibeam bathymetry systems, by transmitting multiple acoustic signals and receiving their reflected signals, can quickly acquire three-dimensional water depth data of large areas of seabed topography, providing crucial data support for marine scientific research, seabed resource exploration, and seabed engineering construction. However, in actual measurement processes, multibeam bathymetry data is often affected by various factors, such as instrument noise, environmental interference, and data acquisition errors, leading to outliers. These outliers not only affect data quality but can also mislead subsequent data analysis and decision-making. Therefore, how to efficiently and accurately detect and remove outliers in multibeam bathymetry data has become an important research topic in the field of marine surveying.
[0003] Detecting outliers in multibeam bathymetry data typically employs filtering methods, with interactive filtering and automatic filtering being the two main approaches. Interactive filtering offers high filtering quality through human-computer interaction, but it suffers from long processing times and low efficiency when dealing with massive datasets. To address this, researchers have proposed several automatic filtering methods, including: 1. A two-stage filtering method based on relative density. This method uses a small threshold to remove outliers, then partitions the removed data and employs a clustering algorithm combined with least-squares surface fitting and cluster boundary selection rules to balance the trade-off between outlier removal and the preservation of effective information. However, this method is computationally inefficient when processing large-scale data, is sensitive to clustering parameters, and may exhibit blurred cluster boundaries in complex terrains such as steep slopes or gullies, leading to misidentification of effective points. 2. A trend surface filtering method that uses a polynomial surface function to fit the seabed topography based on water depth data and planar position coordinates. This method is effective in detecting outliers in flat seabed conditions; however, it fails to accurately reflect the seabed topography when conditions are complex. 3. A method for detecting gross errors in multibeam bathymetry data based on a BP neural network is proposed. This method uses a multi-layer feedforward network training and learning algorithm to fit complex curves of single ping data. It combines correlation analysis and vertical checks of adjacent ping data to locate and remove gross errors. However, it relies heavily on sample quality, has high computational complexity, and ignores many topographic features in areas with significant topographic relief. To address the challenge of outlier detection in underwater topographic data, Li et al. proposed an optimized DBSCAN-IForest stepwise anomaly detection algorithm. By automatically determining the neighborhood radius and minimum number of points in the DBSCAN algorithm, the influence of human intervention on the results is reduced. Their results show that the anomaly detection rate of this algorithm is significantly better than traditional methods. However, the impact of sample limitations and environmental variables on the algorithm's generalization ability still needs further investigation.
[0004] Significant progress has been made in existing methods for denoising and outlier detection of multibeam bathymetry data, but several challenges remain. Most traditional algorithms rely heavily on empirical parameter adjustments, making it difficult to adapt to complex and ever-changing seabed topography; while intelligent algorithms, although possessing adaptive characteristics, often struggle to strike a balance between detection accuracy and computational efficiency. Summary of the Invention
[0005] The purpose of this invention is to provide a method for detecting outliers in multibeam echo sounding data, so as to solve the above-mentioned problems.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for detecting outliers in multibeam bathymetry data, comprising the following specific steps: S1. Calculate the statistics of the neighborhood of each point in the 3D data using CUBE filtering. Based on these statistics, set a threshold and use the set threshold to identify outliers, and then remove obviously abnormal points. S2. Then, by inputting the data features after CUBE filtering into the isolated forest algorithm, the isolated forest algorithm is used to randomly select features and split points to construct a decision tree, isolate the data points, and further refine the detection of hidden and clustered anomalies. The step-by-step strategy under varying terrain conditions improves the detection accuracy and robustness.
[0007] Preferably, S1 determines anomalies based on a set threshold. The determination rule is that for multiple depth measurements within the same grid cell, each measurement value is set as follows: Uncertainty is CUBE filtering calculates the desired depth using either maximum likelihood estimation or minimum variance estimation. ; When an observation point deviates from the current depth estimate by more than a set multiple, that point is identified as an outlier.
[0008] Preferably, in S1, a sliding window is selected to set the threshold, dividing the data into multiple windows. The threshold is calculated independently in each window, and anomalies are judged based on the threshold in the window. Local thresholds are used in dense areas, and global thresholds are used in sparse areas to avoid misjudgment of global thresholds when the data distribution is uneven.
[0009] Preferably, the construction of the decision tree in S2 is specifically to generate multiple isolated trees by randomly selecting subsamples and recursively splitting them; in each isolated tree, a feature and a splitting value are recursively and randomly selected to divide the data into left and right subtrees.
[0010] Calculate the path length and, for each data point, calculate the average number of splits required to isolate it across all trees; Anomalies are determined based on anomaly score s: s≈1: anomaly detected; s≈0: normal data.
[0011] Preferably, the anomaly score S is specifically calculated using the average path length of multiple isolated trees: s( , )= ; ): The average path length of a data point across all trees; C( ): Normalization factor, related to dataset size Related; Scoring range: s∈[0,1], the closer to 1, the more abnormal it is.
[0012] Preferably, the path length in S2 is the number of splits required to get from the root node to the leaf node.
[0013] Preferably, in step S2, the data features input into the isolated forest algorithm after CUBE filtering need to be used to construct feature vectors, and finally the constructed feature vectors are input into the isolated forest algorithm.
[0014] Preferably, the feature vector construction process involves taking the CUBE-filtered data features, including the average distance of k nearest neighbors. Distance variance Local density reciprocal Planarity curvature The original coordinates X, Y, Z are used to construct a feature vector, which serves as the input feature of the isolated forest.
[0015] The technical effects and advantages of this invention are as follows: It utilizes CUBE filtering to eliminate gross anomalies based on local uncertainty estimation, and then employs the Isolation Forest algorithm to further identify subtle anomalies in the finely filtered data, enabling stratified detection of outliers. Experimental results using measured multibeam bathymetry data from a certain ocean show that, compared with traditional filtering methods, this method significantly improves detection accuracy, reducing both false positive and false negative rates by approximately 30%, verifying its efficiency and reliability in seabed mapping and analysis. CUBE filtering removes significant outliers based on local uncertainty estimation, followed by the Isolation Forest algorithm for refined detection of concealed and clustered anomalies. The step-by-step strategy under varying terrain conditions improves detection accuracy and robustness. Verification results based on actual multibeam bathymetry data demonstrate that this method can effectively identify both significant and concealed anomalies, with significantly higher detection accuracy than traditional methods and a lower false positive rate. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a diagram of the original data of the present invention; Figure 3 This is the histogram of the sliding threshold distribution of the present invention; Figure 4 This is a trend graph showing the change of the outlier threshold with depth according to the present invention. Figure 5 This is a diagram showing the detection effect of the algorithm of the present invention; Figure 6 This is a diagram showing the detection effect of the CUBE filter in this invention; Figure 7 This is a diagram showing the DBSCAN detection effect of the present invention; Figure 8 This is a diagram showing the detection effect of the Neural Network in this invention; Figure 9 This is a CIF-processed image of the present invention. Figure 10This is a statistical chart of the evaluation indicators of the present invention; Figure 11 This is a contour map of water depth according to the present invention; Figure 12 This is a diagram illustrating the seamount treatment effect of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] This invention provides, for example Figures 1-12 One method for detecting outliers in multibeam bathymetry data is shown below: For multiple bathymetry observations within the same grid cell (i.e., the dataset to be processed), let each measurement value be... Uncertainty is CUBE calculates the expected depth using either maximum likelihood estimation or minimum variance estimation. ; When an observation point deviates from the current depth estimate by more than a set multiple, that point is identified as an outlier.
[0019] Isolation Forest (IForest) is an efficient outlier detection algorithm based on unsupervised learning. It constructs a decision tree by randomly selecting features and split points to isolate data points. Outliers are characterized by their small number and significant differences in features compared to normal data, making them easier to identify quickly through random splitting. Outliers represent a small percentage of the dataset, and their feature distribution deviates from normal data, which tends to cluster in dense regions. Due to their large feature differences, outliers can be isolated into independent subspaces with only a small amount of random splitting. Unlike traditional distance- or density-based methods, Isolation Forest directly utilizes the ease with which outliers can be isolated to detect anomalies, giving it a significant advantage when handling high-dimensional and large-scale datasets.
[0020] Implementation steps of the Isolation Forest algorithm: a. Construct an isolated forest by randomly selecting subsamples and recursively splitting them into multiple isolated trees. In each isolated tree, recursively and randomly select a feature and a split value to divide the data into left and right subtrees.
[0021] b. Calculate the path length (the number of splits required from the root node to the leaf node). For each data point, calculate the average number of splits required to isolate it across all trees. Outliers typically have shorter path lengths (because they are easier to isolate quickly).
[0022] c. Anomaly detection is based on anomaly score s. s≈1: Anomaly detected. s≈0: Normal data.
[0023] Anomaly scoring: Scores are calculated based on the average path length of multiple isolated trees. s( , )= ; ): The average path length of a data point across all trees.
[0024] C( ): Normalization factor, related to dataset size Related.
[0025] Scoring range: s∈[0,1], the closer to 1, the more likely it is to be an anomaly; For each point in the 3D data, CUBE filtering is used to calculate the statistics of its surrounding neighborhood. Based on these statistics, a threshold is set, which is three times the mean standard error, to remove points that are obviously outliers (i.e., points greater than the threshold). The Isolation Forest algorithm is combined with CUBE filtering, and feature-level fusion is used to efficiently detect outliers. For each point p, the covariance matrix of points within its neighborhood radius is constructed. Then, the filtered data features, including the average distance of the k-nearest neighbors, are used. Distance variance Local density reciprocal Planarity curvature The original coordinates X, Y, Z are used to construct a feature vector, which serves as the input feature of the isolated forest.
[0026] ; Weight Configure the grid search as follows: ; ; in, is the eigenvalue of the neighborhood covariance matrix.
[0027] In the isolation forest stage, different weights are assigned to different types of features. Geometric features have high weights. Submarine topographic anomalies often manifest as abrupt changes in curvature or planar anomalies; assigning higher weights to these anomalies can enhance the model's sensitivity to structural distortions. Weights in statistical features... It suppresses random noise caused by sea state fluctuations while preserving density-gradient anomalies. Coordinates are given low weight to avoid spatial coordinates dominating the segmentation process and obscuring local patterns. Local statistical and geometric characteristics are emphasized while preserving spatial location information. The present invention does not simply feed the output of CUBE to IFOres, but rather makes targeted modifications to the structural characteristics of both to enable them to truly work together.
[0028] 1. Constructing CUBE high-order features adapted to IFORS. The Isolation Forest algorithm is combined with CUBE filtering, using feature-level fusion to efficiently detect outliers. For each point p, the points within its neighborhood radius are counted, and a covariance matrix is constructed. Then, the filtered data features, including the average distance of the k-nearest neighbors, are used. Distance variance Local density reciprocal Planarity curvature The original coordinates X, Y, Z are used to construct a feature vector, which serves as the input feature of the isolated forest (K is the number of nearest neighbors. When making a prediction for a data point, the k known data points that are closest to this point in the feature space).
[0029] ; Weight Configure the grid search as follows: ; ; in, is the eigenvalue of the neighborhood covariance matrix.
[0030] In the isolation forest stage, different weights are assigned to different types of features. Geometric features have high weights. Submarine topographic anomalies often manifest as abrupt changes in curvature or planar anomalies; assigning higher weights to these anomalies can enhance the model's sensitivity to structural distortions. Weights in statistical features... It suppresses random noise caused by sea state fluctuations while preserving density gradient anomalies. Coordinates are given low weight to avoid spatial coordinates dominating the segmentation process and obscuring local patterns. Local statistical and geometric characteristics are emphasized while preserving spatial location information.
[0031] 2. When dealing with complex seabed topographic bathymetry data, an adaptive sliding window is introduced. In anomaly detection, an adaptive sliding window is selected for threshold setting. The data is divided into multiple windows, and the threshold is calculated independently within each window. Anomalies are determined based on the threshold within each window. Local thresholds are used in dense areas, while sparse areas revert to global thresholds, avoiding misjudgments by global thresholds when data distribution is uneven.
[0032] When faced with complex seabed topography The technical barriers that prevent the two from being directly linked have been overcome, achieving true integration. Through the above-mentioned targeted modifications, by constructing a CUBE feature system adapted to IFOres, multi-scale modeling, and adaptive parameter mechanism, the above barriers have been successfully eliminated, forming a high-precision anomaly detection method for complex seabed topography, improving the accuracy and robustness of detection. Verification 1: Data description and parameter settings: The data selected for the experiment came from measured water depth data of the R2 Sonic2024 multibeam echo sounder in a sea area of the Northeast Pacific Ocean, covering an area of approximately 190m × 160m, totaling 158,720 discrete water depth data points. After installation deviation correction, attitude correction, sound velocity correction, and tide level correction, as shown... Figure 3 As shown, a large number of abnormal data were found, and the false terrain caused by them seriously affected the accurate seabed DEM and isobath mapping, so they need to be filtered out. The experiment used data from a region with significant topographic relief in the Northeast Pacific Ocean. CUBE filtering was employed for data cleaning, feature extraction, and data standardization. Missing values, duplicates, and obviously erroneous data were removed. Abnormal records with depth measurements exceeding reasonable ranges were excluded. Necessary features, including spatial coordinates and depth values, were extracted for each depth measurement point. These features were then used as input to the Isolation Forest algorithm to identify potential outliers. Considering the potential significant differences in the dimensions and numerical ranges of different features, Z-score standardization was used to process the data. This process eliminates the influence of dimensions, allowing all features to be compared under the same standard, which improves the performance and stability of the Isolation Forest algorithm.
[0033] After preprocessing the raw data, a feature selection method is used to determine the subset of features that contribute most to anomaly identification. The isolated forest algorithm is then trained on the dataset with the selected features. During training, various parameters of the isolated forest are systematically adjusted to find the optimal configuration, including key parameters such as the number of trees, the maximum depth of each tree, and the number of samples for segmentation. Cross-validation is used to ensure that the selected parameters maintain good performance on different datasets. Finally, the degree of anomaly of the data points is quantified according to a set scoring threshold, thereby more accurately identifying and judging potential outliers.
[0034] In anomaly detection, the threshold setting directly affects the accuracy and reliability of the final result. This paper's method uses a sliding window for threshold setting, dividing the data into multiple windows. The threshold is calculated independently within each window, and anomalies are determined based on the threshold within that window. Local thresholds are used in dense regions, while sparse regions revert to a global threshold, avoiding misjudgments by the global threshold when the data distribution is uneven. Figure 4As shown, the histogram exhibits two main peaks, indicating a clear grouping phenomenon in the threshold values at different depths. The right-skewed data distribution is due to the increased topographic relief in the area, with landslide areas exhibiting localized high curvature and low density characteristics, and a gradual transition between the edges and normal terrain. The anomaly threshold showed significant differences with seabed depth. Figure 4 As can be seen, the single-point thresholds indicated by orange dots fluctuate significantly, while the stratified average thresholds represented by green lines clearly show a trend with depth. In deeper regions (below -5500m), the thresholds are relatively low, possibly due to sparse anomaly distribution or relatively flat terrain. Between -5500m and -4500m, the thresholds rise rapidly, indicating potentially complex geomorphic structures in this area, thus requiring stricter anomaly detection criteria. Beyond -4500m, the thresholds gradually decrease, possibly reflecting a flatter terrain or reduced anomaly density. After setting the thresholds, anomaly detection is performed on the data, such as... Figure 5 As shown in the figure, blue dots represent normal values, while red dots indicate values identified as outliers. To verify the feasibility of the CIF algorithm for detecting outliers in multibeam bathymetry data, this paper selects three commonly used outlier detection methods as comparison objects: CUBE filtering method, neural network filtering method, and DBSCAN method. CUBE filtering, based on spatial neighborhood analysis, is widely used for outlier removal in ocean bathymetry data. By analyzing the distribution characteristics of bathymetry values within a neighborhood, outliers deviating from the normal range are identified and removed. Specifically, this paper uses three times the average elevation difference as a threshold to detect outliers. By fitting a seabed surface model, the number of outliers detected by the CUBE filtering method is as follows: Figure 6 As shown (blue indicates normal values, red indicates abnormal values).
[0035] DBSCAN is a density-based spatial clustering algorithm that uses the neighborhood radius (Eps) and minimum number of points (MinPts) to distinguish between normal and outlier values. This paper uses K-distance to determine the parameters, selecting an inflection point of 0.22 as the neighborhood radius and setting the minimum number of points to 10. Experimental results are as follows: Figure 7 As shown.
[0036] Neural network-based filtering methods utilize deep learning technology to train models that learn the complex nonlinear characteristics of data, thereby achieving automatic outlier identification and removal. This paper selects an autoencoder as a representative neural network filtering method. This method trains a surface model that fits the data and also uses three times the average elevation difference as a threshold for outlier detection. Experimental results are as follows: Figure 8As shown (blue indicates normal values, red indicates abnormal values); This paper's model first utilizes the CUBE algorithm to estimate depth values and their uncertainties at a local scale, achieving spatial gridding and physical constraint pre-screening of the data. The output is a "preliminary smoothed terrain" containing uncertainty indicators, reducing extreme outliers. Building upon this, an isolated forest model is introduced to meticulously identify and remove residual and hidden outliers. Simultaneously, a sliding window mechanism and dynamic threshold setting address the instability of detection thresholds caused by differences in data distribution under different terrain environments. By moving the window spatially and statistically analyzing local anomaly distribution characteristics in real time, the judgment threshold of the IFOres anomaly score is dynamically adjusted, enabling the detection process to adapt to local changes in complex terrains (such as seamounts and canyons).
[0037] To compare the efficiency of our method with traditional methods in detecting outliers, precision is introduced as an evaluation metric. Precision measures the proportion of truly correct predictions among all correctly predicted samples. It reflects how many of the model's positive predictions are actually correct. Recall measures the proportion of correctly predicted positive samples among all correct samples, reflecting how many actual positive examples the model found. The F1-score is the harmonic mean of precision and recall. It combines the two metrics; a higher F1-score indicates a better balance between precision and recall. Precision focuses on the accuracy of predicting positive samples. Recall focuses on how many actual positive samples the model can find. The F1-score is a comprehensive evaluation of precision and recall.
[0038] ; Where TP represents correctly identified outliers, FP represents normally identified outliers, and FN represents outliers that were incorrectly identified as normal values. It can be used to evaluate the effectiveness of outlier detection; the larger the value, the better the effect.
[0039] The effects of the four methods were compared experimentally, and the specific results are shown in the table below: Table 1 Comparison of the evaluation effects of the four methods method TP FP FN Accuracy Recall rate F-score runtime CUBE-IForest 362 23 112 94.02% 76.37% 84.26% 4.76s CUBE 471 136 253 77.59% 65.06% 70.77% 2.28s DBSCAN 463 153 346 73.96% 57.23% 64.4% 4.53s Neural Networks 556 123 263 81.89% 67.89% 74.35% 23.35s pass Figure 10 The results clearly demonstrate the performance of the CUBE-IForest method in outlier detection of multibeam bathymetry data. It exhibits superior performance in terms of accuracy, recall, and F-score, while maintaining reasonable computational efficiency. Although its recall rate is only 76.37%, the extremely low F-score indicates that it can achieve high-quality anomaly detection with minimal false positives. In applications sensitive to false positives and requiring a balance between quality and efficiency, CUBE-IForest demonstrates significant technical advantages.
[0040] In addition, by comparing the original data with the processed survey lines and water depth contour maps, the effectiveness of this method in detecting outliers in multibeam bathymetry data in areas with large seabed topographic relief and its impact on data quality can be seen more intuitively. In the contour map of the original data, some areas showed obvious abnormal fluctuations and abrupt changes in local contour lines. Figure 11 a). These abnormal fluctuations are outliers caused by noise or environmental interference during data acquisition, resulting in an inability to accurately reflect the continuity and smoothness of the seabed topography. The contour map generated from the data processed by the CUBE-IForest method ( Figure 11 b) The model effectively constrains the edge regions, preventing dense isobaths and accurately reflecting topographic features, minimizing the impact of outliers. This allows the isobath map to more accurately reflect the true characteristics of the seabed topography. Compared to the original data, the CUBE method smooths the surface well, but it is still affected by some outliers, leading to local distortions in the isobaths. Figure 11 c). DBSCAN effectively marks isolated outliers, reduces the impact of noise, and generates contour maps that are smoother than the original data and the CUBE method; however, some minor outliers still remain. Figure 11 d). Although the neural network provides a good fit, there are outliers that are not removed, resulting in dense local contour lines (d). Figure 11 e). Overall, the CUBE-IForest method performs best in removing outliers and generates the clearest and smoothest contour maps.
[0041] To further verify the detection accuracy and computational efficiency of the proposed CUBE-IForest method in complex seamount topography, the method was applied to a set of multibeam bathymetry data from a typical seamount area. Multiple anomalous spikes and noise artifacts were observed on the seamount slope and adjacent seabed in the original topography. Figure 12 a). These outliers are mainly caused by multiple reflections of sound waves on steep slopes and inconsistencies in beam footprints, which significantly affect the accuracy and continuity of topographic data; The original data presented significant flaws in its depiction of the seabed topography: firstly, the areas circled in red contain abnormally raised noise points, interfering with the judgment of the actual topography; secondly, the topographic details are rather blurry, and the layers of the seabed landforms are not clearly defined. After processing with CUBE-IForest ( Figure 12 (b) The visual representation of the terrain has been significantly improved. Noise has been largely filtered out, the color gradations of the seabed terrain are more distinct, and the overall outline is more regular, clearly showcasing the general shape of the seabed topography. CUBE processing ( Figure 12c) While it has some effect on noise suppression, local anomalies still remain, and the restoration of terrain details is somewhat lacking. Compared to CUBE-IForest processing, its optimization effect on seabed terrain is slightly insufficient. DBSCAN processing ( Figure 12 d) Although it improved the terrain representation to some extent and reduced some noise, noise residue was still quite noticeable, and the preservation of seabed topographic details was not ideal, resulting in a poor overall representation of the terrain. Neural network processing ( Figure 12 e) While most noise interference has been effectively removed, the terrain in the lower left corner is still affected by noise. Although the undulations and structural details of the seabed topography are well preserved, the overall boundaries of the seabed topography are not clear.
[0042] In summary, different data processing methods show significant differences in their optimization effects on seabed topography. Considering noise filtering, topographic detail preservation, and overall contour rendering, the CUBE-IForest method performs best in optimizing seabed topography data. It clearly identifies both the undulations of the terrain and local features, achieving a good balance between noise filtering and detail preservation. This method can provide more reliable visual and data support for the accurate analysis and research of seabed topography. This study addresses the challenges of outlier identification in multibeam bathymetry data and the insufficient efficiency and accuracy of traditional algorithms. It proposes an outlier detection method (CUBE-IForest) combining CUBE filtering and the Isolation Forest algorithm. The aim is to leverage the spatial consistency advantage of CUBE filtering and the statistical discriminative power of Isolation Forest to achieve stratified detection of different types of anomalies, thereby improving the efficiency and reliability of data cleaning. Validation using measured multibeam data from the Northeast Pacific region demonstrates that this method exhibits good overall performance in terms of detection accuracy, computational efficiency, and adaptability.
[0043] Overall, the CUBE-IForest algorithm demonstrates good performance in anomaly detection accuracy, efficiency, and adaptability, providing an effective approach to improve the reliability of multibeam bathymetry data and the accuracy of subsequent terrain analysis. This research not only has reference value for multibeam measurement data processing but also provides new ideas for anomaly identification in complex spatial data. At the application level, its quality control and automatic outlier removal of multibeam bathymetry data also support seafloor topography modeling, geological hazard identification, and seafloor structural feature analysis. However, the algorithm still has some limitations. The validation in this study was mainly based on offline data processing and has not fully considered the algorithm's stability and operating efficiency under real-time data streams. Future research directions will focus on adaptive parameter optimization, the integration of the algorithm with deep learning models, and embedded applications in real-time measurement systems to further improve the robustness and practicality of the method.
[0044] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting outliers in multibeam bathymetry data, characterized in that: The specific steps are as follows: S1. Calculate the statistics of the neighborhood around each point P in the three-dimensional data by using CUBE filtering. Based on the statistics, set a threshold value of three times the mean error. Determine outlier points based on the set threshold value and then remove points that are greater than the threshold value. S2. Then, by inputting the data features after CUBE filtering into the isolated forest algorithm, the isolated forest algorithm is used to randomly select features and split points to construct a decision tree, isolate the data points, and further refine the detection of hidden and clustered anomalies. The step-by-step strategy under varying terrain conditions improves the detection accuracy and robustness.
2. The method for detecting outliers in multibeam bathymetry data according to claim 1, characterized in that: S1 determines outliers based on a set threshold. The rule for determining outliers is that for multiple depth measurements within the same grid cell, each measurement value is set to... Uncertainty is CUBE filtering calculates the desired depth using either maximum likelihood estimation or minimum variance estimation. ; The interval i is [1, n], where n represents the total number of depth observations of the grid cells; When an observation point deviates from the current depth estimate by more than the threshold, the point is identified as an outlier.
3. The method for detecting outliers in multibeam bathymetry data according to claim 1, characterized in that: In S1, a sliding window is selected to set the threshold, dividing the data into multiple windows. The threshold is calculated independently within each window, and anomalies are judged based on the threshold within the window. Local thresholds are used in dense areas, while global thresholds are used in sparse areas to avoid misjudgment of global thresholds when the data distribution is uneven.
4. The method for detecting outliers in multibeam bathymetry data according to claim 2, characterized in that: In S2, the decision tree is constructed by randomly selecting subsamples and recursively splitting them to generate multiple isolated trees. In each isolated tree, a feature and a splitting value are recursively and randomly selected to divide the data into left and right subtrees. Calculate the path length and, for each data point, calculate the average number of splits required to isolate it across all trees; Anomalies are determined based on anomaly score s, where s≈1: anomaly detected; s≈0: normal data.
5. The method for detecting outliers in multibeam bathymetry data according to claim 4, characterized in that: The anomaly score S is specifically calculated using the average path length of multiple isolated trees: s( , )= ; ): The average path length of a data point across all trees; C( ): Normalization factor, related to dataset size Related; Scoring range: s∈[0,1], the closer to 1, the more abnormal it is.
6. The method for detecting outliers in multibeam bathymetry data according to claim 4, characterized in that: The path length in S2 is the number of splits required to get from the root node to the leaf node.
7. The method for detecting outliers in multibeam bathymetry data according to claim 1, characterized in that: S2 inputs the CUBE-filtered data features into the Isolation Forest algorithm. The data features then need to be used to construct feature vectors, which are then input into the Isolation Forest algorithm.
8. The method for detecting outliers in multibeam bathymetry data according to claim 7, characterized in that: The process of constructing the feature vector involves taking the features of the CUBE-filtered data, including the average distance of the k nearest neighbors. Distance variance Local density reciprocal Planarity curvature The original coordinates X, Y, Z are used to construct a feature vector, which serves as the input feature of the isolated forest.