Acoustic characterization method for deep-sea manganese nodule coverage based on multi-classifier decision fusion

By employing a multi-classifier decision fusion method, combined with multibeam echo sounding system and optical towed body data, and optimizing feature selection and robust estimation, the problems of low sampling efficiency and noise interference in deep-sea manganese nodule coverage measurement are solved, achieving high-precision quantitative characterization of coverage and ecological protection.

CN120877082BActive Publication Date: 2025-11-28SHANDONG UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511394103.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-28
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing technologies for measuring the coverage of deep-sea manganese nodules suffer from problems such as low sampling efficiency, limited coverage, severe noise interference, poor model stability, and insufficient feature selection, making it difficult to achieve high-precision quantitative characterization of coverage and ecological environmental protection.

Method used

A multi-classifier decision fusion method is adopted, combining the Kongsberg EM122 multibeam echo sounder system and a 6000-meter integrated optical towed body. The Boruta algorithm is used to select features, and an iterative robust estimation algorithm with a sliding window and a stacking mechanism are introduced to integrate the advantages of multiple classifiers at the decision level to generate a spatial distribution map of nodule coverage.

Benefits of technology

It has achieved high-precision and continuous spatial prediction of deep-sea manganese nodule coverage, which has improved exploration efficiency, reduced sampling costs, supported resource reserve assessment and ecological protection, and improved the stability and generalization ability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877082B_ABST
    Figure CN120877082B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of deep-sea mineral resource exploration, and discloses a deep-sea manganese nodule coverage acoustic characterization method based on multi-classifier decision fusion. By extracting 14 backscattering texture features and 4 seabed topography features, and using Boruta algorithm to optimize the extracted features, a refined feature set for characterizing the seabed is constructed. An iterative robust estimation algorithm based on a sliding window is introduced to identify and suppress gross errors in the feature image by dynamically adjusting the observation weight through continuous smoothing of the data in different directions. In the model construction stage, the Stacking mechanism in ensemble learning is used to integrate the advantages of multiple classifiers for decision-level integration and generate a spatial distribution map of nodule coverage. The Boruta feature selection algorithm is introduced to screen out the key features with the strongest discrimination ability for nodule coverage from the original acoustic features, avoiding redundant interference and improving the accuracy and prediction discrimination ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of deep-sea mineral resource exploration, and particularly relates to an acoustic characterization method for deep-sea manganese nodule coverage based on multi-classifier decision fusion. BACKGROUND

[0002] In order to deeply understand the distribution characteristics of manganese nodules on the deep-sea seabed, it is crucial to obtain high-quality seabed information. With the development of the industry, the investigation method for benthic mineral resources is constantly iterating. The direct sampling method (such as using a grab to obtain samples) is the most direct and effective method to obtain seabed information. However, in the deep-sea environment, large-scale data acquisition through direct sampling faces great challenges. On the one hand, data collection at each sampling point takes a long time, and the overall sampling efficiency is low; on the other hand, large-scale high-density sampling may cause irreversible damage to the seabed ecological environment. The emergence of integrated optical towed bodies (near-bottom photography and video) provides a more efficient and more detailed solution for seabed sediment investigation. Through continuous photography or video recording during the process of towing the body, a large amount of high-definition sampling information can be obtained. However, compared with the shallow water environment, the physical characteristics of the deep sea make the coverage range of the optical towed body extremely limited (only a few meters). Its investigation area is usually limited to the cross section along the survey line, thereby affecting the full coverage investigation capability of the seabed sediment. Compared with the above method, the underwater acoustic method provides an economical and efficient means of measuring seabed topography. The ship-borne multi-beam echo sounder (MBES) can cover a large area of the seabed (generally 3-4 times the water depth), and since it can record both the seabed topography and the backscatter intensity, it has the characteristics of high precision, high efficiency and full coverage, and has become a mainstream choice for mapping seabed topography, detecting seabed sediment type, detecting seabed current flow and detecting benthic organisms.

[0003] Deep-sea manganese nodules are rich in various rare metal elements, and are important mineral resources for alleviating the contradiction between supply and demand of land resources and supporting energy transformation. In addition to the metal resources contained therein, manganese nodules are also considered to be the core of the deep-sea ecosystem. They are one of the few hard bases in the typical large-scale scene based on soft silty clay in deep-sea basins. Research shows that the nodule seabed area coverage rate is closely related to the abundance, spatial distribution, species richness and community composition of benthic animals and microorganisms. As a key indicator for resource evaluation, high-precision nodule coverage prediction is of great significance for resource evaluation and development. Due to the complexity and spatial heterogeneity of the deep-sea environment, traditional surface sediment and manganese nodule detection mainly faces attribute classification, and there are difficulties in quantifying resource evaluation indicators such as the coverage rate of seabed polymetallic nodules.

[0004] Through the above analysis, the problems and defects of the prior art are:

[0005] (1) Traditional large-scale data acquisition methods through direct sampling have long data collection time for each sampling point, low overall sampling efficiency, and large-scale high-density sampling may cause irreversible damage to the seabed ecological environment. The coverage range of integrated optical towed bodies is extremely limited, and the investigation area is usually limited to the cross section along the survey line, thereby affecting the full coverage investigation capability of the seabed sediments.

[0006] (2) Acoustic detection results are limited to binary discrimination and lack quantitative coverage characterization capability. The shipborne multi-beam echo sounding system (MBES) has become the mainstream detection means and can provide wide-range water depth and backscatter intensity data. However, existing researches in the field of bottom type classification / resource exploration mostly stop at binary classification of "existence / absence", and cannot realize quantitative analysis of coverage, so as to be difficult to support fine resource reserve evaluation and ecological environment protection. And the existing classification modeling relies on a single model, and the result stability is poor. The existing researches mostly use a single classifier for prediction, and different algorithms have significant differences in complex seabed environment, and lack stability and robustness. The single model excessively depends on specific assumptions or sample distribution, and the prediction result is biased, and it is difficult to adapt to the spatial heterogeneity of manganese nodule coverage.

[0007] (3) The feature construction and optimization are insufficient. The existing method is limited in the use of acoustic features, and often ignores the comprehensive effect of backscatter texture and seabed topography and other multi-dimensional information. At the same time, the feature selection depends on artificial experience, which is easy to introduce redundancy and noise, and weaken the model discrimination ability. Lack of systematic feature optimization mechanism leads to poor stability and generalization of the prediction result.

[0008] (4) Noise and gross error interference is difficult to effectively suppress. There are ship noise, reverberation, bubble and biological activity interference in deep sea environment, which often causes gross error and abnormal value in water depth and backscatter intensity observation. The existing method is mostly one-way or local processing, and lacks robustness, and is difficult to fully eliminate the influence of gross error on prediction. SUMMARY

[0009] To overcome the problems in the related art, the embodiment of the present application provides a deep-sea manganese nodule coverage acoustic characterization method based on multi-classifier decision fusion. The technical solution is as follows:

[0010] The present application is realized in the following way: a deep-sea manganese nodule coverage acoustic characterization method based on multi-classifier decision fusion, comprising the following steps:

[0011] S1, based on Kongsberg EM122 multi-beam sounding system, obtain water depth and backscatter intensity data, combined with 6000-meter integrated optical tow body to obtain near-bottom images, the positioning of the near-bottom optical tow body uses ultra-short baseline USBL combined with inertial navigation system INS and Doppler velocity recorder DVL to obtain the positioning of the underwater optical sampling point, the shipborne multi-beam obtains the precise underwater topographic model through the global navigation satellite system GNSS combined with inertial navigation system INS; by constructing the sampling points and the ship-measured terrain in the same projection manner, the water depth and backscatter intensity data at the location of the sampling points are obtained;

[0012] S2, by extracting 14 backscatter texture features and 4 seafloor topographic features, using Boruta algorithm to optimize the extracted features, constructing a feature set that finely represents the seafloor;

[0013] S3, introduce the iterative robust estimation algorithm based on sliding window, by continuously smoothing the data in different directions to dynamically adjust the observation weight, identify and suppress the gross errors in the feature image;

[0014] S4, in the model construction stage, using the Stacking mechanism in ensemble learning, integrating the advantages of multiple classifiers for decision-level integration to generate the spatial distribution map of nodule coverage.

[0015] The training of the base classifier uses the k-fold cross-validation strategy to ensure the robustness of the prediction features (Meta-features). Each base classifier generates out-of-fold results through cross-prediction on the training set, which are used as new features input to the secondary classifier (Gradient Boosting), thereby realizing the nonlinear fusion of the decision layer. The final output is a discrete coverage level (low, medium, high), and the spatial distribution of deep-sea manganese nodules is presented in the form of a classification map.

[0016] In step S2, the backscatter texture features include water depth, slope, curvature, aspect, backscatter intensity, 、 、 、 、 、 、 、 , standard deviation, kurtosis, skewness, energy, entropy, and seafloor topographic features include depth, slope, aspect, and planar curvature.

[0017] In step S2, the Boruta algorithm is used to optimize the features, including:

[0018] The Boruta algorithm expands the feature space by using the method of supplementing random attributes to eliminate the correspondence between sample attribute values and labels, and uses the expanded feature space for classification test to calculate the importance index of all attributes; the importance index of the shadow attribute is used as a reference to evaluate the importance degree of each attribute;

[0019] The importance of each feature variable is evaluated by using the Boruta algorithm, including:

[0020] (1) Initialization and data preparation: the data set is , wherein is the th feature, is the total number of features; the data set also contains the target variable , which represents the label of each sample;

[0021] (2) Generate shadow features: the Boruta algorithm generates a set of random "shadow features" to simulate noise to determine whether the real features are important; for each original feature , a shadow feature is generated, and the value of the shadow feature is obtained by randomly disturbing the sample labels in , and only represents noise;

[0022] (3) Merge original features and shadow features: merge the original features and the shadow features to form a data set containing original features and shadow features; use the merged feature set to train a random forest model to estimate the importance of each feature; the random forest model outputs the importance score of each feature and compares the importance of the original features and the shadow features.

[0023] Further, the Boruta algorithm uses the Gini index to measure the importance of the features;

[0024] For each feature , a random forest model is trained and the importance is calculated to determine the contribution to the prediction of the target variable ; the random forest model evaluates the importance of the features according to the contribution of each feature.

[0025] Further, the random forest model compares the importance of the original features and the shadow features, including:

[0026] Calculate the maximum importance of all shadow features :

[0027] ;

[0028] For each original feature , the importance is compared with ; if , the feature is considered important; otherwise, the feature is considered useless; the Boruta algorithm runs multiple times by constantly iterating the image feature importance comparison step until the classification of each feature is stable.

[0029] In step S3, the sliding window-based iterative robust estimation includes:

[0030] (1) sliding on the feature image with a window of a specified size;

[0031] (2) in each window, local pixel sequences are extracted along multiple directions (including horizontal, vertical, and diagonal directions) respectively; by constructing one-dimensional pixel chains in different directions, the spatial correlation of acoustic scattering signals in the nodule coverage area is utilized to enhance the sensitivity of the feature extraction process to spatial texture and nodule edge structure.

[0032] (3) based on the Huber loss function, weights are assigned to the observation points in the pixel sequences of each direction, and the iterative weighted least squares method is used to update the local regression model parameters of the direction until convergence, obtaining the robust estimation result of the direction; robust regression is performed on the sequence data of each direction, and the Huber function is used to reduce the influence of gross errors caused by noise pulses, deep water bubbles, or local abnormal reflections, so that the estimation result of the nodule coverage rate related feature is more stable.

[0033] (4) the direction weights are dynamically calculated according to the robust standard deviation of the residuals of each direction, and the estimation results of different directions are weighted and fused to obtain the final robust estimation value of the center pixel of the window;

[0034] (5) move the window to the next position and repeat steps (1)-(4) until the entire feature image is traversed; through the multi-direction weighted residual and iterative optimization process, robust parameter estimation is obtained which is not sensitive to noise and gross errors.

[0035] The sliding window-based iterative robust estimation algorithm: the core idea is to use robust estimation (Huber) to fit a local model in each window, predict the estimation value of the center point of the window, and replace the center pixel

[0036] (1) for a pixel , a local sliding window is defined, with a fixed size of 5x5 and a step size of 1 pixel (pixel-by-pixel processing), to ensure that the boundary prediction is concentrated and complete.

[0037] (2) extract a 1D pixel chain (length same as window size, 5 pixel units in this method) along Four directions are used for each chain, and a robust regression is performed on the 1D pixel chain using a Huber kernel function to obtain the center estimate in the direction and the residual scale in the direction .

[0038] (3) A weighted average of the four estimates obtained in the above directions is taken as the final output; the expression is:

[0039] ;

[0040] wherein, is the weight coefficient in the direction at the position is a very small positive number (regularization term) to prevent the denominator from being zero; is the final estimate value at the position is the normalized weight to ensure that the weight sum of all directions is 1. In step S4, the multi-classifier decision fusion based on the Stacking mechanism includes:

[0041] Five supervised classification algorithms are selected as the first layer classifiers, and the data is independently trained and predicted; the prediction results of the first layer model are used to evaluate the individual performance of each model and are passed to the second layer model as input features; a performance screening threshold is set, and if the prediction accuracy of a certain first layer classifier is lower than the threshold, the prediction result is not passed to the second layer meta-learning model; a gradient boosting classifier is used as the final decision model to integrate the prediction information of the first layer classifiers.

[0042] Further, in the spatial prediction of the coverage of the tubercle, the first layer model of the Stacking ensemble learning method classifies each pixel to obtain several independent prediction results;

[0043] After the low-precision models are screened out by the performance screening mechanism, the remaining prediction results are combined to form a new feature set, which is input to the meta-model for final decision; the decision process is based on the majority voting principle to ensure that the final category determination integrates the advantages of different classifiers.

[0044] Further, the supervised classification algorithms include random forest, decision tree, BP network, support vector machine, and K-nearest neighbor classifier.

[0045] Further, the supervised classification algorithms include random forest, decision tree, BP network, support vector machine, and K-nearest neighbor classifier.

[0046] ​Another object of the present application is to provide a deep-sea manganese nodule coverage acoustic characterization system based on multi-classifier decision fusion, which is used to regulate the deep-sea manganese nodule coverage acoustic characterization method based on multi-classifier decision fusion, and the system comprises:

[0047] A data acquisition module is configured to acquire water depth and backscattering intensity data based on a Kongsberg EM122 multi-beam sounding system, acquire near-bottom images by a 6000-meter integrated optical towed body, acquire the positioning of the near-bottom optical towed body by an ultra-short baseline (USBL) combined inertial navigation system (INS) and a Doppler velocity log (DVL) to obtain the positioning of an underwater optical sampling point, and acquire a precise underwater topographic model by a shipborne multi-beam through a global navigation satellite system (GNSS) combined INS.

[0048] A feature optimization module is configured to extract 14 backscattering texture features and 4 seafloor topographic features, optimize the extracted features by using a Boruta algorithm, and construct a feature set for fine characterization of the seafloor.

[0049] An iterative robust estimation module is configured to introduce an iterative robust estimation algorithm based on a sliding window, dynamically adjust the observation weight by continuously smoothing the data in different directions, and identify and suppress gross errors in the feature image.

[0050] A multi-classifier decision fusion module is configured to use a Stacking mechanism in ensemble learning, integrate the advantages of multiple classifiers for decision-level integration, and generate a spatial distribution map of the nodule coverage.

[0051] In combination with all the above technical solutions, the present application has the following beneficial effects:

[0052] Firstly, the present application focuses on the high-precision and continuous spatial prediction of deep-sea polymetallic nodule coverage, constructs a prediction framework that integrates multi-source data and robust modeling, and emphasizes the improvement of the spatial interpretation capability of manganese nodule coverage. The present application acquires water depth and backscattering intensity data based on a Kongsberg EM122 multi-beam system, acquires near-bottom images by a 6000-meter integrated optical towed body; extracts 14 backscattering texture features and 4 seafloor topographic features, optimizes the extracted features, and constructs a feature set for fine characterization of the seafloor; introduces a robust estimation method to identify and suppress gross errors in the feature image, thereby improving the robustness and usability of the data; in the model construction stage, a Stacking mechanism in ensemble learning is used to integrate the advantages of multiple classifiers for decision-level integration, and a spatial distribution map of the nodule coverage is generated.

[0053] Secondly, the Boruta feature selection algorithm is introduced to screen out the key features with the strongest discrimination ability for nodule coverage from the original acoustic features, thereby improving the precision and generalization ability of the model. The Boruta feature selection algorithm can improve the feature effectiveness, avoid redundant interference, and enhance the prediction discrimination ability of the model in the deep-sea nodule coverage prediction task, thereby providing a solid feature foundation for subsequent manganese nodule coverage prediction. Compared with the single classifier model optimized based on the Boruta feature selection and robust estimation, the Stacking method improves the prediction accuracy by 0.38% (compared with RF), 36.21% (compared with SVM), 12.49% (compared with BP neural network), and 4.32% (compared with KNN), respectively, further verifying the effectiveness and applicability of the ensemble strategy in the prediction of manganese nodule distribution in the deep-sea complex environment.

[0054] Thirdly, the present application can quantitatively characterize the deep-sea manganese nodule coverage with high precision under the condition of a wide range and non-contact, thereby providing an efficient and low-cost solution for international seabed resource exploration and development. The application can not only significantly improve the exploration efficiency and reduce the sampling cost, but also support the evaluation of deep-sea mineral resources reserves and the collaborative decision-making of ecological protection. The present application realizes the multi-level quantitative prediction of deep-sea manganese nodule coverage for the first time and proposes a systematic framework integrating Boruta feature selection, robust estimation, and Stacking ensemble learning.

[0055] Fourthly, the present application effectively solves the problem of quantitative prediction of coverage in a high-noise environment by eliminating gross errors through robust estimation, screening key features through Boruta, and using multi-classifier fusion. The present application realizes the previously unattainable goal. Traditional methods generally rely on a single classifier or empirical feature selection, rely on specific model assumptions or subjective judgments, and result in large prediction deviations. The present application breaks through the limitations of single models and empirical judgments, overcomes the bias of traditional technologies, and realizes more objective, robust, and universal coverage prediction by introducing multi-classifier Stacking fusion and data-driven Boruta feature optimization. BRIEF DESCRIPTION OF DRAWINGS

[0056] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure;

[0057] Figure 1 is a schematic diagram of the iterative robust estimation algorithm based on the sliding window provided by the embodiments of the present application;

[0058] Figure 2 is a principle diagram of the deep-sea manganese nodule coverage acoustic characterization method based on multi-classifier decision fusion provided by the embodiments of the present application;

[0059] Figure 3 is an importance score figure of different features provided by the embodiment of the present application;

[0060] Figure 4 is a coverage feature correlation figure of manganese nodule provided by the embodiment of the present application;

[0061] Figure 5 (a) is an RF figure of nodule coverage prediction value provided by the embodiment of the present application;

[0062] Figure 5 (b) is an SVM figure of nodule coverage prediction value provided by the embodiment of the present application;

[0063] Figure 5 (c) is a BP figure of nodule coverage prediction value provided by the embodiment of the present application;

[0064] Figure 5 (d) is a KNN figure of nodule coverage prediction value provided by the embodiment of the present application;

[0065] Figure 5 (e) is a DT figure of nodule coverage prediction value provided by the embodiment of the present application;

[0066] Figure 6 (a) is an RF figure of nodule coverage prediction value after robust estimation provided by the embodiment of the present application;

[0067] Figure 6 (b) is an SVM figure of nodule coverage prediction value after robust estimation provided by the embodiment of the present application;

[0068] Figure 6 (c) is a BP figure of nodule coverage prediction value after robust estimation provided by the embodiment of the present application;

[0069] Figure 6 (d) is a KNN figure of nodule coverage prediction value after robust estimation provided by the embodiment of the present application;

[0070] Figure 6 (e) is a DT figure of nodule coverage prediction value after robust estimation provided by the embodiment of the present application;

[0071] Figure 7 is a nodule coverage prediction result figure provided by the embodiment of the present application using Stacking mechanism. DETAILED DESCRIPTION

[0072] In order to make the above objectives, features and advantages of the present application more apparent, specific embodiments of the present application are described in detail below with reference to the accompanying drawings. In the following description, a lot of specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many different ways other than the ones described herein, and one of ordinary skill in the art can make similar improvements without departing from the spirit of the present application, so the present application is not limited to the specific implementations disclosed below.

[0073] Embodiment 1, the methodology of the present application mainly consists of two parts. The first part revolves around the feature-oriented correlation analysis, focusing on the feature mining and screening of the acoustic data collected by MBES. In order to ensure the stability and accuracy of the subsequent classification process, the present application introduces the robust estimation method to eliminate gross errors in the data on the basis of the preferred features, thereby improving the quality of the feature data.

[0074] The second part focuses on the estimation of the nodule coverage. The present application first adopts five kinds of classification for preliminary classification, and further combines the Stacking integrated learning mechanism to integrate the advantages of different models, so as to improve the robustness and generalization ability of the prediction model. The overall technical process is as shown in Figure 1

[0075] As shown in Figure 2 The acoustic characterization method for deep-sea manganese nodule coverage based on multi-classifier decision fusion provided by the embodiment of the present application specifically comprises the following steps:

[0076] S1, based on Kongsberg EM122 multi-beam sounding system, water depth and backscattering intensity data are obtained, and 6000-meter integrated optical towed body is used to obtain near-bottom image, the positioning of the near-bottom optical towed body uses ultra-short baseline USBL combined inertial navigation system INS and Doppler velocity recorder DVL to obtain the positioning of the underwater optical sampling point, the shipborne multi-beam obtains the precise underwater topographic model through the global navigation satellite system GNSS combined inertial navigation system INS, the sampling point and the ship-surveyed topography are constructed in the same projection mode to obtain the water depth and backscattering intensity data at the position of the sampling point;

[0077] ​Different features perform differently in the task of characterizing the spatial distribution of manganese nodule coverage, so multi-dimensional feature information needs to be extracted to improve the accuracy and reliability of the prediction. Backscattering intensity is closely related to the physical properties of seabed sediments, such as roughness, sediment grain size, porosity, saturation, and incident angle. At the same time, seabed topography changes the hydrodynamic model in the region, thereby affecting the spatial distribution pattern of nodules. Integrating these information effectively will effectively improve the accuracy and interpretability of the prediction of the spatial distribution of nodules. Therefore, before training and predicting the spatial distribution information of nodules using MBES data, systematic feature extraction is needed. The present application extracts 18 dimensions of features, including 14 dimensions of backscattering intensity features based on multi-beam and 4 dimensions of seabed topography features. However, due to the complexity of the deep-sea environment, the underwater acoustic signal is disturbed by seabed reverberation, ship noise, biological activity noise and other factors during long-distance propagation, resulting in a large amount of noise in the data. Therefore, before feature extraction, all MBES data is manually removed. Table 1 details the extracted feature factors.

[0078] (1) Backscattering intensity features: Backscattering intensity features mainly include backscattering angle response curve features and backscattering intensity image features. The present application mainly focuses on backscattering intensity image features, i.e. sonar image texture features. Texture features can effectively distinguish the spatial distribution information of different coverage of nodules.

[0079] (2) Seabed topography features: MBES can obtain high-quality water depth data and obtain seabed topography information, which can be extracted from the processed DEM data. The present application extracts four seabed topography features from water depth raster data, including depth, slope, aspect, and planar curvature.

[0080] Table 1: Extraction summary table of texture and topography features extracted based on multi-beam backscattering intensity and bathymetric data and calculation method of each feature

[0081]

[0082] S2, by extracting 14 backscattering texture features and 4 seabed topography features, using Boruta algorithm to optimize the extracted features, and constructing a refined feature set to represent the seabed;

[0083] Backscattering texture features include water depth, slope, curvature, aspect, backscattering intensity, , , , , , , , Standard deviation, kurtosis, skewness, energy, entropy, and seafloor topographic features include depth, slope, aspect, and planar curvature.

[0084] In learning tasks, feature selection aims to filter out the most discriminative feature variables from all attributes for a specific task, thereby constructing a concise and efficient feature space. This process effectively eliminates redundant and erroneous features, avoiding the adverse effects of high-dimensional data on classifier performance. In this invention, feature engineering is particularly crucial for predicting the spatial distribution of manganese nodule coverage. The extracted acoustic feature datasets often suffer from high noise and nonlinearity; excessively high dimensionality or inaccurate features can weaken the performance of subsequent classifiers. Feature selection and optimization strongly influence the design and performance of the classifier. This invention selects the Boruta algorithm for feature selection, extracting the most representative feature variables to provide a solid data foundation for accurately predicting the spatial distribution of deep-sea manganese nodules.

[0085] The Boruta algorithm is a "fully relevant" feature selection method that identifies all predictor variables that may be relevant to classification. It is a wrapper-style supervised feature selection algorithm built on a random forest model. It expands the feature space by adding random attributes to eliminate the correspondence between sample attribute values ​​and labels. Then, it uses this expanded feature space for classification testing and calculates the importance index of all attributes. Due to random fluctuations, the importance index of shadow attributes may be non-zero. Using the importance index of shadow attributes as a reference can assess the importance of each attribute. Since the distribution level of the importance index varies with the randomness of the classifier and is related to the presence of unimportant features and the specific realization of shadow attributes, the process of assigning random values ​​to shadow attributes needs to be repeated multiple times to obtain statistically significant results. The steps for evaluating the importance of each feature variable using the Boruta algorithm are as follows:

[0086] First, initialization and data preparation are required. The dataset is... ,in, It is the first One characteristic, It is the total number of features; the dataset also contains the target variable. , used to represent the label of each sample.

[0087] Next, image features need to be generated. Unlike traditional feature selection methods, Boruta generates a set of random "image features" to simulate noise, thereby determining whether the true features are important enough. For each original feature... Generate an image feature Its value is obtained by randomly shuffling The image features generated in this way are uninformative and should only represent noise.

[0088] Next, the original features and the image features will be combined. The original features and the image features will be merged to form a dataset containing both original features and image features. The combined feature set will be used to train a random forest model to estimate the importance of each feature. Typically, Boruta uses the Gini index to measure the importance of a feature. For each feature , the present invention trains a random forest and calculates its importance (i.e. its contribution to the prediction of the target variable ). The random forest model assesses the importance of a feature based on its contribution.

[0089] (1);

[0090] In the Gini index, is the number of classes, is the proportion of samples of the th class in the node.

[0091] For each feature , the importance is calculated using the random forest. Here, is the model's assessment of the Gini impurity of the feature with respect to the target , i.e. a measure of its contribution.

[0092] (2);

[0093] The importance score of each feature output by the random forest is compared, and the importance of the original features and the image features is compared. The specific steps are as follows: calculate the maximum importance of all image features :

[0094] (3);

[0095] For each original feature , its importance is compared with . If , the feature is considered important; otherwise, the feature is considered useless. The above steps are iterated until the classification of each feature ("important" or "useless") is stable. Typically, Boruta is run multiple times to ensure the stability and accuracy of feature selection. Only important variables are retained for subsequent spatial distribution prediction of manganese nodule coverage.

[0096] S3, an iterative robust estimation algorithm based on sliding window is introduced, which can identify and suppress the gross errors in the feature image by dynamically adjusting the observation weight through continuous smoothing of the data in different directions;

[0097] In the process of deep-sea multi-beam measurement, the bottom reverberation and environmental noise will inevitably produce gross errors in the observation of water depth and backscattering intensity. In the modern measurement adjustment theory, if the gross error is too different from the prior random model and the actual model, the gross error can be classified as a random model. At this time, it can be explained as a variance inflation model. The processing idea is to constantly modify the weight or variance of the observation value according to the results of the iterative adjustment, so that the weight of the observation value containing the gross error tends to zero or the variance tends to infinity, so as to ensure that the estimated parameters are less affected by the model error, especially the gross error. Robust estimation is an effective method to solve this problem. In the prior art, Zhang et al. group the BS data according to the incident angle, and for each group, the robust estimation is applied to the sliding window in the along-track direction to process the along-track heterogeneity. Although this method can effectively eliminate the gross error in each group of data, due to the limitation of the grouping strategy, the gross error elimination process is mainly aimed at a single direction, and the multi-directional harmful data interference from the surrounding area is ignored. Therefore, the present application further considers the possible influence of the seabed gross error in the along-track and cross-track directions. For this purpose, the present application proposes an iterative robust estimation algorithm based on sliding window. This method can dynamically adjust the observation weight by continuously smoothing the MBES observation data in different directions, so as to reduce the influence of local outliers on the final estimation result, thereby achieving more comprehensive gross error elimination.

[0098] Iterative robust estimation algorithm based on sliding window: the core idea is to use robust estimation (Huber) to fit a local model in each window, predict the estimated value of the center point of the window and replace the center pixel

[0099] (1) For the pixel , a local sliding window is defined, the size is fixed at 5x5, and the step is 1 pixel (pixel-by-pixel processing), so as to ensure the central prediction and complete coverage;

[0100] (2) Extract the 1D pixel chain (length same as window size, 5 pixel units in this method) where the center pixel is located, and use four directions to each chain, do robust regression on this 1D pixel chain, and the kernel function adopts Huber function; obtain the center estimate of this direction and the residual scale of this direction.

[0101] (3) Weighted average of the four estimated values obtained in the above directions is taken as the final output; the expression is:

[0102] ;

[0103] In the formula, For in position Location, along direction The weighting coefficients, It is a very small positive number (regularization term) to prevent the denominator from being zero; For in position The final estimated value at that location, To normalize the weights, the sum of the weights in all directions is guaranteed to be 1.

[0104] like Figure 1 As shown, in robust estimation, the robust estimation method adopted in this invention is Iteratively Reweighted Least Squares (IRLS). "Outlier removal" is not an independent step completely separate from the main process, but is implemented gradually through a weighting function during the iterative fitting process. Outliers are not hard-removed points during multiple iterations in the Huber function, but rather soft-removed. Points with residuals exceeding a threshold are not directly discarded, but assigned a very small weight.

[0105] With the development and progress of modern measurement adjustment theory, various robust estimation methods have been developed to handle potential outliers. When some observations contain gross errors, robust estimation is superior to least squares estimation. The fundamental difference between robust estimation and classical estimation theory lies in that robust estimation is based on the actual distribution pattern of the observed data, rather than on some ideal distribution pattern. This invention uses the M-estimation theory proposed by Huber in 1966, a method introduced into the measurement community by Krarup and Kubik of Denmark in 1980. First, the parameters of the regression model are initialized, and an initial parameter estimate can be obtained using the least squares (OLS) method. Let the model parameters be:

[0106] (4);

[0107] In the formula, It is the number of features of the model. It is the first The actual observed values ​​of each data point That is the corresponding predicted value.

[0108] In the initialization phase, this invention obtains the initial parameters by minimizing the traditional squared error:

[0109] (5);

[0110] where, is the feature vector of the th data point.

[0111] The second step is to calculate the residual error for each point in the sliding window, i.e., the difference between the actual value and the predicted value. The residual error reflects the fitting error of the model to the data. For each data point, its residual error is calculated based on the prediction result of the current model.

[0112] (6);

[0113] where, is the parameter estimate of the current iteration.

[0114] The third step is to introduce the Huber loss function to calculate the weighted residual error for each observation point in the sliding window. A weight is assigned to the weighted residual error. The Huber loss function for the calculation of the weighted residual error is as follows:

[0115] (7);

[0116] where, is the threshold value that determines when to switch from the square loss to the linear loss.

[0117] The fourth step is to update the weight matrix, which is used to update the weights of the regression model. Specifically, the weight matrix is a diagonal matrix, where each element corresponds to the weighted residual error of each sample:

[0118] (8);

[0119] The fifth step is to minimize the weighted loss. Next, the parameters are updated by the weighted least squares method. The optimization objective is to minimize the weighted loss function:

[0120] (9);

[0121] This optimization problem can be solved by an analytical solution or a numerical method (such as gradient descent). For the weighted least squares method, the analytical solution is:

[0122] (10);

[0123] The sixth step is the process of iterative updating, which will continue until the change in the parameters is small enough, i.e.,

[0124] (11);

[0125] The above process is repeated for each position of the sliding window. Through the weighted residual and iterative optimization process, the Huber robust estimate can effectively reduce the influence of outliers and obtain a robust parameter estimate. In this section, the invention estimates the rough error problem of the processed MBES water depth data and intensity data. To this end, the characteristics extracted from the MBES water depth data and intensity data are fused using the robust estimation method, and the rough error is removed through the sliding window to ensure the purity of the extracted characteristics.

[0126] S4, in the model construction stage, the advantages of multiple classifiers are fused by using the Stacking mechanism in ensemble learning to perform decision level integration to generate a spatial distribution map of the coverage of the tubercle.

[0127] The training of the base classifier uses the k-fold cross-validation strategy to ensure the robustness of the prediction features (Meta-features). Each base classifier generates out-of-fold results through cross-prediction on the training set, and these results are used as new features input into the secondary classifier (Gradient Boosting) to realize the nonlinear fusion of the decision layer. The final output is a discrete coverage level (low, medium, high), and the spatial distribution of the deep-sea manganese nodule is presented in the form of a classification map.

[0128] Machine learning has shown strong applicability in seafloor classification and manganese nodule spatial distribution prediction, and can establish a complex nonlinear mapping relationship based on the input feature data. However, different machine learning algorithms may give different prediction results even for the same data set due to different modeling methods, assumptions and parameter selection. A single model is often limited by certain assumptions and may perform well in some areas but poorly in others. The multi-model method can help reduce prediction error and better generalize models for a wide range of geographic areas or more complex feature spaces, and improve the confidence of locations with good spatial consistency. Therefore, in order to improve the stability and generalization ability of the prediction, the invention introduces a multi-classifier decision fusion method based on the Stacking mechanism to better cope with the spatial variability of seafloor features in complex geographic areas.

[0129] Stacking (stacking generalization) is an efficient ensemble learning method, which trains a meta model by taking the prediction results of multiple base classifiers as input, further optimizing the final decision result. In the field of remote sensing, the ensemble learning method based on stacking mechanism has achieved satisfactory results. In the present invention, five classic supervised classification algorithms are selected as the first layer classifier: random forest (RF), decision tree (DT), BP network (BP), support vector machine (SVM) and K nearest neighbor classifier (KNN), which are independently trained and predicted. The prediction results of the first layer model are not only used to evaluate the individual performance of each model, but also as input features to the second layer meta model to improve the accuracy and stability of the overall classification. At the same time, the invention also takes into account the performance limit of the model on some data, and sets a performance screening threshold: if the prediction accuracy of a first layer classifier is less than 60%, the prediction result will not be passed to the second layer meta learning model, so as to avoid the interference of the wrong prediction result of the low performance model to the final decision.

[0130] The selection of the meta model determines the final prediction result, which has an important influence on resource evaluation. In the present invention, the gradient boosting classifier (GBC) is used as the final decision model. GBC can optimize the residual error through iteration, construct a more robust classification boundary, and further integrate the prediction information of the first layer classifier, thereby improving the overall performance of the model.

[0131] In the spatial prediction of the coverage of the tubercle, the first layer model of the stacking ensemble learning method classifies each pixel to obtain several independent prediction results. After the low-precision model is screened out by the performance screening mechanism, the remaining prediction results are combined to form a new feature set, which is input to the meta model for final decision. The decision process is based on the majority voting principle, which ensures that the final class determination can integrate the advantages of different classifiers and improve the spatial consistency.

[0132] Embodiment 2 provides a deep-sea manganese nodule coverage acoustic characterization system based on multi-classifier decision fusion.

[0133] The data acquisition module is used for acquiring water depth and backscattering intensity data based on a Kongsberg EM122 multi-beam sounding system, acquiring near-bottom images by a 6000-meter integrated optical towed body, and acquiring the positioning of the near-bottom optical towed body by using an ultra-short baseline (USBL) combined inertial navigation system (INS) and a Doppler velocity log (DVL) to obtain the positioning of the underwater optical sampling point, and acquiring a precise underwater topographic model by a shipborne multi-beam combined with a global navigation satellite system (GNSS) and an inertial navigation system (INS); the water depth and backscattering intensity data at the position of the sampling point are acquired by constructing the sampling point and the ship-measured terrain in the same projection manner.

[0134] The feature optimization module optimizes the extracted features by using a Boruta algorithm, and constructs a feature set for finely representing the seabed by extracting 14 backscattering texture features and 4 seabed topographic features.

[0135] The iterative robust estimation module introduces an iterative robust estimation algorithm based on a sliding window, dynamically adjusts the observation weight by continuously smoothing the data in different directions, and identifies and suppresses the gross errors in the feature image.

[0136] The multi-classifier decision fusion module adopts a Stacking mechanism in ensemble learning, integrates the advantages of multiple classifiers for decision-level integration, and generates a spatial distribution map of nodule coverage.

[0137] To further prove the positive effect of the above embodiment, the present application based on the above technical solution carries out the following experiment.

[0138] 1. Experimental data introduction: the area to be analyzed is located in the western part of the Clarion-Clipperton Zone (CCZ) in the eastern Pacific Ocean, about 1150 km away from Honolulu, Hawaii, USA. The CCZ is bounded by the Clipperton fault zone to the south, the Clarion fault zone to the north, the eastern Pacific rise to the east, and the Line Islands to the west. It is an intermediate block of the Pacific plate, with Mesozoic strata covering the oceanic basalt basement. The western part of the CCZ is mainly a deep-sea plain environment, which receives a large amount of deposition of biological remains from the ocean surface, terrestrial detritus, and cosmic dust. These sediments gradually accumulate over a long period of geological history, forming sedimentary strata of different thicknesses and types, covering the oceanic crust basement and fault structures.

[0139] The area to be analyzed is about 1181 km 2, the geological era of the basement formation is Late Cretaceous (about 95-65 Ma), the low positive gravity anomaly in the region shows high-density mantle anomaly and regional uplift of the seafloor, the water depth ranges from 4574 to 5451 m. In addition to the multi-beam data, in-situ observation data of the seafloor are also collected. A 6000-meter integrated optical tow body LH-GT6000G is used to collect a deep-sea optical tow line, and the seafloor surface matrix is collected by underwater high-definition video. The original optical data collected are screened, and the images with poor quality are removed. Through analysis of the screened images, the coverage index of manganese nodule is obtained. The optical tow body uses global positioning navigation system (GPS) combined with ultra-short baseline acoustic navigation (USBL) for navigation and positioning, each system has its own limitations, and these limitations have a cumulative effect on the total navigation error, and the position offset is a small error that is difficult to quantify. In addition to the optical data collected by the optical tow body, box samplers are used, which are dragged by the geological cable, penetrate the seafloor by gravity, release the bottom shovel at the moment of touching the bottom, and the sampler performs sampling and records the bottom matrix on site. A total of 16 box sampling points and 4753 optical tow line sampling points are collected.

[0140] 2. Experimental results and analysis;

[0141] 2.1 Feature extraction and preferred result evaluation;

[0142] 2.1.1 Feature extraction and preferred results: Based on multi-beam sounding and backscatter intensity data, the present application extracts 14 backscatter intensity features and 4 terrain features with the same resolution (i.e. 150 m resolution). Table 1 shows all feature types and abbreviations, and the calculation method of each feature is presented in mathematical expressions. In the process of Boruta feature selection, the importance score of each feature is calculated, and the specific score is visualized in the form of a column chart as shown in Figure 3 .

[0143] The Boruta algorithm runs under the condition of maxRuns=500 iterations and p value of 0.05. According to Boruta analysis, 16 of the 18 initial prediction variables included in the model are considered important, none of the prediction variables is considered a weak feature, and the features contrast and variance are determined as invalid features. The present application excludes the two features contrast and variance, which not only ensures high model performance but also reduces the number of prediction variables as much as possible.

[0144] 2.1.2 Preferred feature evaluation: Figure 4For the correlation analysis between different features, it can be seen that irrelevant features have been screened out, and all the remaining features have certain correlation with the target variable or each other.

[0145] The analysis shows that the dominance of the topographic skeleton features is embodied here. Water depth plays an important role in the index as the primary predictor, while topographic curvature and slope direction are the second and third important indicators in the evaluation index. They control the local hydrodynamic strength and form sediment trapping "traps" in the micro-topographic uplift area (positive curvature), which explains the strong spatial coupling phenomenon between the nodule enrichment zone and the seamount slope. The above topographic skeleton features again confirm that the seafloor topography controls the nodule distribution pattern through the topography-hydrodynamic coupling mechanism, supporting the "topographic pump" metallogenic hypothesis. Unlike the traditional bottom classification method, the importance of backscattering intensity is not as large as that in the traditional bottom classification. The invention believes that the reason is that the distribution of nodules on the seafloor shows the characteristics of being buried or semi-buried by sediments, and backscattering intensity focuses more on reflecting the physical properties of the surface layer of sediments, thereby significantly increasing the difficulty of exploration.

[0146] Simply relying on the feature correlation relationship between features and categories for feature effectiveness evaluation often has strong subjectivity and limited discrimination ability. To further evaluate the change of model separability before and after feature selection, the Jeffries-Matusita (JM) distance index is used for comparative analysis of the separation degree between different categories. Therefore, the invention further introduces Jeffries-Matusita (JM) index as a separability measurement index to quantitatively evaluate the feature set before and after feature selection. JM distance is based on the assumption of normal distribution of data to obtain the separation degree of different categories, and is widely used in the field of pattern recognition and feature selection. It measures the separation degree between categories by calculating the Bhattacharyya distance between categories, and the value is normalized to 0 to , the value closer to , the stronger the separability between categories.

[0147] Table 2 JM distance matrix of different nodule coverage

[0148]

[0149] Table 3 JM distance matrix of different nodule coverage after feature selection

[0150]

[0151] Table 2 shows the JM distance matrix between different nodule coverage rates without feature selection. The results show that due to the presence of highly redundant or linearly correlated features in the original features, the covariance matrix is singular, resulting in ineffective consideration of the degree of separation between classes. In contrast, Table 3 shows the corresponding JM distance matrix after feature selection using the Boruta algorithm. The JM distance between all classes is greater than 1, and the JM distance between multiple class pairs is close to 2, indicating that the Boruta-selected feature set significantly improves the discriminability of the feature set and effectively enhances the separability between different classes. Table 3 shows the JM distance matrix between different nodule coverage rates after feature selection using the Boruta algorithm.

[0152] In summary, the Boruta feature selection algorithm can improve feature effectiveness, avoid redundant interference, and enhance the predictive discriminability of the model in the deep-sea nodule coverage prediction task, providing a solid feature foundation for subsequent manganese nodule coverage prediction.

[0153] 2.2 Nodule coverage prediction results

[0154] 2.2.1 Comparison of feature optimization effects from the perspective of features: According to the natural distribution characteristics of coverage, it is divided into three categories: low coverage (0-20%), medium coverage (20-40%), and high coverage (>40%). Subsequently, five mainstream classifier models are constructed, including random forest (RF), support vector machine (SVM), back propagation neural network (BP), K nearest neighbor (KNN), and decision tree (DT). From multiple dimensions such as overall accuracy (OA), Kappa coefficient, and user accuracy (UA) and producer accuracy (PA) of each class, the performance of different models in the coverage classification task is compared and analyzed, and the results are shown in Table 4.

[0155] Table 4 shows the prediction results of different classifiers for nodule coverage

[0156]

[0157] For coverage prediction, as shown in FIG. 5(a)-5(e), the RF model presents a relatively smooth and coherent spatial distribution, especially in the middle and eastern regions of the analysis area, with clear boundaries between high coverage (yellow) and low coverage (blue) areas. The spatial heterogeneity and the transition zone between coverage levels are effectively captured. The results obtained by DT have some similarity with RF in spatial distribution, but the texture is coarser. Although it can better distinguish high and low coverage areas, the class boundary is handled roughly, showing the characteristics of "hard division" of decision trees in the feature space, which is difficult to finely depict the classification boundary in complex environments. The BP neural network model retains the overall spatial structure while introducing more detailed information. It is similar to the spatial pattern of the RF model to some extent, but there is some overestimation in the medium coverage area (green). This model performs well in identifying local anomalies, but there is some risk of misclassification at the coverage boundary. The spatial distribution predicted by the SVM model and the KNN model is scattered, and the classification result presents a strong pixelization feature, with poor spatial continuity and large spatial noise. In the complex seafloor topography background, the generalization ability is weak, and the classification effect is not satisfactory.

[0158] To further improve the robustness and generalization ability of the nodule coverage classification model, the Boruta feature selection method is introduced to evaluate and simplify the original features. After completing the Boruta feature selection, the classification precision comparison experiment of manganese nodule coverage is carried out to evaluate the influence of feature optimization on model performance. By comparing the prediction performance of five mainstream classifiers before feature selection (see Table 4) and after feature selection (see Table 5), it can be clearly observed that Boruta selection has a positive effect in multiple dimensions.

[0159] Overall, as shown in Table 5, Boruta feature selection significantly improves the overall accuracy (OA) and Kappa coefficient of most classifiers. The Kappa coefficient of the RF model increases from 0.8808 to 0.9014, and the overall accuracy increases from 92.16% to 93.54%, further consolidating its dominant position in spatial prediction. The model improves the UA and PA indicators of the three coverage classes, indicating that it has stronger balance and stability in various discrimination tasks. The classification results show that the Kappa coefficients of all classifiers are improved after feature selection, indicating that the stability and discrimination ability of the model are enhanced.

[0160] At the same time, except for the BP neural network, the overall classification accuracy of the remaining classifiers also improves to varying degrees; although the accuracy of the BP neural network does not change significantly, its performance does not decrease significantly, indicating that Boruta selection has a positive effect on the overall robustness of the model.

[0161] Table 5 Prediction results of different classifiers on nodule coverage after Boruta feature selection

[0162]

[0163] 2.2.2 Comparison of accuracy from the perspective of feature error interference: In this experiment, the application of robust estimation theory in the prediction scenario of deep-sea nodule resource evaluation indicators will be focused on, especially its effectiveness in data gross error rejection. As mentioned earlier, the M-estimation method in robust estimation is adopted in the present application, and the Huber function is selected as the p function for gross error rejection. In specific implementation, the present application uses a sliding window with a size of 7 for robust estimation, and shows the effect of nodule coverage prediction under this method, as shown in Table 6.

[0164] Table 6 Prediction results of different classifiers on nodule coverage after robust estimation

[0165]

[0166] Figures 6(a)-6(e) show the spatial distribution of the classification prediction of manganese nodule coverage using five typical classification algorithms after robust estimation processing of the original feature set. Compared with the previous classification results based on the un-robustly estimated feature set, it can be found that after robust estimation processing, the recognition ability of each model for different coverage areas has been improved to some extent, especially in the spatial continuity and boundary recognition of the medium coverage (20-40%) area. Specifically, the distribution of each classification scheme in the medium coverage area (green) is more coherent, significantly better than the fragmented medium coverage distribution trend before processing.

[0167] At the same time, from the quantitative evaluation results, almost all classification models have obtained different degrees of accuracy improvement after introducing robust processing, which is 1.97% higher than random forest (RF), 2.85% higher than support vector machine (SVM), 1.51% higher than BP neural network, and 2.55% higher than decision tree (DT). Compared with the nodule coverage prediction results after feature selection using Boruta, robust estimation also brings obvious accuracy improvement, which is 1.86% higher than SVM, 1.55% higher than BP neural network, and 0.83% higher than DT.

[0168] The above experimental results show that the robust estimation enhances the anti-interference ability and generalization ability of the model at the feature level, and can effectively eliminate gross errors in the prediction process of deep-sea manganese nodule resource evaluation indicators. These gross errors are often represented as discrete abnormal points in the prediction map. Due to the influence of the complex deep-sea environment and the resolution limitation of the multi-beam system, such gross errors are difficult to completely avoid in actual detection. However, after the robustness processing by the M-estimation method, the discrete noise points in the nodule coverage rate prediction map are significantly reduced from the image effect; from the quantitative evaluation index, the prediction accuracy of almost all classifiers is also improved to different degrees, further verifying the effectiveness of the robust estimation in the deep-sea application scenario.

[0169] 2.2.3 Compare the effect of the proposed method from the perspective of classification method: The present application carries out spatial prediction analysis on the coverage rate of manganese nodule based on a variety of mainstream classifiers. The experimental results show that different classifiers have significant differences in classification performance. For example, random forest (RF) and decision tree (DT) show high prediction accuracy and good classification ability. However, some models such as support vector machine (SVM) perform relatively poorly in the current task. In addition, it is observed that some classifiers have inconsistent prediction consistency and accuracy: for example, K nearest neighbor (KNN) performs well in total accuracy, but its classification image shows great fragmentation; while the BP neural network has slightly lower overall accuracy, but its prediction result has high consistency with most other models, showing a more stable spatial pattern. In order to overcome the limitations of a single classifier and enhance the adaptability of the model to complex data, on the basis of completing the coverage rate prediction of a single classification model, the Stacking ensemble learning method is further introduced in order to integrate the advantages of multiple models and improve the generalization ability and robustness of the overall model.

[0170] In the prediction task of nodule coverage rate, the Stacking ensemble learning method significantly improves the classification performance. Specifically, as shown in Table 2, the overall accuracy of the Stacking model is 0.87, which is 0.05 higher than that of the best single classifier (RF), and the Kappa coefficient is 0.82, which is 0.03 higher than that of the best single classifier (RF). In addition, the Stacking model also has the highest F1 score and the second highest precision among all classifiers, which further verifies the effectiveness of the Stacking method in the prediction task of nodule coverage rate. Figure 7As shown, the low coverage area (blue) in the prediction map of the Stacking ensemble learning method shows higher internal consistency, smoother distribution boundary, and significantly reduced abnormal discrete points, showing strong anti-noise robustness; the medium coverage area (green) has more structural distribution, and the spatial aggregation feature is obviously enhanced, which better captures the transition zone characteristics of the manganese nodule coverage; while the high coverage area (yellow) has a relatively clear boundary and better spatial coherence than the previous two methods. As shown in Table 7, the total accuracy of the prediction results based on the ensemble model reached 93.21%, and the Kappa coefficient was 89.62%. Compared with the prediction using all features, the Stacking method improved the prediction accuracy by 2.35% (compared with RF), 39.06% (compared with SVM), 14.00% (compared with BP neural network), 7.33% (compared with DT) and 3.69% (compared with KNN), respectively. Compared with the single classifier model based on Boruta feature selection and robust estimation optimization, the Stacking method improved the prediction accuracy by 0.38% (compared with RF), 36.21% (compared with SVM), 12.49% (compared with BP neural network) and 4.32% (compared with KNN), respectively, further verifying the effectiveness and applicability of the ensemble strategy in the prediction of manganese nodule distribution in deep-sea complex environment.

[0171] Table 7 Confusion matrix of nodule coverage using Stacking mechanism

[0172]

[0173] The above describes only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any modification, equivalent replacement and improvement made by those skilled in the art within the technical range disclosed by the present application, as long as it is within the spirit and principle of the present application, should be covered within the protection scope of the present application.

Claims

1. A method for acoustic characterization of deep-sea manganese nodule coverage based on multi-classifier decision fusion, characterized in that, The method comprises the following steps: S1, based on the Kongsberg EM122 multi-beam sounding system, the water depth and backscattering intensity data are obtained, and the 6000-meter integrated optical tow body is combined to obtain the near-bottom image, the positioning of the near-bottom optical tow body is obtained by using the ultra-short baseline USBL combined inertial navigation system INS and the Doppler velocity recorder DVL to obtain the positioning of the underwater optical sampling point, and the shipborne multi-beam obtains the precise underwater topographic model by combining the global navigation satellite system GNSS combined inertial navigation system INS; the water depth and backscattering intensity data at the position of the sampling point are obtained by constructing the sampling point and the ship-borne topography in the same projection manner; S2, by extracting backscattering texture features and seabed topographic features, the Boruta algorithm is used to optimize the extracted features, and a feature set finely representing the seabed is constructed; S3, an iterative robust estimation algorithm based on a sliding window is introduced, the observation weight is dynamically adjusted by continuously smoothing the data in different directions, and the gross error in the feature image is identified and suppressed; S4, in the model construction stage, the advantages of multiple classifiers are fused by using the Stacking mechanism in ensemble learning to perform decision level integration, and a spatial distribution map of the nodule coverage rate is generated; In step S3, the iterative robust estimation algorithm based on the sliding window comprises: (1) sliding on the feature image with a specified size window; (2) in each window, local pixel sequences are extracted along the horizontal, vertical and diagonal directions respectively; by constructing one-dimensional pixel chains in different directions, the spatial correlation of the acoustic scattering signal in the nodule coverage area is utilized to enhance the sensitivity of the feature extraction process to spatial texture and nodule edge structure; (3) based on the Huber loss function, weights are assigned to the observation points in each direction pixel sequence, and the iterative weighted least squares method is used to update the local regression model parameters of the direction until convergence, and the robust estimation result of the direction is obtained; the sequence data of each direction is subjected to robust regression, the Huber function is used to reduce the influence of gross errors caused by noise pulses, deep water bubbles or local abnormal reflections, and the estimation result of the nodule coverage rate related feature is more stable; (4) the direction weight is dynamically calculated according to the robust standard deviation of the residual error of each direction, and the estimation results of different directions are weighted and fused to obtain the final robust estimation value of the window center pixel; (5) move the window to the next position, repeat steps (1)-(4) until the entire feature image is traversed; through the multi-direction weighted residual error and the iterative optimization process, the robust parameter estimation which is not sensitive to noise and gross error is obtained.

2. The deep-sea nodule coverage acoustic characterization method based on multi- classifier decision fusion according to claim 1, characterized in that, In step S2, backscattering texture features include water depth, slope, curvature, aspect, backscattering intensity, GLCM 均值 , GLCM 方差 , GLCM 角二阶矩 , GLCM 熵 , GLCM 对比度 , GLCM 同质性 , GLCM 异质性 , GLCM 相关性 , standard deviation, kurtosis, skewness, energy, entropy; The seabed topographic features include depth, slope, slope direction and plane curvature.

3. The method for deep-sea manganese nodule coverage acoustic characterization based on multi-classifier decision fusion according to claim 1, characterized in that, In step S2, the Boruta algorithm is used to optimize the extracted features, which comprises: the Boruta algorithm expands the feature space by using the method of supplementing random attributes, and eliminates the correspondence between the sample attribute value and the label; classification test is carried out by using the expanded feature space, and the importance index of all attributes is calculated; the importance degree of each attribute is evaluated by taking the importance index of the shadow attribute as a reference; The importance of each feature variable is evaluated using the Boruta algorithm, including: (1) Initialization and data preparation: The dataset is x = {x1, x2,..., x p}, where xi is the i-th feature, and p is the total number of features; the dataset also contains a target variable y, which represents the label of each sample; i ​ (2) Image feature generation: The Boruta algorithm generates a set of random image features to simulate noise, thereby determining whether the real features are important; for each original feature x i Generate an image feature x' i Image features x' i The value is obtained by randomly shuffling x. i The sample labels in the data are obtained and represent only noise; (3) Merge original features and image features: merge original features X = {x1, x2, …, x p} and image features X' = {x'1, x'2, …, x' p} to form a dataset X combined containing original features and image features; use the merged feature set X combined to train a random forest model to estimate the importance of each feature; the random forest model outputs an importance score for each feature, and the importance of original features and image features is compared.

4. The deep-sea nodule coverage acoustic characterization method based on multi-classifier decision fusion according to claim 3, characterized in that, The Boruta algorithm uses the Gini index to measure the importance of features; For each feature x i , a random forest model is trained and the importance is computed, determining the contribution to the prediction of the target variable y; The random forest model evaluates the importance of features based on the contribution of each feature.

5. The deep-sea nodule coverage acoustic characterization method based on multi-classifier decision fusion according to claim 3, characterized in that, The random forest model compares the importance of original features and image features, including: calculating the maximum importance I of all image features max : For each original feature x i , compare its importance I(x i ) with I max ; if I(x i )>I max , then the feature is important; otherwise, the feature is useless; the Boruta algorithm runs multiple times by constantly iterating the image feature importance comparison step until the classification of each feature is stable.

6. The method for deep-sea manganese nodule coverage acoustic characterization based on multi-classifier decision fusion according to claim 1, characterized in that, The sliding window-based iterative robust estimation algorithm uses robust estimation Huber to fit a local model in each window, predicts the estimated value of the center point of the window, and replaces the center pixel. The specific steps are as follows: (1) For pixel p = (i, j), define a local sliding window with a fixed size of 5x5 and a step size of 1 pixel. Process each pixel to ensure that the boundary prediction is centralized and complete; (2) Extract the 1D pixel chain where the center pixel is located, use Huber function as the kernel function to do robust regression along the four directions Θ = {0°, 90°, 45°, 135°} for each chain, get the center estimation of this direction and the residual scale of this direction (3) Take a weighted average of the four estimated values obtained in the above directions as the final output; the expression is: where α θ (p) is the weight coefficient at position p in direction θ, ε is a positive number to prevent the denominator from being zero; is the final estimate at position p, is the normalized weight to ensure that the sum of the weights of all directions is 1.

7. The method for deep-sea manganese nodule coverage acoustic characterization based on multi-classifier decision fusion according to claim 1, characterized in that, In step S4, the multi-classifier decision fusion based on the Stacking mechanism includes: Select five supervised classification algorithms as the first layer of classifiers, independently train and predict the data; the prediction results of the first layer model are used to evaluate the individual performance of each model and are passed as input features to the second layer model; set a performance screening threshold, if the prediction accuracy of a first layer classifier is lower than the threshold, do not pass the prediction result to the second layer meta-learning model; use gradient boosting classifier as the final decision model to integrate the prediction information of the first layer classifier.

8. The deep-sea nodule coverage acoustic characterization method based on multi-classifier decision fusion according to claim 6, characterized in that, In the spatial prediction of tuberculosis coverage, the first layer model of the Stacking ensemble learning method classifies each pixel to obtain several independent prediction results; After excluding low-precision models through the performance screening mechanism, the remaining prediction results are combined to form a new feature set, which is input to the meta-model for final decision; The decision-making process is based on the majority voting principle to ensure that the final classification integrates the advantages of different classifiers.

9. The deep-sea nodule coverage acoustic characterization method based on multi- classifier decision fusion according to claim 6, characterized in that, The supervised classification algorithms include random forest, decision tree, BP network, support vector machine, and K nearest neighbor classifier.

Citation Information

Patent Citations

  • Substrate classification method and system based on deep sea multi-beam water body bottom echo information

    CN118277842A

  • Multi-beam seabed sediment classification method based on multi-classifier decision fusion mechanism

    CN119150234A