Method for evaluating stability of seabed sediment sound speed prediction model based on performance attenuation rate and related equipment
Patent Information
- Application Number
- CN202610762746.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-28
AI Technical Summary
在实际海洋工程与科学调查中,沉积物声速通常难以在原位直接大规模测量,主要依赖取样后的实验室测定,耗时耗力,数据获取成本较高
[0015] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method and related equipment for evaluating the stability of a seabed sediment acoustic velocity prediction model based on performance degradation rate. This scheme obtains a sediment sample dataset to provide a data foundation for subsequent steps; each sediment sample in the sediment sample dataset includes several parameter features and a corresponding target value for P-wave acoustic velocity; based on all parameter features, a first evaluation is performed on the preset acoustic velocity prediction model to obtain the baseline root mean square error of the acoustic velocity prediction model, providing a comparable benchmark for subsequent quantification of performance degradation caused by missing features; according to a preset continuous missing rate gradient, a reduction operation is performed on individual parameter features to obtain the missing features at the current missing rate; through a progressive feature information reduction process, the structural stability of the internal feature dependency ordering of the model is tested under different feature information compression intensities, thereby eliminating the random deviations that may be caused by a single extreme missing case; based on the remaining parameter features and the missing features at the current missing rate, a second evaluation is performed on the acoustic velocity prediction model to obtain the average root mean square error of the acoustic velocity prediction model; based on the baseline root mean square error and the average root mean square error, the missing features are obtained. The first performance decay rate of the sound velocity prediction model under the current missing rate is used to transform the absolute error into a dimensionless relative decay ratio, eliminating the differences in dimensions and numerical scales between different models and enabling cross-model feature dependency comparison. Based on the first performance decay rate, the second performance decay rate of the sound velocity prediction model under extreme missing rates is extracted for each missing feature. The missing features are then ranked according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence, quantifying the actual damage to model performance when each feature is completely missing. Based on the feature dependency ranking sequence, a feature dependency structure map is constructed, thus transforming the abstract numerical ranking into a visual structural expression. Based on a preset global missing rate, all parameter features are randomly reduced to obtain the third performance decay rate of the sound velocity prediction model under the preset global missing rate, evaluating the overall stability of the model under multiple simultaneous missing features. Based on the feature dependency structure map and the third performance decay rate, the stability evaluation results of the sound velocity prediction model are obtained. By comprehensively considering the model selection criteria from the dual dimensions of performance decay rate and root mean square error, the robustness and reliability of the sound velocity prediction model under data missing conditions can be improved.
Smart Images

Figure CN122654584A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a method and related equipment for stability evaluation of a seabed sediment sound velocity prediction model based on performance decay rate. Background Technology
[0002] The velocity of sound in seabed sediments is a crucial fundamental parameter for ocean acoustic field modeling, sonar performance prediction, underwater target detection, and ground acoustic inversion. Accurately obtaining the spatial distribution of sediment sound velocity is of great significance for understanding the propagation patterns of sound waves in the ocean and improving the accuracy of underwater detection. In practical marine engineering and scientific surveys, sediment sound velocity is usually difficult to measure directly on a large scale in situ, relying mainly on laboratory measurements after sampling, which is time-consuming, labor-intensive, and costly. Furthermore, in actual marine survey operations, instrument malfunctions, sample damage, or incomplete historical data frequently occur. Current technologies cannot quantitatively assess the degree of performance degradation caused by missing features, lacking a reliable basis for engineering applications.
[0003] Furthermore, in the evaluation and application of seabed sediment sound velocity prediction models, existing indicators such as Gini Importance and Permutation Importance are insufficient to directly reflect the true impact of features on model performance in real-world missing feature imputation scenarios. They also cannot be compared across different machine learning models, making it difficult to provide quantitative support for prioritizing feature acquisition in practical engineering. Moreover, the internal mechanisms of different types of machine learning models, such as linear models, tree models, and kernel methods, differ significantly. Existing technologies lack a unified quantitative standard to compare the sensitivity of different models to feature missing features, making it difficult to provide quantitative basis for model selection in scenarios with data incompleteness risks. Summary of the Invention
[0004] In view of this, the main objective of the embodiments of the present invention is to provide a method and related equipment for stability evaluation of a sound velocity prediction model for seabed sediments based on performance decay rate, in order to solve at least one of the problems in the prior art. The present invention can improve the robustness and reliability of the sound velocity prediction model.
[0005] To achieve the above objectives, one aspect of the present invention provides a method for stability evaluation of a seabed sediment sound velocity prediction model based on performance degradation rate, the method comprising: Obtain a sediment sample dataset; each sediment sample in the sediment sample dataset includes several parametric features and a corresponding target value for P-wave sound velocity; Based on all the aforementioned parameter characteristics, a first evaluation is performed on the preset sound speed prediction model to obtain the benchmark root mean square error of the sound speed prediction model. Based on the preset continuous missing rate gradient, a reduction operation is performed on a single parameter feature to obtain the missing features at the current missing rate; Based on the remaining parameter features and the missing features under the current missing rate, the sound speed prediction model is evaluated a second time to obtain the average root mean square error of the sound speed prediction model. Based on the benchmark root mean square error and the average root mean square error, obtain the first performance degradation rate of the sound speed prediction model under the current missing feature rate. Based on the first performance decay rate, the second performance decay rate of the sound speed prediction model for each missing feature under extreme missing rate is extracted, and the missing features are sorted according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence. Construct a feature dependency structure graph based on the sorted sequence of feature dependencies; Based on a preset global missing rate, all the parameter features are randomly reduced to obtain the third performance decay rate of the sound speed prediction model under the preset global missing rate. Based on the feature-dependent structure map and the third performance attenuation rate, the stability evaluation result of the sound speed prediction model is obtained.
[0006] In some embodiments, the first evaluation of the preset sound velocity prediction model based on all the parameter characteristics to obtain the baseline root mean square error of the sound velocity prediction model includes the following steps: The sediment sample dataset is divided into a training set and a test set; Based on the training set, the pre-trained model is optimized and trained to obtain the sound speed prediction model; Input all the parameter features in the test set into the sound velocity prediction model to obtain the first longitudinal wave sound velocity prediction value; The root mean square error of the sound velocity prediction model is obtained based on the target value of the longitudinal wave sound velocity and the predicted value of the first longitudinal wave sound velocity.
[0007] In some embodiments, the step of performing a reduction operation on a single parameter feature according to a preset continuous missing rate gradient to obtain the missing feature at the current missing rate includes the following steps: The sediment sample dataset is divided into a training set and a test set; Based on the training set, obtain the median of a single parameter feature; A sequential missing rate gradient from 0% to 100% is preset, and the current missing rate is determined based on the sequential missing rate gradient; Preprocessed samples are obtained by randomly selecting sediment samples from the test set that represent the current missing rate. The value of a single parameter feature of the preprocessed sample is replaced with the median to obtain the missing feature under the current missing rate.
[0008] In some embodiments, the second evaluation of the sound speed prediction model based on the remaining parameter features and the missing features at the current missing rate, to obtain the average root mean square error of the sound speed prediction model, includes the following steps: The parameter features and the missing features under the current missing rate are input into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value. Based on the target value of the longitudinal wave sound velocity and the predicted value of the second longitudinal wave sound velocity, the initial root mean square error of the sound velocity prediction model is obtained. Return to the step of inputting the parameter features and the missing features under the current missing rate into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value, until the preset number of repeated executions is reached to obtain several initial root mean square errors; The average root mean square error of the sound speed prediction model is obtained based on several initial root mean square errors.
[0009] In some embodiments, obtaining the first performance degradation rate of the sound speed prediction model under the current missing feature rate based on the benchmark root mean square error and the average root mean square error includes the following steps: Obtain the evaluation difference between the average root mean square error and the baseline root mean square error; The first performance degradation rate is obtained by dividing the evaluation difference by the root mean square error of the benchmark.
[0010] In some embodiments, the step of randomly reducing all the parameter features according to a preset global missing rate to obtain the third performance degradation rate of the sound speed prediction model under the preset global missing rate includes the following steps: The sediment sample dataset is divided into a training set and a test set; Based on the training set, obtain the median of each parameter feature; Perform an independent and identically distributed Bernoulli mask on all the parameter features in the test set, and randomly generate missing positions according to the preset global missing rate; The median of the corresponding parameter feature is filled into the missing position to generate a global perturbation test set; By performing several Monte Carlo simulations, the global disturbance test set is input into the sound speed prediction model to obtain the third performance degradation rate.
[0011] To achieve the above objectives, another aspect of this invention proposes a stability evaluation device for a seabed sediment sound velocity prediction model based on performance degradation rate, the device comprising: The data acquisition module is used to acquire a sediment sample dataset; each sediment sample in the sediment sample dataset includes several parametric features and a corresponding target value for P-wave sound velocity; The benchmark root mean square error acquisition module is used to perform a first evaluation on the preset sound speed prediction model based on all the parameter features, and obtain the benchmark root mean square error of the sound speed prediction model. The single-feature failure module is used to perform a reduction operation on a single parameter feature according to a preset continuous missing rate gradient to obtain the missing feature under the current missing rate. The average root mean square error acquisition module is used to perform a second evaluation on the sound speed prediction model based on the remaining parameter features and the missing features under the current missing rate, so as to obtain the average root mean square error of the sound speed prediction model. The performance decay rate calculation module is used to obtain the first performance decay rate of the sound speed prediction model under the current missing rate based on the benchmark root mean square error and the average root mean square error. The feature dependency ranking determination module is used to extract the second performance decay rate of the sound speed prediction model for each missing feature under extreme missing rate according to the first performance decay rate, and to rank the missing features according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence. The feature dependency structure graph generation module is used to construct a feature dependency structure graph based on the sorted sequence according to the feature dependency degree. The global feature failure module is used to randomly reduce all the parameter features according to a preset global missing rate, and obtain the third performance decay rate of the sound speed prediction model under the preset global missing rate. The evaluation result output module is used to obtain the stability evaluation result of the sound speed prediction model based on the feature-dependent structure map and the third performance decay rate.
[0012] To achieve the above objectives, another aspect of the present invention provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described above.
[0013] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0014] To achieve the above objectives, another aspect of the present invention provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions to cause the computer device to perform the aforementioned method.
[0015] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method and related equipment for evaluating the stability of a seabed sediment acoustic velocity prediction model based on performance degradation rate. This scheme obtains a sediment sample dataset to provide a data foundation for subsequent steps; each sediment sample in the sediment sample dataset includes several parameter features and a corresponding target value for P-wave acoustic velocity; based on all parameter features, a first evaluation is performed on the preset acoustic velocity prediction model to obtain the baseline root mean square error of the acoustic velocity prediction model, providing a comparable benchmark for subsequent quantification of performance degradation caused by missing features; according to a preset continuous missing rate gradient, a reduction operation is performed on individual parameter features to obtain the missing features at the current missing rate; through a progressive feature information reduction process, the structural stability of the internal feature dependency ordering of the model is tested under different feature information compression intensities, thereby eliminating the random deviations that may be caused by a single extreme missing case; based on the remaining parameter features and the missing features at the current missing rate, a second evaluation is performed on the acoustic velocity prediction model to obtain the average root mean square error of the acoustic velocity prediction model; based on the baseline root mean square error and the average root mean square error, the missing features are obtained. The first performance decay rate of the sound velocity prediction model under the current missing rate is used to transform the absolute error into a dimensionless relative decay ratio, eliminating the differences in dimensions and numerical scales between different models and enabling cross-model feature dependency comparison. Based on the first performance decay rate, the second performance decay rate of the sound velocity prediction model under extreme missing rates is extracted for each missing feature. The missing features are then ranked according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence, quantifying the actual damage to model performance when each feature is completely missing. Based on the feature dependency ranking sequence, a feature dependency structure map is constructed, thus transforming the abstract numerical ranking into a visual structural expression. Based on a preset global missing rate, all parameter features are randomly reduced to obtain the third performance decay rate of the sound velocity prediction model under the preset global missing rate, evaluating the overall stability of the model under multiple simultaneous missing features. Based on the feature dependency structure map and the third performance decay rate, the stability evaluation results of the sound velocity prediction model are obtained. By comprehensively considering the model selection criteria from the dual dimensions of performance decay rate and root mean square error, the robustness and reliability of the sound velocity prediction model under data missing conditions can be improved. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the stability evaluation method for a seabed sediment sound velocity prediction model based on performance decay rate provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall steps for stability assessment of the seabed sediment sound velocity prediction model based on performance decay rate provided in an embodiment of the present invention; Figure 3 This is a comparison diagram of the PDR structure of four types of models under the condition of complete failure of a single feature provided in the embodiments of the present invention; Figure 4 This is a comparison chart of model performance degradation and PDR ranking under a 50% random missing condition provided in the embodiments of the present invention; Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0019] It should be noted that although functional modules are divided in the system diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims, and the foregoing drawings may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of the embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to a determination," or "in the event of a determination."
[0020] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0022] Before providing a detailed description of the embodiments of the present invention, some of the nouns and terms involved in the embodiments of the present invention will be explained first. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations.
[0023] Performance Degradation Rate (PDR): This is the core quantitative metric proposed in this invention, defined as the normalized increment of the model's root mean square error (RMSE) relative to the baseline RMSE with complete features under feature loss conditions. PDR is a dimensionless metric, eliminating the dimensional differences in the accuracy of different model baselines and enabling a unified comparison of feature dependency strength across models. A larger PDR indicates a more severe impairment to model prediction performance due to the missing feature, and a higher degree of model dependence on that feature.
[0024] Seafloor sediment acoustic velocity prediction: This method uses physical properties (such as wet density and porosity) and grain size parameters (such as sand content, silt content, clay content, gravel content, average grain size, and median grain size) of sediments as input features, and employs a machine learning model to predict the longitudinal wave velocity (Vp) of sediments. Accurate prediction of acoustic velocity is of great significance for ocean acoustic field modeling, sonar performance prediction, and ground acoustic inversion.
[0025] Feature failure experiment: A systematic evaluation method that involves masking each input feature one by one (replacing it with statistical values from the training set) and observing the degree of degradation in the model's predictive performance under different missing rate gradients, thereby quantifying the independent contribution and dependence of each feature on the model.
[0026] Bernoulli Mask: A random missing feature simulation mechanism that independently and randomly sets each element in the feature matrix to be missing with a given probability, used to simulate the overall scenario of multiple features being missing simultaneously in actual marine surveys.
[0027] Monte Carlo repeated experiments: By conducting multiple independent random sampling experiments, the mean and standard deviation are used as performance estimates to eliminate the random bias caused by a single random sampling and improve the statistical reliability of the evaluation results.
[0028] Among related technologies, there are sound velocity prediction methods with the core objective of improving accuracy and neural network prediction methods with physical information fusion as their core. Methods that evaluate and validate model accuracy by optimizing model structure and feature combinations under conditions of complete data and co-domain distribution do not address the robustness of models in scenarios with missing data or feature failure. Methods that improve the physical consistency of models by embedding physical constraints into neural network structures also only optimize the prediction structure of the model itself, failing to consider common data incompleteness scenarios in actual ocean surveys, such as single feature failure, random missing multiple features, and instrument malfunctions, and also do not provide quantitative analysis methods for feature dependencies.
[0029] In terms of feature importance evaluation, existing methods mainly fall into two categories: Gini importance and permutation importance. Gini importance reflects the average contribution of a feature to node splitting on the complete dataset, but it cannot quantify the performance degradation when a feature is completely missing. Permutation importance evaluates feature values by randomly shuffling them, and its perturbation method differs from the actual impact of missing features on the data distribution. Furthermore, both methods lack cross-model comparability, making it difficult to uniformly compare the degree of feature dependence under different algorithmic structures.
[0030] In view of this, this invention provides a method and related equipment for stability assessment of a seabed sediment sound velocity prediction model based on performance degradation rate (PDR). The method first constructs a benchmark model including parameters such as physical properties and grain size, and establishes a benchmark root mean square error (RMSE). Then, by simulating a progressive loss of a single feature from 0% to 100%, imputation is performed using the number of bits in the training set, and the dimensionless PDR value causing model performance degradation is quantified. By extracting and sorting the PDR values under extreme loss conditions, a feature dependency structure map is generated, providing a clear data acquisition priority for surveying work, thereby reducing the acquisition cost and resource waste of unnecessary parameters. Furthermore, by combining stability assessment with global random loss, the model structure with the highest tolerance for feature information deficiencies can be selected, significantly improving the robustness and engineering applicability of the sound velocity prediction model under complex real-world conditions.
[0031] The stability evaluation method for a seabed sediment acoustic velocity prediction model based on performance decay rate provided in this invention relates to the fields of marine acoustic detection, seabed sediment parameter inversion, and machine learning. This method can be applied to terminals, servers, or software running on either. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the stability evaluation method for the seabed sediment acoustic velocity prediction model based on performance decay rate, but is not limited to these forms.
[0032] Figure 1 This is an optional flowchart of a stability evaluation method for a seabed sediment sound velocity prediction model based on performance decay rate, provided in an embodiment of the present invention. Figure 1 The method may include, but is not limited to, steps S100 to S900: Step S100: Obtain a sediment sample dataset; each sediment sample in the sediment sample dataset includes several parametric features and the corresponding target value of P-wave sound velocity; Step S200: Based on all parameter characteristics, perform a first evaluation on the preset sound velocity prediction model to obtain the benchmark root mean square error of the sound velocity prediction model. Step S300: Based on the preset continuous missing rate gradient, perform a reduction operation on the individual parameter features to obtain the missing features under the current missing rate. Step S400: Based on the remaining parameter features and the missing features under the current missing rate, perform a second evaluation of the sound velocity prediction model to obtain the average root mean square error of the sound velocity prediction model. Step S500: Based on the baseline root mean square error and the average root mean square error, obtain the first performance degradation rate of the sound speed prediction model under the current missing feature rate. Step S600: Based on the first performance decay rate, extract the second performance decay rate of the sound speed prediction model for each missing feature under extreme missing rate, and sort the missing features according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence. Step S700: Sort the sequences according to feature dependencies and construct a feature dependency structure map; Step S800: Based on the preset global missing rate, all parameter features are randomly reduced to obtain the third performance decay rate of the sound speed prediction model under the preset global missing rate. Step S900: Based on the feature-dependent structure map and the third performance attenuation rate, obtain the stability evaluation results of the sound speed prediction model.
[0033] In step S100 of some embodiments, a sediment sample dataset is obtained. Each sediment sample in the dataset contains several parametric features and corresponding target values for P-wave velocity. The parametric feature system is divided into three categories: the first category is basic physical property parameters, including wet density and porosity, which directly characterize the compactness of the sediment skeleton and are the strongest univariate predictors of sound velocity; the second category is grain size composition, including gravel content, sand content, silt content, and clay content, the sum of which is 100%, which together describe the grain size structure of the sediment; the third category is grain size statistics, including average grain size and median grain size, which provide a compressed grain size expression.
[0034] In step S200 of some embodiments, using the preprocessed sediment sample dataset, while maintaining feature integrity, hyperparameter optimization and training are performed on the training set for machine learning algorithms such as Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), and ElasticNet. The model is then run under ideal conditions without missing feature information to obtain a benchmark for model performance, namely the benchmark root mean square error. This metric serves as the normalized denominator for subsequent quantification of performance degradation caused by missing features, establishing a unified metric for horizontal comparison between different models and effectively eliminating evaluation bias caused by differences in algorithm benchmark accuracy. Optionally, auxiliary evaluation metrics such as the coefficient of determination (R²) and mean absolute error (MAE) can also be used to evaluate the model.
[0035] In some embodiments, step S200 may include, but is not limited to, steps S210 to S240: Step S210: Divide the sediment sample dataset into a training set and a test set; Step S220: Based on the training set, perform hyperparameter optimization and training on the pre-trained model to obtain the sound speed prediction model; Step S230: Input all parameter features in the test set into the sound velocity prediction model to obtain the first longitudinal wave sound velocity prediction value; Step S240: Based on the target value of the longitudinal wave sound velocity and the predicted value of the first longitudinal wave sound velocity, obtain the benchmark root mean square error of the sound velocity prediction model.
[0036] In step S210 of some embodiments, the dataset is divided into a training set and a test set according to a certain ratio. Sediment samples with missing values are directly removed to ensure data quality. For example, a random sampling or stratified sampling strategy is used to divide the entire sediment sample into a model training set and an independent test set according to a preset ratio. Data isolation is performed during the partitioning process to ensure that the test set data does not participate in model construction and parameter optimization throughout the process. This simulates prediction scenarios under real unknown environments, ensuring the objectivity and generalization of subsequent evaluation results. Simultaneously, abnormal sediment samples containing missing values are removed to maintain data quality.
[0037] In step S220 of some embodiments, machine learning algorithms such as Random Forest, XGBoost, Support Vector Regression, and ElasticNet are selected to construct pre-trained models. Hyperparameter optimization and training are performed on each pre-trained model on the training set to obtain various sound speed prediction models. For example, different combinations of hyperparameters (i.e., the configurable parameters of the machine learning algorithm itself) are traversed on the training set to find the optimal model structure through cross-validation. The selected optimal parameters are then fed back into a preset machine learning algorithm for full fitting, completing the training of the sound speed prediction model. This allows the sound speed prediction model to fully learn the nonlinear mapping relationship between the input parameter features and the target value of the P-wave sound speed, forming a fixed model with predictive capabilities.
[0038] In step S230 of some embodiments, all sediment samples and their complete multidimensional parameter features from the independent test set are retrieved and input in batches into the trained sound velocity prediction model. The sound velocity prediction model performs inference operations on the input data based on the learned mapping rules, and outputs the first P-wave sound velocity prediction value corresponding to the sediment sample, providing the necessary numerical basis for subsequent error quantification calculations.
[0039] In step S240 of some embodiments, the residual between the predicted first P-wave velocity and the actual target P-wave velocity is calculated. This residual is then substituted into the root mean square error (RMSE) formula to obtain the baseline RMSE of the velocity prediction model. The obtained baseline RMSE reflects the theoretical minimum error of the model under complete parameter characteristics and can be used as a reference for subsequent calculation of the performance degradation rate (PDR), quantifying the relative damage caused by feature loss.
[0040] In step S300 of some embodiments, for each input parameter feature, the entire process of the parameter from slight loss to complete loss is simulated. By setting a continuous gradient of the missing proportion, the corresponding parameter feature values in the test set samples are masked and replaced, constructing a perturbation dataset under different degrees of feature information loss. This allows for the accurate extraction and quantification of the independent contribution of a single feature to the model's predictive ability in controlled experiments, eliminating evaluation noise caused by multi-feature interference.
[0041] In some embodiments, step S300 may include, but is not limited to, steps S310 to S350: Step S310: Divide the sediment sample dataset into a training set and a test set; Step S320: Based on the training set, obtain the median of a single parameter feature; Step S330: Preset a continuous missing rate gradient from 0% to 100%, and determine the current missing rate based on the continuous missing rate gradient; Step S340: Randomly select sediment samples from the test set that account for the current missing rate to obtain preprocessed samples; Step S350: Replace the values of individual parameter features of the preprocessed sample with the median to obtain the missing features under the current missing rate.
[0042] In step S310 of some embodiments, the dataset is divided into a training set and a test set according to a certain ratio, and a direct removal strategy is adopted for sediment samples with missing values to ensure data quality.
[0043] In step S320 of some embodiments, the median of a single parameter feature is pre-calculated from the training set, providing a data basis for subsequent simulation of low-cost, low-precision conventional measurement filling methods.
[0044] In step S330 of some embodiments, for each evaluated parameter feature, a continuous missing rate gradient from 0% to 100% is set. The gradient step size can be flexibly adjusted according to actual calculation needs. Commonly used step sizes include 3%, 5%, and 10%. The current missing rate is selected in the continuous missing rate gradient according to the gradient step size.
[0045] In steps S340 to S350 of some embodiments, at each current missing rate level, a corresponding proportion of sediment samples are randomly selected from the test set as preprocessing samples. The values of the individual parameter features being evaluated in these preprocessing samples are replaced with the median pre-calculated from the training set. This simulates the conventional processing method in real-world applications where historical statistics are used instead of missing data measurements, thus obtaining the missing features at the current missing rate. The remaining parameter features remain unchanged from their original observations. The purpose of setting continuous missing gradients is not to characterize the precise functional relationship between the missing rate and prediction performance, but rather to examine whether the internal feature dependency ranking of each model remains structurally stable under different information compression intensities through a gradual information reduction process, thereby eliminating the random bias that may be caused by a single extreme missing case.
[0046] In step S400 of some embodiments, the remaining parameter features that keep the original observations unchanged, as well as the missing features under the current missing rate, are input into the sound speed prediction model. The performance of the model under different missing features and different missing rates is collected through multiple independent repeated experiments. Optionally, each experimental condition is independently repeated no less than 20 times, and the mean and standard deviation are taken as the performance estimate under that condition, so as to provide reliable data support for quantifying the impact of single feature missingness on the model.
[0047] In some embodiments, step S400 may include, but is not limited to, steps S410 to S440: Step S410: Input the parameter features and the missing features under the current missing rate into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value. Step S420: Obtain the initial root mean square error of the sound velocity prediction model based on the target value of the longitudinal wave sound velocity and the predicted value of the second longitudinal wave sound velocity. Step S430: Return to the step of inputting the parameter features and the missing features under the current missing rate into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value, until the preset number of repetitions is reached to obtain several initial root mean square errors. Step S440: Obtain the average root mean square error of the sound speed prediction model based on several initial root mean square errors.
[0048] In step S410 of some embodiments, parameter features and missing features under the current missing rate are batch input into the sound velocity prediction model. The model performs forward inference calculation based on the current feature distribution state and outputs the second P-wave sound velocity prediction value corresponding to the sediment sample. The second P-wave sound velocity prediction value reflects the model's ability to restore the P-wave sound velocity under a specific single parameter feature missing rate.
[0049] In step S420 of some embodiments, the root mean square of the sum of squared residuals between the target P-wave velocity and the predicted second P-wave velocity is calculated to obtain the initial root mean square error. This initial root mean square error quantifies the degree of direct deviation between the model prediction result and the actual geological conditions under the current specific missing rate and the current specific single missing feature, and serves as the error of a single experiment.
[0050] In step S430 of some embodiments, under the same missing rate and single missing feature conditions, the step of inputting the parameter features and the missing features under the current missing rate into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value is returned. Multiple sets of initial root mean square errors are accumulated through no less than 20 independent cyclic experiments.
[0051] In step S440 of some embodiments, the obtained initial root mean square errors are arithmetically averaged to calculate the average root mean square error of the sound speed prediction model. This average root mean square error eliminates noise interference from a single random experiment and can obtain the performance loss of the model due to the missing specific feature at the current missing rate, which can be used as an input variable for subsequent calculation of the performance degradation rate.
[0052] In step S500 of some embodiments, the root mean square error of the benchmark is used as a performance reference benchmark, and the average root mean square error is used as the performance observation value under the current feature missing. By constructing the calculation logic of the relative error ratio, the accuracy change of the model caused by the missing specific features is converted into a dimensionless percentage form, thereby eliminating the dimensional barrier caused by the difference in the benchmark accuracy of different sound speed prediction models, and obtaining the first performance decay rate.
[0053] In some embodiments, step S500 may include, but is not limited to, steps S510 to S520: Step S510: Obtain the evaluation difference between the average root mean square error and the baseline root mean square error; Step S520: Obtain the quotient of the evaluation difference divided by the root mean square error of the benchmark to obtain the first performance degradation rate.
[0054] In steps S510 to S520 of some embodiments, the first performance degradation rate of the sound velocity prediction model at the current missing rate is calculated based on the baseline root mean square error and the average root mean square error. The formula used includes: ; In the formula, Representing missing features At the current missing rate The first performance degradation rate of the lower sound velocity prediction model; Represents the current missing rate; This represents the average root mean square error. This represents the root mean square error of the reference.
[0055] The performance degradation rate (PDR) defined in this embodiment of the invention is a dimensionless index, which eliminates the dimensional differences in the baseline accuracy of different models and makes the feature dependency strength between different models comparable. The larger the PDR, the more severe the damage to the model's predictive performance caused by the lack of that feature, and the higher the model's dependence on that feature.
[0056] In step S600 of some embodiments, for each input parameter feature, its PDR value under the complete failure condition (missing rate of 100%) is extracted to obtain the second performance degradation rate. All parameter features are sorted from largest to smallest according to the second performance degradation rate to obtain the feature dependency ranking sequence of the sound velocity prediction model. This ranking directly quantifies the actual degree of damage to model performance when each feature is completely missing, and can provide a quantitative basis for planning data acquisition priorities in actual surveying.
[0057] The embodiments of the present invention also verify the structural stability of the feature dependency ranking in the range of 0% to 100% total missing rate, that is, the dependency ranking of each feature does not reverse with the change of missing rate, proving that the ranking conclusion is robust under different information reduction degrees.
[0058] In step S700 of some embodiments, a feature dependency structure map of the sound speed prediction model is generated based on the obtained feature dependency ranking sequence. Optionally, the map is presented in the form of a radar chart, bar chart, or polar coordinate graph, with the horizontal axis representing the feature name and the vertical axis representing the PDR value. The PDR ranking results of multiple different machine learning models can be overlaid and compared to generate a cross-model feature dependency structure comparison map, which intuitively presents the differences in the dependence strength of different algorithm structures on the same feature.
[0059] In step S800 of some embodiments, a complex working condition of simultaneous missing data from multiple sources in actual surveys is simulated, and random missing interference is applied to all input parameter features of the full sample of the test set; the relative performance degradation ratio of the model under the overall information missing environment is calculated by Monte Carlo simulation, that is, the third performance degradation rate. This index quantifies the overall sensitivity of the model to the concurrent missing of multiple features, complements the aforementioned single feature missing analysis, and together constitutes a complete dimension for the robustness evaluation of the model.
[0060] In some embodiments, step S800 may include, but is not limited to, steps S810 to S850: Step S810: Divide the sediment sample dataset into training and test sets; Step S820: Based on the training set, obtain the median of each parameter feature; Step S830: Perform an independent and identically distributed Bernoulli mask on all parameter features in the test set, and randomly generate missing positions according to the preset global missing rate; Step S840: Fill in the missing positions with the median of the corresponding parameter features to generate a global perturbation test set; In step S850, the global disturbance test set is input into the sound speed prediction model through several Monte Carlo simulations to obtain the third performance degradation rate.
[0061] In step S810 of some embodiments, the dataset is divided into a training set and a test set according to a certain ratio, and a direct removal strategy is adopted for sediment samples with missing values to ensure data quality.
[0062] In step S820 of some embodiments, the median of each parameter feature in the training set is obtained to provide a data basis for filling in missing values at any position when simulating global random missing values in the future.
[0063] In steps S830 to S840 of some embodiments, an independent and identically distributed Bernoulli mask with a preset missing rate is applied to the complete parametric feature matrix of the test set. That is, each element in the matrix is independently and randomly set to be missing with a given preset global missing rate. All missing values are imputed using the median of the corresponding parametric features in the training set, resulting in a globally perturbed test set. For example, a binary mask matrix with the same dimensions as the complete parametric feature matrix of the test set is constructed. Each element in the matrix is independently set to 1 (representing missing) with a probability of a preset global missing rate p, and to 0 (representing retention) with a probability of 1-p. All feature values in the test set marked as missing (mask value 1) are replaced with the median of the corresponding feature dimension. After the imputation operation, the original complete feature matrix is transformed into a globally perturbed test set containing a large number of median imputed values. This dataset completely preserves the original information of the non-missing positions, introducing statistical noise only at random positions.
[0064] In step S850 of some embodiments, the Monte Carlo method is used to perform multiple repeated simulations, generating different random masks each time to construct a new global perturbation test set. These datasets are then sequentially input into the sound speed prediction model for inference. The mean and standard deviation of the performance decay rate of multiple experiments are statistically analyzed to eliminate the random influence of a single random mask. Finally, the third performance decay rate of the sound speed prediction model is obtained, and the overall stability of each model for the simultaneous loss of multiple features is evaluated.
[0065] The horizontal comparison of PDR reflects the structural sensitivity of each model to the overall lack of information, rather than the absolute superiority or inferiority of the final prediction accuracy. PDR measures the relative degradation ratio, while RMSE measures the final predictive ability. The two reflect different dimensions, and both types of indicators should be considered together when selecting a model.
[0066] In step S900 of some embodiments, the stability evaluation result of the sound speed prediction model can be obtained by combining the feature dependency structure map and the third performance decay rate. The feature dependency structure map can generate a cross-model feature dependency structure comparison map, visually presenting the differences in the strength of dependency on the same feature by different algorithm structures. The mean and standard deviation of the third performance decay rate can serve as quantitative indicators of model fault tolerance.
[0067] In some embodiments, such as Figure 2 As shown, the stability assessment steps for the seabed sediment acoustic velocity prediction model based on performance degradation rate include: Step 1: Baseline Model Construction The dataset is divided into training and test sets. The training set data is used to drive the training of the sound speed prediction model, and the prediction error of the model is calculated on the test set to establish a baseline root mean square error, which serves as a reference for measuring the performance degradation of the model in subsequent feature missing experiments.
[0068] Step 2, Feature Failure Experiment: For each individual parameter feature being evaluated, simulate its missing state in a real-world environment. Randomly select samples from the test set that represent the current missing rate, and replace only the value of that individual parameter feature in these samples with the median of that feature in the training set, while keeping the values of the remaining parameter features unchanged, thus constructing a perturbed test set with missing features.
[0069] Step 3: Calculate PDR: The model prediction results under the feature missing state are compared with the baseline root mean square error. The error difference between the two is calculated and divided by the baseline root mean square error to obtain the first performance decay rate, which quantifies the impact of a single feature missing on the model.
[0070] Step 4, Feature Dependency Ranking: The experimental results for missing parameters and features are summarized and sorted in descending order based on the first performance decay rate. Features with higher decay rates indicate stronger model dependence on them. Based on this, a feature dependency ranking sequence is generated to determine the core parameters.
[0071] Step 5: Generate a dependency graph: Based on the feature dependency ranking sequence, visualization methods such as bar charts or radar charts are used to transform the abstract numerical ranking into an intuitive graphical representation, and to draw the feature dependency structure map of a single model or multiple models superimposed.
[0072] After generating the dependency graph, determine whether it is necessary to evaluate the model's overall robustness to concurrent feature loss. If evaluation is selected, proceed to the overall loss experiment; otherwise, output the final stability evaluation report directly based on the previous single-feature analysis results.
[0073] Step 6: Perform the global missing value experiment: All parameter features are simultaneously subjected to random masking on the test set. After median imputation to generate a global perturbation test set, the overall performance degradation rate of the model is calculated through multiple Monte Carlo simulations.
[0074] Step 7: Output the evaluation results: By integrating the single-feature dependency ranking sequence with the overall random missing data assessment results, the robustness of the model under different data missing conditions is comprehensively analyzed, and the stability assessment results of the sound speed prediction model are finally output.
[0075] In some embodiments, taking comprehensive sediment survey data of a certain sea area as an example, several sets of valid samples are collected after missing value removal and outlier cleaning, forming the data basis of this embodiment. The input parameter feature system contains 8 parameters, covering physical properties (wet density, porosity), grain size composition (gravel, sand, silt, clay content), and grain size statistics (mean grain size, median grain size). The target variable is P-wave velocity. The dataset is divided into training and test sets according to a certain ratio, and a fixed random seed is used to ensure repeatability.
[0076] Under complete parameter features, four types of models—Random Forest, XGBoost, SVR, and ElasticNet—were trained respectively, and their baseline prediction performance, i.e., baseline root mean square error, was obtained. In the single-feature complete failure experiment, all models exhibited a consistent feature dependency ranking structure, such as... Figure 3 As shown: Porosity ranks first, and its PDR value is significantly higher than other features in all models, forming a fault-like leading position and becoming the strongest dominant control variable; Sand content, particle size statistics (average particle size and median particle size), and wet density are in the middle; Silt content and clay content have low dependence; Gravel content PDR is close to zero, indicating that it does not carry identifiable independent predictive information in this dataset. Under the condition of limited survey resources, it is advisable to prioritize reducing the collection of this parameter.
[0077] In cross-model comparisons, the Random Forest model shows the most concentrated dependency weights on porosity, while ElasticNet's dependency distribution is relatively dispersed, with XGBoost and SVR falling in between. This difference corresponds to how different algorithmic structures handle redundant information: the Bagging structure strengthens the proportion of main variables, the gradient boosting framework forms path sharing through multiple rounds of residual correction, and linear sparsity leads to dependency compression.
[0078] like Figure 4 As shown, in the global random missing data experiment, the PDR ranking of the four models maintained a stable and consistent monotonic gradient relationship across all missing data ranges. Random Forest had the highest overall degradation rate, ElasticNet the lowest, and XGBoost and SVR were in the middle. However, in terms of absolute prediction accuracy, models with high expressive power may still maintain a low RMSE even under high missing data conditions, indicating that model selection should be based on both PDR and RMSE dimensions, rather than solely on the degradation rate.
[0079] In the continuous missing gradient experiment, taking the random forest model as an example, the dependency ranking of each feature did not cross or reverse within the range of 10% to 100% full missing rate, proving that the feature dependency ranking conclusion has structural stability in the continuous information reduction process and the evaluation results are reliable.
[0080] This invention also provides a stability evaluation device for a seabed sediment acoustic velocity prediction model based on performance degradation rate, which can implement the above-mentioned stability evaluation method for a seabed sediment acoustic velocity prediction model based on performance degradation rate. The device includes: The data acquisition module is used to acquire sediment sample datasets; each sediment sample in the sediment sample dataset includes several parametric features and the corresponding target value of P-wave sound velocity; The benchmark root mean square error acquisition module is used to perform a first evaluation of the preset sound velocity prediction model based on all parameter features, and obtain the benchmark root mean square error of the sound velocity prediction model. The single-feature failure module is used to perform a reduction operation on a single parameter feature based on a preset continuous missing rate gradient to obtain the missing features at the current missing rate. The mean root mean square error acquisition module is used to perform a second evaluation of the sound velocity prediction model based on the remaining parameter features and the missing features under the current missing rate, and to obtain the mean root mean square error of the sound velocity prediction model. The performance degradation rate calculation module is used to obtain the first performance degradation rate of the sound speed prediction model under the current missing feature rate based on the benchmark root mean square error and the average root mean square error. The feature dependency ranking determination module is used to extract the second performance decay rate of each missing feature under the extreme missing rate of the sound speed prediction model based on the first performance decay rate, and to rank the missing features according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence. The feature dependency structure graph generation module is used to construct a feature dependency structure graph based on the sorted sequences according to the feature dependency degree. The global feature failure module is used to randomly reduce all parameter features according to a preset global missing rate, and obtain the third performance decay rate of the sound speed prediction model under the preset global missing rate. The evaluation result output module is used to obtain the stability evaluation results of the sound speed prediction model based on the feature-dependent structure map and the third performance decay rate.
[0081] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0082] This invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including a tablet computer, an in-vehicle computer, or similar device.
[0083] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0084] refer to Figure 5 , Figure 5 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1001 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0085] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0086] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0087] This invention also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions to cause the computer device to perform the aforementioned method.
[0088] In summary, this invention provides a method and related equipment for evaluating the stability of a seabed sediment acoustic velocity prediction model based on performance decay rate (PDR). By defining PDR as a dimensionless index, this invention establishes a systematic stability evaluation system for sediment acoustic velocity prediction models. This system combines single-feature failure experiments with global random missing feature experiments, using Monte Carlo repeated simulations to eliminate random bias. This achieves a unified quantification of the robustness of different machine learning models under feature missing feature conditions, and the evaluation results can be directly converted into feature importance ranking, model selection recommendations, and data acquisition priority guidance.
[0089] Specifically, the embodiments of the present invention have the following advantages: 1. This invention proposes PDR, a dimensionless quantitative metric, for the first time, overcoming the shortcomings of existing technologies that only focus on model accuracy and do not evaluate robustness to missing features. Unlike traditional methods that only perform performance evaluation under complete data conditions, this invention can systematically quantify the stability of different models under data incomplete scenarios such as single feature failure and random missing multiple features.
[0090] 2. This invention achieves unified quantification and horizontal comparability of feature dependency strength under different model structures. Addressing the problem that existing methods such as Gini importance and permutation importance cannot compare across models, this invention employs dimensionless normalization to enable direct comparison of feature dependency strength between linear models, tree models, and kernel methods, thereby revealing the differences in feature sensitivity among different algorithm structures.
[0091] 3. This embodiment of the invention does not require retraining the model; it directly performs evaluation calculations based on the test set, resulting in high computational efficiency and low deployment costs. Furthermore, its output can directly support engineering decisions, identifying key and reducible features, and guiding model selection in scenarios with high data loss risk, demonstrating clear engineering application value.
[0092] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.
[0093] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.
[0094] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0095] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0096] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0097] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0098] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0099] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.
Claims
1. A method for stability evaluation of a seabed sediment acoustic velocity prediction model based on performance decay rate, characterized in that, Includes the following steps: Obtain a sediment sample dataset; each sediment sample in the sediment sample dataset includes several parametric features and a corresponding target value for P-wave sound velocity; Based on all the aforementioned parameter characteristics, a first evaluation is performed on the preset sound speed prediction model to obtain the benchmark root mean square error of the sound speed prediction model. Based on the preset continuous missing rate gradient, a reduction operation is performed on a single parameter feature to obtain the missing features at the current missing rate; Based on the remaining parameter features and the missing features under the current missing rate, the sound speed prediction model is evaluated a second time to obtain the average root mean square error of the sound speed prediction model. Based on the benchmark root mean square error and the average root mean square error, obtain the first performance degradation rate of the sound speed prediction model under the current missing feature rate. Based on the first performance decay rate, the second performance decay rate of the sound speed prediction model for each missing feature under extreme missing rate is extracted, and the missing features are sorted according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence. Construct a feature dependency structure graph based on the sorted sequence of feature dependencies; Based on a preset global missing rate, all the parameter features are randomly reduced to obtain the third performance decay rate of the sound speed prediction model under the preset global missing rate. Based on the feature-dependent structure map and the third performance attenuation rate, the stability evaluation result of the sound speed prediction model is obtained.
2. The method according to claim 1, characterized in that, The first evaluation of the preset sound speed prediction model based on all the parameter characteristics to obtain the benchmark root mean square error of the sound speed prediction model includes the following steps: The sediment sample dataset is divided into a training set and a test set; Based on the training set, the pre-trained model is optimized and trained to obtain the sound speed prediction model; Input all the parameter features in the test set into the sound velocity prediction model to obtain the first longitudinal wave sound velocity prediction value; The root mean square error of the sound velocity prediction model is obtained based on the target value of the longitudinal wave sound velocity and the predicted value of the first longitudinal wave sound velocity.
3. The method according to claim 1, characterized in that, The step of performing a reduction operation on a single parameter feature based on a preset continuous missing rate gradient to obtain the missing features at the current missing rate includes the following steps: The sediment sample dataset is divided into a training set and a test set; Based on the training set, obtain the median of a single parameter feature; A sequential missing rate gradient from 0% to 100% is preset, and the current missing rate is determined based on the sequential missing rate gradient; Preprocessed samples are obtained by randomly selecting sediment samples from the test set that represent the current missing rate. The value of a single parameter feature of the preprocessed sample is replaced with the median to obtain the missing feature under the current missing rate.
4. The method according to claim 1, characterized in that, The second evaluation of the sound velocity prediction model, based on the remaining parameter features and the missing features at the current missing rate, to obtain the average root mean square error of the sound velocity prediction model, includes the following steps: The parameter features and the missing features under the current missing rate are input into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value. Based on the target value of the longitudinal wave sound velocity and the predicted value of the second longitudinal wave sound velocity, the initial root mean square error of the sound velocity prediction model is obtained. Return to the step of inputting the parameter features and the missing features under the current missing rate into the sound velocity prediction model to obtain the second longitudinal wave sound velocity prediction value, until the preset number of repeated executions is reached to obtain several initial root mean square errors; The average root mean square error of the sound speed prediction model is obtained based on several initial root mean square errors.
5. The method according to claim 1, characterized in that, The step of obtaining the first performance degradation rate of the sound speed prediction model under the current missing feature rate based on the benchmark root mean square error and the average root mean square error includes the following steps: Obtain the evaluation difference between the average root mean square error and the baseline root mean square error; The first performance degradation rate is obtained by dividing the evaluation difference by the root mean square error of the benchmark.
6. The method according to claim 1, characterized in that, The step of randomly reducing all parameter features according to a preset global missing rate to obtain the third performance degradation rate of the sound speed prediction model under the preset global missing rate includes the following steps: The sediment sample dataset is divided into a training set and a test set; Based on the training set, obtain the median of each parameter feature; Perform an independent and identically distributed Bernoulli mask on all the parameter features in the test set, and randomly generate missing positions according to the preset global missing rate; The median of the corresponding parameter feature is filled into the missing position to generate a global perturbation test set; By performing several Monte Carlo simulations, the global disturbance test set is input into the sound speed prediction model to obtain the third performance degradation rate.
7. A stability evaluation system for a seabed sediment acoustic velocity prediction model based on performance decay rate, characterized in that, include: The data acquisition module is used to acquire sediment sample datasets; Each sediment sample in the sediment sample dataset includes several parametric features and a corresponding target value for longitudinal wave velocity; The benchmark root mean square error acquisition module is used to perform a first evaluation on the preset sound speed prediction model based on all the parameter features, and obtain the benchmark root mean square error of the sound speed prediction model. The single-feature failure module is used to perform a reduction operation on a single parameter feature according to a preset continuous missing rate gradient to obtain the missing feature under the current missing rate. The average root mean square error acquisition module is used to perform a second evaluation on the sound speed prediction model based on the remaining parameter features and the missing features under the current missing rate, so as to obtain the average root mean square error of the sound speed prediction model. The performance decay rate calculation module is used to obtain the first performance decay rate of the sound speed prediction model under the current missing rate based on the benchmark root mean square error and the average root mean square error. The feature dependency ranking determination module is used to extract the second performance decay rate of the sound speed prediction model for each missing feature under extreme missing rate according to the first performance decay rate, and to rank the missing features according to the magnitude of the second performance decay rate to generate a feature dependency ranking sequence. The feature dependency structure graph generation module is used to construct a feature dependency structure graph based on the sorted sequence according to the feature dependency degree. The global feature failure module is used to randomly reduce all the parameter features according to a preset global missing rate, and obtain the third performance decay rate of the sound speed prediction model under the preset global missing rate. The evaluation result output module is used to obtain the stability evaluation result of the sound speed prediction model based on the feature-dependent structure map and the third performance decay rate.
8. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.