Submerged unmanned ship resistance approximation model data fusion method considering uncertainty
By constructing a Gaussian process regression model and weighted fusion algorithm, combining experimental data, simulation data and expansion data, a BP neural network resistance approximation model is built, which solves the problem of multi-source data uncertainty in ship design and improves prediction accuracy and decision reliability.
Patent Information
- Application Number
- CN202510243150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-08
AI Technical Summary
In the field of ship design, how to efficiently utilize multi-source data, especially combining data fusion with approximate modeling technology, improve prediction accuracy and decision-making reliability, has not been effectively solved.
By constructing a Gaussian process regression model, the mean square variance and average absolute error of multi-source data are calculated, the weight is allocated, and combined with the weighted fusion algorithm, the experimental data, simulation data and expansion data are effectively combined to build a BP neural network resistance approximation model.
It effectively reduces the uncertainty brought by multi-source data, improves the generalization ability and prediction stability of the approximate model, and provides technical support for the design optimization of submarine floating unmanned ships.
Smart Images

Figure CN120277798A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data fusion, and in particular, to a data fusion method for a resistance approximation model of a submersible unmanned ship considering uncertainty. Background Art
[0002] The submersible unmanned ship is a new type of semi-submersible vehicle with excellent concealment and flexibility, capable of autonomous navigation on the water surface and underwater. This type of unmanned ship usually has a long endurance and low operating costs, and is suitable for long-term monitoring and data collection tasks. With the development of technology, approximation models, as an efficient alternative technology, have been widely used in the design optimization of complex products. From the initial proposal to the current mature application, approximation modeling technology has achieved remarkable results in solving various engineering problems, especially in the field of ship design, which can effectively shorten the calculation time and improve efficiency.
[0003] Data fusion technology, especially its application in a multi-sensor environment, aims to effectively fuse and aggregate data from different sensors. This technology can be widely applied in various fields. By combining data obtained from different prediction methods, it can make up for their respective deficiencies, thereby improving the accuracy and reliability of the input data of the surrogate model. The fused dataset can significantly improve the prediction accuracy of the surrogate model, effectively reduce the uncertainty brought by a single data source, and provide more reliable data support for the design optimization of the submersible unmanned ship.
[0004] The Bayesian method is based on observed data, regards unknown parameters as random variables through Bayesian statistical theory, and quantitatively describes uncertainty using probability distributions. Gaussian process regression, as a non-parametric Bayesian regression method, uses a Gaussian process as the prior probability distribution to model data, and based on this, makes predictions and inferences on the data. The weight calculation in the Gaussian process regression algorithm is the core of the data fusion method, and the weight reflects the importance of each data source in the fusion result. By reasonably calculating the weight, it can ensure that high-precision data has a greater influence in the model establishment process.
[0005] Currently, in the field of ship design, the research on data fusion technology lags behind and is still in the initial development stage. How to efficiently utilize multi-source data, especially combining data fusion with approximation modeling technology to improve the prediction accuracy and decision-making reliability in the ship design process, is the current research focus. Summary of the Invention
[0006] In response to the above-mentioned technical problems, a data fusion method for a resistance approximation model of a submersible unmanned ship considering uncertainty is provided. The present invention effectively reduces the uncertainty brought by multi-source data through the fusion of multi-source sample data.
[0007] The technical means adopted by the present invention are as follows:
[0008] A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty, comprising:
[0009] S1. Taking the submersible and floating unmanned ship as the target ship type, generating a multi-source data set through ship model tests, numerical simulation tests and cloud computing technology;
[0010] S2. Constructing a Gaussian process regression model, calculating the mean square deviation of each data set in the multi-source data set, and comprehensively evaluating the uncertainty of the data by combining the mean absolute error and variance to obtain the data uncertainty quantification result;
[0011] S3. Based on the data uncertainty quantification result, assigning corresponding weights to each data set, and effectively combining the experimental data, simulation data and extended data through a weighted fusion algorithm to finally obtain a fused data set;
[0012] S4. Using the fused data set to construct a BP neural network resistance approximation model.
[0013] Further, the ship model experimental data and numerical simulation experiments in step S1 take the submersible and floating unmanned ship as the target ship type, only consider the resistance characteristics of the main hull of the ship, ignore the influence of appendages on the resistance, and obtain experimental data and simulation data.
[0014] Further, the cloud computing technology in step S1 is to calculate six digital features by using an inverse cloud generator, perform data expansion by using a forward cloud generator, set the expansion multiple to 100 times, and obtain the speed, submergence depth and resistance insufficient information supplement of the submersible and floating unmanned ship.
[0015] Further, in step S2, the constructed Gaussian process regression model is as follows:
[0016] f*|X,y,X*~N(m*(x),cov(f*))
[0017] Where f * represents the predicted value, X represents the input of the training set, y represents the observed value, X * represents the test value, N represents the normal distribution, m * (x) represents the expected value of the predicted value, Where K(x*,x) represents the n×1 order covariance matrix between the test value X * and the input X of the training set, K(X,X) represents the n×n order symmetric positive definite covariance matrix, represents the variance of the Gaussian white noise in the observed value y, I represents the n×n identity matrix, n represents the number of training samples, cov(f * ) represents the covariance of the predicted output,
[0018] Further, in step S2, the uncertainty of the data is comprehensively evaluated by combining the mean absolute error and the variance, including: repeating the processing for each data point in the simulation data set and the experimental data set to increase the occurrence frequency of the samples.
[0019] Further, in step S2, the uncertainty of the data is comprehensively evaluated by combining the mean absolute error and the variance, including: performing farthest point sampling on the augmented data, and the specific steps are as follows:
[0020] Given the data sample S, set the number of sampling points to N;
[0021] Initialize the set, randomly select a point and delete this point from the original set;
[0022] For each point in the data sample S, calculate its distance to all points in the data sample S t , and each point retains its minimum distance to the nearest point in the sampled point set;
[0023] Among the points in the data sample S, select the point farthest from the sampled point set, add it to the data sample S t , and delete this point from the data sample S;
[0024] Repeat the above steps until the number of the data sample S t reaches the number of sampling points N.
[0025] Further, step S3 specifically includes:
[0026] S31. Use Gaussian process regression to fit the enhanced simulation data set, the enhanced experimental data set, and the sampled augmented data set;
[0027] S32. According to the variance and root mean square error obtained after fitting, obtain the weights of each data source, and the formula is as follows:
[0028]
[0029] where MAE(X, h) represents the variance obtained after fitting, m represents the number of data, h(x i ) represents the predicted value, y i represents the actual value, RMSE(X, h) represents the root mean square error obtained after fitting, ω i represents the weight of the output of the i-th data source, represents the mean square deviation of the output of the i-th data source, and j represents the traversal index for n elements;
[0030] S33. Combine the three groups of data through a weighted fusion algorithm, effectively integrate the experimental data, simulation data, and augmented data, and finally obtain a fused data set as follows:
[0031] y = W T X = [ω1, ω2, ··· ω n [x1, x2, ··· x n T
[0032] Among them, y represents the fused data set, and W T represents the weight vector, ω1, ω2, ··· ω n represents the weight of the output of the i-th data source, and x1, x2, ··· x n represents the output of the i-th data source.
[0033] Furthermore, in step S31, the K-fold cross-validation method is used to fit the enhanced simulation data set, which specifically includes:
[0034] S311. Randomly divide the data into K groups, select one training fold as the test data set, and the remaining K - 1 as the training set;
[0035] S312. Use the training data set to train the model and evaluate the model performance using the test data set;
[0036] S313. Repeat the K-fold cross-validation t times to obtain more stable and reliable model evaluation results, and calculate the predicted values of each sample point in the sample set. The formula is as follows:
[0037]
[0038] Among them, represents the predicted value of each sample point in the sample set, t represents the number of times of K-fold cross-validation, represents the predicted result of the i-th category of the j-th sample x j for the v-th prediction, and v represents the v-th prediction being calculated currently.
[0039] Furthermore, step S4 specifically includes:
[0040] S41. Adopt a single hidden layer model and improve the prediction accuracy by increasing the number of hidden layer neuron nodes;
[0041] S42. The input variables are the speed and diving depth of the unmanned ship, and the number of input layer nodes is 2;
[0042] S43. The transfer function of the hidden layer is tansig, the transfer function of the output layer is logsig, and the learning function is learngdm;
[0043] S44. Use the mean square error as the performance function to evaluate the network prediction error;
[0044] S45. Determine the optimal number of hidden layer nodes through comparative experiments;
[0045] S46. Randomly sample the three datasets respectively. For the sampled part of the dataset, the target variable R is weighted and fused according to the calculated weights, while the feature variables D and V are fused through simple arithmetic mean to maintain feature consistency. For the remaining part of the dataset, average fusion is performed without considering weights, and only the arithmetic means of D, V, and R are taken;
[0046] S47. Under multiple combinations of data source ratios, train several BP neural network approximation models, and select the optimal fusion ratio to construct an approximation model by comparing and analyzing the evaluation indicators of each model.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] 1. A data fusion method for the resistance approximation model of a submersible unmanned ship considering uncertainty provided by the present invention effectively reduces the uncertainty brought by multi-source data through the fusion of multi-source sample data.
[0049] 2. A data fusion method for the resistance approximation model of a submersible unmanned ship considering uncertainty provided by the present invention trains several BP neural network approximation models under multiple combinations of data source ratios, and selects the optimal fusion ratio to construct an approximation model by comparing and analyzing the evaluation indicators of each model, improving the generalization ability and prediction stability of the approximation model, and providing technical support for the massive sample reduction strategy in the hull form optimization design of submersible unmanned ships.
[0050] For the above reasons, the present invention can be widely promoted in the fields of data fusion and the like. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a flow chart of the method of the present invention.
[0053] Figure 2 It is a comparison chart of the resistance of each working condition of the ship model test provided by the embodiment of the present invention.
[0054] Figure 3This is the comparison chart of resistances under various working conditions in the numerical simulation provided by the embodiments of the present invention.
[0055] Figure 4 This is the comparison chart of the cloud expansion model when the diving depth is 0m provided by the embodiments of the present invention.
[0056] Figure 4 In the figure: (a) Before data expansion: Working condition 0#; (b) After data expansion: Working condition 0#
[0057] Figure 5 This is the sampling result chart of the farthest point provided by the embodiments of the present invention.
[0058] Figure 6 This is the three-dimensional scatter plot of the fusion dataset and the multi-source dataset provided by the embodiments of the present invention.
[0059] Figure 7 This is the comparison chart of the accuracy of the BP neural network provided by the embodiments of the present invention. Detailed implementation manners
[0060] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0061] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0062] As Figure 1 shown, the present invention provides a method for data fusion of the resistance approximate model of a submersible and floating unmanned ship considering uncertainty, including:
[0063] S1. Taking the submersible and floating unmanned ship as the target ship type, generating a multi-source dataset through ship model tests, numerical simulation tests and cloud computing technology;
[0064] S2. Construct a Gaussian process regression model, calculate the mean square error of each dataset in the multi-source dataset, and comprehensively evaluate the uncertainty of the data by combining the mean absolute error and variance to obtain the quantification result of data uncertainty.
[0065] S3. Based on the quantification result of data uncertainty, assign corresponding weights to each dataset, and effectively combine the experimental data, simulation data, and augmented data through a weighted fusion algorithm to finally obtain a fused dataset, thereby reflecting the contribution of high-precision data in the data fusion process and reducing the uncertainty brought by multi-source data.
[0066] S4. Use the fused dataset to construct a BP neural network drag approximation model.
[0067] In specific implementation, as a preferred implementation manner of the present invention, the ship model experiment data and numerical simulation experiment in step S1 take the submersible and floating unmanned ship as the target ship type, only consider the drag characteristics of the main hull of the ship, ignore the influence of appendages on the drag, and obtain the experimental data and simulation data.
[0068] In this embodiment, the drag of the submersible and floating unmanned ship is obtained through three methods: ship model experiment, numerical simulation experiment, and cloud computing augmentation. In order to ensure the accuracy and reliability of the research data, during the process of collecting data through the ship model drag test, only the drag characteristics of the main hull of this type of submersible and floating unmanned ship are considered, and the influence of appendages on the drag is ignored. The scale of the water tank is 160.0m×7.0m×3.7m (length×width×water depth). Taking the submersible and floating unmanned ship as the target ship type, ship model tests and numerical simulation tests are carried out, and six different submergence conditions from 0# to 5# are set, and these conditions correspond to submergence depths of 0m, 0.054m, 0.32m, 0.48m, 0.64m, and 0.96m respectively. The test measures the still water drag at multiple speed points from 0.4m / s to 1.7m / s for each condition, which are 0.4m / s, 0.6m / s, 0.8m / s, 1.0m / s, 1.1m / s, 1.2m / s, 1.3m / s, 1.4m / s, 1.5m / s, 1.6m / s, and 1.7m / s respectively. The corresponding Froude number (Fr) range is 0.1010 to 0.4291, and the Reynolds number (Re) range is 0.5284×106 to 2.2459×106. The comparison of the drag of the ship model test and numerical simulation under each condition is respectively represented by Figure 2 and Figure 3 to indicate.
[0069] In specific implementation, as a preferred implementation manner of the present invention, the cloud computing technology in step S1 calculates six digital features by using the reverse cloud generator, performs data augmentation by using the forward cloud generator, sets the augmentation multiple to 100 times, and obtains the insufficient information supplementation of the speed, submergence depth, and drag of the submersible and floating unmanned ship. Such as Figure 4As shown, it is a comparison diagram of the cloud expansion model when the diving depth is 0m;
[0070] In this embodiment, the fused dataset is used as the target sample for data augmentation. To prevent the augmented data from having too large a deviation from the original real data, a constraint condition is set based on the extreme values of V and D in the original dataset and allowing them to float by 10%. The specific value range is: 0.2000 < V < 1.800 (m / s), 0.0000 < D < 1.000 (m). Any augmented sample outside the following range will be excluded to ensure the accuracy and reliability of the augmented data. First, taking the 55 groups of "speed - resistance" data in the fused sample set as the target event, its six digital features are calculated using the reverse cloud generator. Subsequently, the forward cloud generator is used for data augmentation to solve the problem of small initial data volume and sparse distribution. The augmentation multiple is set to 100 times to obtain the insufficient information supplement of the speed, diving depth, and resistance of the submersible unmanned ship. After augmenting the data to 100 times, 5500 augmented data samples of the resistance performance of the submersible unmanned ship are obtained. Compared with the original fused sample set, the data samples in the speed range of 0.8 - 1 m / s and the diving depth range of 0 - 0.5 m in the augmented sample set are significantly enhanced, making the overall data distribution more reasonable and optimized.
[0071] Specifically in implementation, as a preferred implementation manner of the present invention, in step S2, the constructed Gaussian process regression model is as follows:
[0072] f*|X,y,X*~N(m*(x),cov(f*))
[0073] where f * represents the predicted value, X represents the input of the training set, y represents the observed value, X * represents the test value, N represents the normal distribution, m * (x) represents the expected value of the predicted value, where K(x*,x) represents the n×1 order covariance matrix between the test value X * and the input X of the training set, K(X,X) represents the n×n order symmetric positive definite covariance matrix, represents the variance of the Gaussian white noise in the observed value y, I represents the n×n identity matrix, n represents the number of training samples, cov(f * ) represents the covariance of the predicted output,
[0074] Specifically in implementation, as a preferred implementation manner of the present invention, in step S2, the uncertainty of the data is comprehensively evaluated by combining the mean absolute error and variance, including: repeating the processing of each data point in the simulation dataset and the experimental dataset to increase the occurrence frequency of the samples.
[0075] In this embodiment, the basis for selecting duplicate data points is mainly based on their importance in model training, especially those points that exhibit high uncertainty under specific conditions. This approach aims to improve the model's fitting ability and reduce potential overfitting by enhancing the diversity and coverage of the dataset.
[0076] During specific implementation, as a preferred implementation manner of the present invention, in step S2, the uncertainty of the data is comprehensively evaluated by combining the mean absolute error and variance, including: performing farthest point sampling on the augmented data, and the specific steps are as follows:
[0077] Given a data sample S, set the number of sampling points to N;
[0078] Initialize the set, randomly select a point and delete this point from the original set;
[0079] For each point in the data sample S, calculate its distance to all points in the data sample S t Among them, each point retains the minimum distance to the nearest point in the sampled point set;
[0080] Among the points in the data sample S, select the point farthest from the sampled point set, add it to the data sample S t , and delete this point from the data sample S;
[0081] Repeat the above steps until the number of the data sample S t reaches the number of sampling points N. As Figure 5 shown, it is a graph of the farthest point sampling result.
[0082] During specific implementation, as a preferred implementation manner of the present invention, step S3 specifically includes:
[0083] S31. Use Gaussian process regression to fit the enhanced simulation dataset, the enhanced experimental dataset, and the sampled augmented dataset;
[0084] S32. According to the variance and root mean square error obtained after fitting, obtain the weight of each data source, and the formula is as follows:
[0085]
[0086] wherein, MAE(X,h) represents the variance obtained after fitting, m represents the number of data, h(x i ) represents the predicted value, y i represents the actual value, RMSE(X,h) represents the root mean square error obtained after fitting, ω i represents the weight of the output of the i-th data source, represents the mean square deviation of the output of the i-th data source, and j represents the traversal index for n elements;
[0087] S33. Combine the three groups of data through a weighted fusion algorithm, effectively integrate the experimental data, simulation data, and augmented data, and finally obtain a fused data set as follows:
[0088] y = W T X = [ω1, ω2, ··· ω n [x1, x2, ··· x n T
[0089] Among them, y represents the fused data set, and W T represents the weight vector, ω1, ω2, ··· ω n represents the weight of the output of the i-th data source, and x1, x2, ··· x n represents the output of the i-th data source. As Figure 6 shown, it is the three-dimensional scatter plot of the fused data set and the multi-source data set.
[0090] In specific implementation, as a preferred implementation manner of the present invention, in step S31, the K-fold cross-validation method is used to fit the enhanced simulation data set to reduce the overfitting problem caused by the characteristics of the data itself, specifically including:
[0091] S311. Randomly divide the data into K groups, select one training fold as the test data set, and the remaining K - 1 as the training set;
[0092] S312. Use the training data set to train the model and evaluate the model performance using the test data set;
[0093] S313. Repeat the K-fold cross-validation t times (usually t = 5 or t = 10) to obtain more stable and reliable model evaluation results, and calculate the predicted values of each sample point in the sample set. The formula is as follows:
[0094]
[0095] Among them, represents the predicted value of each sample point in the sample set, t represents the number of times of K-fold cross-validation, represents the predicted result of the i-th class of the j-th sample x j for the v-th prediction, and v represents the v-th prediction being calculated currently.
[0096] In specific implementation, as a preferred implementation manner of the present invention, step S4 specifically includes:
[0097] S41. Adopt a single hidden layer model and improve the prediction accuracy by increasing the number of hidden layer neuron nodes rather than the number of network layers;
[0098] S42. The input variables are the speed and diving depth of the unmanned ship, and the number of input layer nodes is 2;
[0099] S43. The transfer function of the hidden layer is tansig, the transfer function of the output layer is logsig, and the learning function is learngdm;
[0100] S44. The mean square error is used as the performance function to evaluate the network prediction error;
[0101] S45. The optimal number of hidden layer nodes is determined through comparative experiments; in this embodiment, the optimal number of nodes of the multi-factor prediction model is 5, and the minimum value of MSE is 0.0126.
[0102] S46. Random samples are taken from the three data sets respectively. The sampled data sets are weighted and fused for the target variable R according to the calculated weights. The feature variables D and V are fused by simple arithmetic mean to maintain feature consistency. The remaining data sets are averaged and fused without considering weights, and only the arithmetic means of D, V, and R are taken;
[0103] S47. Under various combinations of data source ratios, several BP neural network approximation models are trained, and the optimal fusion ratio is selected through comparative analysis of the evaluation indexes of each model to construct an approximation model. As Figure 7 shown, it is a precision comparison chart of the approximation models constructed for six cases.
[0104] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data fusion method for the resistance approximate model of a submersible and floating unmanned ship considering uncertainty, characterized in that, Including: S1. Taking the submersible and floating unmanned ship as the target ship type, generating a multi-source data set through ship model tests, numerical simulation tests and cloud computing technology; S2. Constructing a Gaussian process regression model, calculating the mean square error of each data set in the multi-source data set, and comprehensively evaluating the uncertainty of the data by combining the mean absolute error and variance to obtain the data uncertainty quantification result; S3. Based on the data uncertainty quantification result, assigning corresponding weights to each data set, and effectively combining the experimental data, simulation data and augmented data through a weighted fusion algorithm to finally obtain a fused data set; S4. Using the fused data set to construct a BP neural network resistance approximation model.
2. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 1, characterized in that The ship model experimental data and numerical simulation experiments in step S1 take the submersible and floating unmanned ship as the target ship type, only considering the resistance characteristics of the main hull of the ship and ignoring the influence of appendages on the resistance to obtain experimental data and simulation data.
3. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 1, characterized in that The cloud computing technology in step S1 is to calculate six digital features by using the reverse cloud generator, perform data augmentation by using the forward cloud generator, set the augmentation multiple to 100 times, and obtain the insufficient information supplement of the speed, submergence depth and resistance of the submersible and floating unmanned ship.
4. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 1, characterized in that In step S2, the constructed Gaussian process regression model is as follows: f*|X,y,X*~N(m*(x),cov(f*)) Among them, f * represents the predicted value, X represents the input of the training set, y represents the observed value, and X * represents the test value, N represents the normal distribution, and m * (x) represents the expected value of the predicted value, Among them, K(x*, x) represents the n×1 order covariance matrix between the test value X * and the input X of the training set, K(X, X) represents the n×n order symmetric positive definite covariance matrix, represents the variance of the Gaussian white noise in the observed value y, I represents the n×n identity matrix, n represents the number of training samples, and cov(f * ) represents the covariance of the predicted output, 5. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 1, characterized in that In step S2, comprehensively evaluating the uncertainty of the data by combining the mean absolute error and variance includes: repeatedly processing each data point in the simulation data set and the experimental data set to increase the occurrence frequency of the samples.
6. A data fusion method for a resistance approximate model of a submersible and floating unmanned vehicle considering uncertainty according to claim 5, characterized in that In step S2, comprehensively evaluating the uncertainty of the data by combining the mean absolute error and variance includes: using the farthest point sampling for the augmented data, and the specific steps are as follows: Given the data sample S, set the number of sampling points to N; Initialize the set, randomly select a point and delete this point from the original set; For each point in the data sample S, calculate its distance to all points in the data sample S t and each point retains the minimum distance to the nearest point in the sampled point set; Among the points in the data sample S, select the point that is farthest from the set of sampled points and add it to the data sample S t , and delete this point from the data sample S; Repeat the above steps until the number of data samples S t reaches the number of sampling points N.
7. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 1, characterized in that, Step S3 specifically includes: S31. Using Gaussian process regression to fit the enhanced simulation data set, the enhanced experimental data set and the sampled augmented data set; S32. According to the variance and mean square error obtained after fitting, obtain the weights of each data source, and the formula is as follows: Among them, MAE(X, h) represents the variance obtained after fitting, m represents the number of data, h(x i ) represents the predicted value, y i represents the actual value, RMSE(X, h) represents the root mean square error obtained after fitting, ω i represents the weight of the output of the i-th data source, represents the mean square error of the output of the i-th data source, and j represents the traversal index for n elements; S33. Fuse the three groups of data through a weighted fusion algorithm, effectively combine the experimental data, simulation data and augmented data, and finally obtain a fused data set, as follows: y = W T X = [ω1, ω2, ··· ω n [x1, x2, ··· x n T Among them, y represents the fused dataset, and W T is represented as the weight vector, ω1, ω2, ··· ω n represents the weight of the output of the i-th data source, and x1, x2, ··· x n represents the output of the i-th data source.
8. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 7, characterized in that In step S31, using the K-fold cross-validation method to fit the enhanced simulation data set specifically includes: S311. Randomly divide the data into K groups, select one training fold as the test data set, and the remaining K - 1 as the training set; S312. Use the training data set to train the model and evaluate the model performance using the test data set; S313. Repeat the K-fold cross-validation t times to obtain a more stable and reliable model evaluation result, and calculate the predicted value of each sample point in the sample set, and the formula is as follows: Among them, represents the predicted value of each sample point in the sample set, t represents the number of times of K-fold cross-validation, represents the prediction result of the i-th class of the j-th sample x j for the v-th prediction, and v represents the v-th prediction being currently calculated.
9. A data fusion method for the resistance approximation model of a submersible and floating unmanned ship considering uncertainty according to claim 1, characterized in that Step S4 specifically includes: S41. Using a single hidden layer model and improving the prediction accuracy by increasing the number of hidden layer neuron nodes; S42. The input variables are the speed and submergence depth of the unmanned ship, and the number of input layer nodes is 2; S43. The transfer function of the hidden layer is tansig, the transfer function of the output layer is logsig, and the learning function is learngdm; S44. The mean squared error is used as the performance function to evaluate the network prediction error; S45. The optimal number of hidden layer nodes is determined through comparative experiments; S46. Random sampling is performed on the three datasets respectively. For the sampled part of the dataset, the target variable R is weighted and fused according to the calculated weights, while the feature variables D and V are fused through simple arithmetic mean to maintain feature consistency. For the remaining part of the dataset, average fusion is performed without considering weights, and only the arithmetic means of D, V, and R are taken; S47. Under multiple combinations of data source ratios, several BP neural network approximation models are trained, and the optimal fusion ratio is selected through comparative analysis of the evaluation indicators of each model to construct an approximation model.