Ocean sound velocity distribution forecasting method based on pruning mixed basis function KAN

Through the neural network model based on pruning mixed basis function KAN, the problem of insufficient generalization ability in the existing technology is solved, and accurate prediction of ocean sound velocity distribution and improvement of training efficiency is achieved.

CN120296361AActive Publication Date: 2025-07-11OCEAN UNIV OF CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510772339.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing models need to readjust the structure when dealing with different types of ocean sound speed distribution tasks. The generalization capability is insufficient, and the remote sensing sea surface temperature has limited impact on the entire ocean sound speed, which cannot accurately reflect the changes in the deep-sea area.

Method used

Using a method based on the pruning mixed basis function Kolmogolov-Arnold neural network (KAN), the hybrid basis function KAN network model is constructed, combined with the pruning strategy, the best basis function is adaptively selected to predict the ocean sound velocity distribution.

Benefits of technology

It realizes wide applicability to data of different distribution types, improves forecast accuracy and training efficiency, and can adapt to multiple time series data without specific optimization for a certain data type.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296361A_ABST
    Figure CN120296361A_ABST
Patent Text Reader

Abstract

The invention discloses an ocean sound velocity distribution forecasting method based on a pruning mixed basis function KAN, and belongs to the technical field of ocean observation. According to the method, a mixed primary function KAN network model is constructed, the model comprises an input layer, a mixed primary function KAN layer, a full connection layer and an output layer, n primary functions exist in the mixed primary function KAN layer, and the optimal primary function can be adaptively selected through a pruning strategy in the model training process. The method can adapt to various time series data, does not need to carry out specific optimization for a certain data type, is high in forecasting accuracy, and provides important guiding significance for marine scientific research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ocean observation, and specifically relates to a method for predicting ocean sound speed distribution based on a pruned hybrid basis function KAN. Background Art

[0002] The underwater sound speed distribution affects the propagation mode of underwater acoustic signals and is one of the important parameters for underwater positioning and navigation. In order to efficiently obtain accurate underwater sound speed distribution, researchers have proposed various methods, including methods for constructing spatial sound speed distribution using remote sensing data, historical sound speed data, and longitude and latitude coordinates, as well as prediction methods using historical sound speed time series data. Dr. Wu Pengfei et al. from Ocean University of China proposed a multi-source data fusion method, mainly using model learning of the influence of remote sensing sea surface temperature on the characteristics of underwater sound speed distribution. However, the influence of remote sensing sea surface temperature on the sound speed of the entire ocean is limited and cannot reflect the changes in sound speed in the deep sea area. In order to estimate the sound speed distribution of the entire depth of the ocean, Lu Jiajun et al. proposed a hierarchical long short-term memory neural network (LSTM) for rapid sound speed profile prediction. However, this model can only process data for a single task and requires readjusting the model structure when switching tasks, with insufficient generalization ability.

[0003] Existing models can only handle single-type tasks. For time series type prediction tasks, such as predicting sound speed, temperature, salinity, etc., it is necessary to optimize the model structure according to different data distributions, and the model does not have wide applicability. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for predicting ocean sound speed distribution based on a pruned hybrid basis function Kolmogorov–Arnold neural network (KAN) to make up for the deficiencies of the existing technology.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A method for predicting ocean sound speed distribution based on a pruned hybrid basis function KAN, the specific steps are as follows: S1: Extract the spatial historical ocean data set: Determine the spatial coordinates as longitude M° and latitude N° according to the data prediction location, and extract the ocean data at different historical times at this location to form a historical data set; S2: Determine the time resolution of data prediction, the start and end times of prediction, and the maximum time length: Use to represent the time resolution required by the prediction task, use to represent the start time of prediction, and use to represent the maximum time length of prediction; S3: Secondary screening of the spatial historical ocean data set: According to the time resolution , screening data from the historical dataset in S1 to form a reference dataset; S4: preprocessing the reference dataset; S5: constructing a hybrid basis function KAN network model, which includes an input layer, a hybrid basis function KAN layer, a fully connected layer, and an output layer. There are n basis functions in the hybrid basis function KAN layer, and the best basis function can be adaptively selected through a pruning strategy during the model training process. The number of neurons for the basis function is N, and the number of neurons for the fully connected layer is N; according to different depth stratifications, the ocean data distribution forecast is executed layer by layer, and a hybrid basis function KAN network is constructed for each depth layer. For the hybrid basis function KAN network constructed for each depth layer, the prediction result is a time series of output with a length of ; S6: training the hybrid basis function KAN network model; S7: model prediction: according to the forecast time resolution, forecast start time, and forecast maximum time length specified by the task, determine the number of forecast iterations, the historical distribution reference data corresponding to the forecast start time, and use the data of 2 or more complete change cycles before the forecast start time as the training set of the model to obtain the prediction result, and then perform denormalization processing on the prediction result to output the forecast future distribution data.

[0006] Further, in the above S3: the time range covered by the training data includes at least the data of the first 2 complete change cycles before the forecast start time For example, when the forecast time resolution unit is "hour", it is necessary to make the sample cover no less than the first 2 full-day cycles before, that is, screen the sample data no less than 2 days before the forecast time; when the forecast time resolution is "month", it is necessary to make the sample cover no less than the first 2 full-year cycles before, that is, screen the sample data no less than 2 years before the forecast time; Let the secondary screened historical dataset be expressed as: ; where represents the i-th data profile, with a total of I, and the i-th data profile is further expressed as: ; where is the data value, and the subscript represents the -th depth layer, represents the transpose of the vector; in formula (1), the S dataset is arranged in chronological order, and the time interval is consistent with the time resolution ; = 0, 1,..., J, in the formula, is the depth layer index label, and the maximum depth layer of the data distribution is dJ; Expanding formula (1) gives the dataset matrix: .

[0007] Furthermore, the said S4 includes: S4-1: Divide the training set and the validation set Take the first I - C columns of the historical dataset S as the training dataset T, as shown in formula (4); the last C columns of S are used as the validation dataset E, as shown in formula (5), where I - C ≥ C; C is the natural change period of the data; ; ; S4-2: Normalize the training dataset T to , and the normalization formula is as shown in formula (6), where is the data after normalization, is the data before normalization, is the mean of the sequence data, is the standard deviation of the sequence data; for a certain depth layer, is obtained from formula (7), represents the sum of the data in the j-th row of matrix T from the first column to the last column; is obtained from formula (8); then the normalized training dataset is specifically divided into the training input set and the training label set ; the training input set X is as shown in formula (10), and the training label set Y is as shown in formula (11); ; ; ; ; ; ; S4-3: Re-segment the training input set and the training label set according to the depth layer The training input set and the training label set are respectively segmented into , ; , are respectively represented by formulas (12) and (13); ; ; It represents the data of the j-th row of matrix X, that is, the data of the j-th deep layer. Each time a window of size C is taken with a window step of 1. It represents the data of the j-th row of matrix Y, that is, the data of the j-th deep layer. Each time a window of size 1 is taken with a window step of 1.

[0008] Furthermore, the S6 includes: S6-1: Before model training, first determine that the learning rate of each deep layer is and the number of iterations is , where j represents the corresponding deep layer, j = 1, 2,..., J; S6-2: During model training, train the mixed basis function KAN network model of each deep layer separately. Use the mixed basis function KAN to represent the mixed basis function KAN network of the j-th deep layer. The input and output of the model are expressed as: ; where represents the output matrix of the model, expressed as: ; After that, calculate the mean square error MSE between the output value of the model and the training label set: ; Then use the value of MSE to update the parameters of the model; the training methods of other deep layers are the same; S6-3: Implement the pruning strategy; S6-4: Model verification. During model verification, it is also carried out for each deep layer separately; use the data of the next cycle of the last time period of the training input set as the input of the model for prediction to predict the data of the next time resolution, that is, a sequence of data with a length of 1. Then take new values by sliding the window (window step is 1) as the input for predicting the next time resolution until a time period is predicted. The other deep layers are similar.

[0009] Even further, the S6-3 includes: S6-31: First pre-train the model. At the beginning of model training, freeze the weight coefficients so that they cannot be trained, and only train the parameters in the basis function. All weight coefficients are initialized as . After pre-training the model for a certain number of times, at this time, the output of the mixed basis function KAN layer is: ; S6-32: Prune the model. First, freeze the trainable parameters in all basis functions, then unfreeze the weight coefficients so that they can be trained. Then, after a certain number of training iterations, observe the weight coefficients of each basis function, find the largest coefficient. At this time, the basis function corresponding to this coefficient is the optimal basis function. Unfreeze its trainable parameters, set its weight coefficient to 1, and set other weight coefficients to 0. Then continue training until the iteration count is completed.

[0010] Furthermore, the S6-4 includes: The prediction result of the first deep layer is a time series of length C, as shown in Equation (17): ; Similarly for the remaining deep layers, then the finally predicted data set is , as shown in Equation (18): ; Since the training data set was normalized before model training, the predicted data should be denormalized (the inverse operation of the original normalization operation), that is, the predicted data is restored to the original data scale range for easy comparison and analysis with the validation data.

[0011] The denormalization formula is: ; The predicted result after denormalization : ; Among them, represents the value predicted at the i-th time resolution of the j-th deep layer.

[0012] After the denormalization operation is completed, calculate the root mean square error (RMSE) between the actual value and the predicted value: ; Among them represents the RMSE of the i-th time resolution of the prediction. If the value of RMSE is not less than the expected threshold, return to S6, increase the iteration count of the training model, and continue to train the model until the model converges, that is, RMSE is less than the expected threshold or there is no obvious change.

[0013] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention proposes a method for predicting ocean sound speed distribution based on a pruned hybrid basis function Kolmogorov - Arnold neural network, which can be widely applied to data fitting of different distribution types. For a given task, the model uses the historical ocean data in the spatial region where the task is located as a reference to accurately predict the future ocean sound speed distribution in the target region. At the same time, in order to improve the training efficiency of the model, a pruning strategy is also proposed to reduce the training time of the model. The present invention can adapt to a variety of time series data without specific optimization for a certain data type, and has a high prediction accuracy, providing important guiding significance for ocean science research. Description of the Drawings

[0014] Figure 1 is the overall flowchart of the present invention.

[0015] Figure 2 is a schematic diagram of the basic structure of the hybrid basis function KAN neural network model.

[0016] Figure 3 is a comparison chart of the predicted sound speed distribution and the actual sound speed.

[0017] Figure 4 is a comparison chart of the predicted temperature distribution and the actual distribution. Detailed Embodiments

[0018] To make the purpose, technical solutions and advantages of the present invention more clear, the following further describes the present invention in detail with reference to specific embodiments and the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative work fall within the scope of protection of the present invention.

[0019] Embodiment 1: A method for predicting ocean time series distribution based on a pruned hybrid basis function Kolmogorov - Arnold neural network is specifically implemented as follows, as Figure 1 shown: I. Predicting Sound Speed Step 1: Extract the spatial historical ocean data set according to the data prediction location According to the data prediction location, the spatial coordinates are determined as 168.5°E longitude and 16.5°N latitude. Select the sound speed profile data measured in 60 months from 2016 to 2020 in this spatial region from the global Argo data set to form a spatial historical sound speed data set as the reference data set of the model.

[0020] Step 2: Determine the data prediction time resolution, prediction start and end times, and maximum time length Use Indicates the time resolution of the forecast task requirements, which is 1 month; use to represent the forecast start time, which is January 2020; use to represent the maximum forecast time length, which is 1 year.

[0021] Step 3: Secondary screening of spatial historical data According to the time resolution , screen the data from the historical dataset in Step 1 to form a training reference dataset. The training data coverage time range should at least include the data of the first 2 complete change cycles before the forecast start time . For example, when the forecast time resolution unit is "hour", it is necessary to make the sample coverage not less than the first 2 full-day cycles, that is, screen the sample data not less than 2 days before the forecast moment; when the forecast time resolution is "month", it is necessary to make the sample coverage not less than the first 2 full-year cycles, that is, screen the sample data not less than 2 years before the forecast moment.

[0022] In this example, the sound speed profile data measured in this spatial region from 2016 to 2020 for five years is used as the spatial historical sound speed dataset. The forecast start time is January 2020, and there are 4 full-year cycles of training reference data before this, meeting the training requirements.

[0023] The historical sound speed distribution dataset is expressed as: ; where i = 1~60 represents the 60 months from January 2016 to December 2020 arranged in chronological order, and is the sound speed distribution data corresponding to the corresponding time.

[0024] The i-th sound speed profile can be further expressed as: = 1,2,..., 58 where is the sound speed value, with the unit: m / s, and the subscript represents the th depth layer, the maximum value of is 58, that is, the maximum depth layer of the sound speed distribution is 58 layers, and " " represents the transpose of the vector.

[0025] Expanding and writing, the collected historical sound speed distribution dataset S is a sound speed value data matrix with 58 rows and 60 columns: ; After the collection of the historical sound velocity distribution data set is completed, proceed to step 4.

[0026] Step 4: (1) The data change time period is C = 12. The first 48 sound velocity distribution data in the historical sound velocity distribution data set S are used as the training data set, and the subsequent 12 sound velocity distribution data are used as the validation data set.

[0027] (2) Specific processing of the training set Divide the first 48 columns of the historical sound velocity data set S into the training data set T, and the last 12 columns of S as the validation data set E.

[0028] Training data set: ; Validation data set: ; Perform normalization processing on the training data set T to be .

[0029] Normalized training set: ; The normalized training data set is specifically divided into the training input set and the training label set , and adopt a training method with an interval of one period, that is, use 12 columns of data to predict the 13th column of data (that is, when the data from January to December 2016 is used as the training input, the data in January 2017 is used as the output of the training; when the data from February 2016 to January 2017 is used as the training input, the data in February 2017 is used as the output of the training), and so on.

[0030] Training input set: ; Training label set: ; After that, X and Y are further divided into , .

[0031] After the preprocessing of the sound velocity distribution sample data is completed, proceed to step 5.

[0032] Step 5: Construct a hybrid basis function KAN network model, and the model structure is as Figure 2 shown.

[0033] According to different depth stratifications, the ocean data distribution forecasting is performed layer by layer. For each depth layer, a hybrid basis function KAN network is constructed. For the hybrid basis function KAN network constructed for each depth layer, the prediction result is an output time series with a length of 12.

[0034] The number of basis functions, the number of basis function neurons, and the number of neurons in the fully connected layer are taken as 4, 256, and 256 respectively.

[0035] The basis functions are four basis functions: B-spline, Jacobi, Taylor, and Wavelet.

[0036] Step 6: Model training Before model training, first determine that the learning rate for each depth layer is = 0.0001 and the number of iterations is , representing the corresponding depth layer, j = 1, 2, …, 58; (1) Model training When training the network of the first depth layer, the input to the model is: = , and the corresponding model output is = ; Then calculate the value of MSE to update the model parameters. The training methods for the remaining layers are similar.

[0037] (2) Implementation of the pruning strategy First step, first perform a certain number of pre-training on the model In the model, the weight coefficients of the mixed basis function KAN layer and the parameters in the basis functions are all trainable. At the beginning of model training, first freeze the weight coefficients so that they cannot be trained, and only train the parameters in the basis functions. All weight coefficients are initialized to . First perform ten pre-training on the model.

[0038] At this time, the output of the mixed basis function KAN layer is ; Second step, perform pruning operation on the model After completing the pre-training, at this time, first freeze the trainable parameters in all basis functions, then unfreeze the weight coefficients so that they can be trained. Then train for another ten times, observe the weight coefficients of each basis function, find the largest coefficient, at this time the basis function corresponding to this coefficient is the best basis function, unfreeze its trainable parameters, set its weight coefficient to 1, and set other weight coefficients to 0.

[0039] Then continue training until the number of iterations is completed.

[0040] (3) Model verification For the training input set , take the data of the last time period of the sliding window one step to get new data, that is As the input of the model for prediction, it is used to predict the data with a time resolution in the future, that is, the sequence data with a length of 1. Then, it slides one step to obtain a new value, which is used as the input for predicting the next time resolution until a time period is predicted. The other deep layers are similar.

[0041] The prediction result of the first deep layer is a time series of sound speed values with a length of 12: ; The same applies to the other deep layers. Then, the finally predicted sound speed data set is : ; The predicted sound speed data set is de-normalized using the inverse operation of the original normalization operation, that is, the predicted data is restored to the scale range of the original data for easy comparison and analysis with the verification data.

[0042] The predicted result after de-normalization is as follows: ; After the de-normalization operation is completed, the root mean square error is calculated. If the value of RMSE is not less than the expected threshold, return to step 6, increase the number of iterations of the training model, and continue to train the model until the model converges, that is, RMSE is less than the expected threshold or there is no obvious change.

[0043] To verify the prediction accuracy of the model, the prediction results of the hybrid basis function KAN are compared with the prediction results of LSTM and the actual sound speed distribution. The intuitive comparison chart is as Figure 3 shown. To reflect the improvement of the pruning strategy on the model training efficiency, Table 1 gives the RMSE of the prediction errors and the training time cost of pruning and non-pruning. The average RMSE of the prediction errors of the hybrid basis function KAN is 0.71, and the average RMSE of the prediction errors of LSTM is 0.95. It can be concluded from Table 1 that the pruning strategy can reduce the training time cost of the model without reducing the prediction accuracy.

[0044] Table 1 Prediction errors and time costs Pruning No Pruning RMSE (m / s) 0.71 0.71 Time Overhead (s) 45 65

[0045] II. Predicting temperature Step 1: Extract the spatial historical ocean data set according to the data prediction location According to the data prediction location, the spatial coordinates are determined as 168.5°E longitude and 16.5°N latitude. The temperature profile data measured in 60 months from 2016 to 2020 in this spatial region are selected from the global Argo data set to form the spatial historical temperature data set as the reference data set of the model.

[0046] Step 2: Determination of Data Forecast Time Resolution, Forecast Start and End Times, and Maximum Time Length Use to represent the time resolution of the forecast task requirements, which is 1 month; use to represent the forecast start time, which is January 2020; use to represent the maximum forecast time length, which is 1 year.

[0047] Step 3: Secondary Screening of Spatial Historical Data According to the time resolution , screen the data from the historical dataset in Step 1 to form a training reference dataset. The training data coverage time range should include at least the data of the first 2 complete change cycles before the forecast start time . For example, when the forecast time resolution unit is "hour", the sample coverage should be no less than the first 2 full-day cycles, that is, screen the sample data no less than 2 days before the forecast moment; when the forecast time resolution is "month", the sample coverage should be no less than the first 2 full-year cycles, that is, screen the sample data no less than 2 years before the forecast moment.

[0048] In this example, the temperature profile data measured in this spatial region from 2016 to 2020 for five years is used as the spatial historical temperature dataset. The forecast start time is January 2020, and there are 4 full-year cycles of training reference data before this, meeting the training requirements.

[0049] The historical temperature distribution dataset is expressed as: ; where i = 1~60 represents the 60 months from January 2016 to December 2020 in chronological order for 5 years, and is the temperature distribution data corresponding to the corresponding time.

[0050] The i-th temperature profile can be further expressed as: = 1,2,...,58 where is the temperature value, with the unit: m / s, and the subscript represents the -th depth layer, the maximum value of is 58, that is, the maximum depth layer of the temperature distribution is 58 layers, and " " represents the transpose of the vector.

[0051] Expand and write the collected historical temperature distribution dataset S as a temperature value data matrix with 58 rows and 60 columns: ; After the historical temperature distribution dataset is collected, go to step 4.

[0052] Step 4: (1) The data change time period is C = 12. The first 48 temperature distribution data in the historical temperature distribution dataset S are used as the training dataset, and the last 12 temperature distribution data are used as the validation dataset.

[0053] (2) Specific processing of the training set Divide the first 48 columns of the historical temperature dataset S into the training dataset T, and the last 12 columns of S as the validation dataset E.

[0054] Training dataset: ; Validation dataset: ; Perform normalization processing on the training dataset T to .

[0055] Normalized training set: ; The normalized training dataset is specifically divided into the training input set and the training label set , and an alternating training method is adopted, that is, 12 columns of data are used to predict the 13th column of data (that is, when the data from January to December 2016 is used as the training input, the data for January 2017 is used as the training output; when the data from February 2016 to January 2017 is used as the training input, the data for February 2017 is used as the training output), and so on.

[0056] Training input set: ; Training label set: ; After that, X and Y are further divided into , .

[0057] After the preprocessing of the temperature distribution sample data is completed, go to step 5.

[0058] Step 5: Construct a hybrid basis function KAN network model According to different depth stratifications, the ocean data distribution forecast is carried out layer by layer. A hybrid basis function KAN network is constructed for each depth layer. For the hybrid basis function KAN network constructed for each depth layer, the prediction result is an output time series with a length of 12.

[0059] The number of basis functions, the number of basis function neurons, and the number of neurons in the fully connected layer are taken as 4, 256, and 256 respectively.

[0060] The basis functions are four basis functions: B-spline, Jacobi, Taylor, and Wavelet.

[0061] Step 6: Model training Before model training, first determine that the learning rate for each depth layer is = 0.0001 and the number of iterations is , where represents the corresponding depth layer, j = 1, 2, …, 58; (1) Model training When training the network of the first depth layer, the input to the model is: = , and the corresponding model output is = ; Then calculate the value of MSE to update the model parameters. The training method for each of the remaining layers is similar.

[0062] (2) Implementation of the pruning strategy First step, first perform a certain number of pre-training on the model In the model, the weight coefficients of the hybrid basis function KAN layer and the parameters in the basis functions are all trainable. At the beginning of model training, first freeze the weight coefficients so that they cannot be trained, and only train the parameters in the basis functions. All weight coefficients are initialized to . First perform ten pre-training on the model.

[0063] At this time, the output of the hybrid basis function KAN layer is ; Second step, perform pruning operation on the model After completing the pre-training, at this time, first freeze the trainable parameters in all basis functions, then unfreeze the weight coefficients so that they can be trained. Then train for another ten times, observe the weight coefficients of each basis function, find the largest coefficient, and at this time the basis function corresponding to this coefficient is the best basis function. Unfreeze its trainable parameters, set its weight coefficient to 1, and set other weight coefficients to 0.

[0064] Then continue training until the number of iterations is completed.

[0065] (3) Model verification For the training input set , take the data of the last time period of the sliding window one step to get new data, that is, As the input of the model for prediction, it is used to predict the data at the next time resolution, that is, the sequence data with a length of 1. Then, a sliding window is taken one step to obtain a new value as the input for predicting the next time resolution until a time period is predicted. The other deep layers are similar.

[0066] The prediction result of the first deep layer is a time series of temperature values with a length of 12: ; Similarly for the other deep layers, then the finally predicted temperature dataset is : ; Denormalize the predicted temperature dataset using the inverse operation of the original normalization operation, that is, restore the predicted data to the scale range of the original data for easy comparison and analysis with the validation data.

[0067] The prediction result after denormalization is as follows: ; After the denormalization operation is completed, calculate the root mean square error. If the value of RMSE is not less than the expected threshold, return to step 6, increase the number of iterations of training the model, and continue to train the model until the model converges, that is, RMSE is less than the expected threshold or there is no obvious change.

[0068] To verify the prediction accuracy of the model, compare the prediction results of the hybrid basis function KAN with the prediction results of LSTM and the actual temperature distribution. The intuitive comparison chart is as Figure 4 shown. To reflect the improvement of the pruning strategy on the model training efficiency, Table 2 gives the RMSE of the prediction errors and the training time cost of pruning and non - pruning. The average RMSE of the prediction errors of the hybrid basis function KAN is 0.27, and the average RMSE of the prediction errors of LSTM is 0.32. It can be seen from Table 2 that the pruning strategy can reduce the training time cost of the model without reducing the prediction accuracy.

[0069] Table 2 Prediction Errors and Time Costs Pruning No Pruning RMSE (m / s) 0.27 0.27 Time Overhead (s) 45 65 The above - described specific embodiments have further detailed the object, technical solution, and beneficial effects of the present disclosure. It should be understood that the above - described are only the specific embodiments of the present disclosure and do not limit the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for predicting the ocean sound speed distribution based on pruning hybrid basis function KAN, characterized in that, The method includes the following specific steps: S1: Extract the spatial historical ocean data set: Determine the spatial coordinates as longitude M° and latitude N° according to the data prediction location, and extract the ocean data at different historical times at this location to form a historical data set; S2: Determine the time resolution of data prediction, the start and end times of prediction, and the maximum time length: Use to represent the time resolution of the prediction task requirements, use to represent the start time of the prediction, and use to represent the maximum time length of the prediction; S3: Secondary screening of the spatial historical ocean dataset: According to the time resolution , screen the data form reference dataset from the historical datasets in S1; S4: Perform preprocessing on the reference data set; S5: Construct a hybrid basis function KAN network model, which includes an input layer, a hybrid basis function KAN layer, a fully connected layer, and an output layer. There are n basis functions in the hybrid basis function KAN layer, and the best basis function can be adaptively selected through a pruning strategy during the model training process. The number of neurons for the basis function is N, and the number of neurons for the fully connected layer is N. According to different depth stratifications, the ocean data distribution prediction is performed layer by layer. A hybrid basis function KAN network is constructed for each depth layer. For the hybrid basis function KAN network constructed for each depth layer, the prediction result is an output time series with a length of ; S6: Train the hybrid basis function KAN network model; S7: Model prediction: According to the task-specified forecast time resolution, forecast start time, and maximum forecast time length, determine the number of forecast iterations and the historical distribution reference data corresponding to the forecast start time. Using the data of two or more complete change cycles before the forecast start time as the training set of the model, the prediction result can be obtained. Then perform denormalization processing on the prediction result and output the forecast future distribution data.

2. The ocean sound speed distribution prediction method according to claim 1, wherein, In S3: The time range covered by the training data shall at least include the forecast start time Data for the first two complete change cycles. When the forecast time resolution unit is "hour", the sample coverage shall be no less than The first two full-day cycles, that is, no less than Sample data for the first two days before the screening forecast time; when the forecast time resolution is "month", the sample coverage shall be no less than The first two full-year cycles, that is, no less than Sample data for the first two years before the screening forecast time; Let the secondary screening historical data set be denoted as: ; Among them represents the i-th data profile, with a total of I. The i-th data profile is further represented as: ; where is the data value, and the subscript represents the th depth layer, represents the transpose of the vector; in formula (1), the S dataset is arranged in chronological order, and the time interval is consistent with the time resolution ; = 0, 1,..., J. In the formula, is the depth layer index label, and the maximum depth layer of the data distribution is dJ; The expansion of formula (1) gives the data set matrix: 。 3. The method for predicting the ocean sound speed distribution according to claim 1, characterized in that, The said S4 includes: S4-1: Divide the training set and the validation set Take the first I-C columns of the historical data set S as the training data set T, as shown in formula (4); the last C columns of S are used as the validation data set E, as shown in formula (5), where I-C≥C; C is the natural change cycle of the data; ; ; S4-2: Normalize the training dataset T to , where the normalization formula is as shown in Equation (6), and is the data after normalization, is the data before normalization, is the mean of the sequence data, is the standard deviation of the sequence data; at a certain depth layer is obtained from Equation (7), represents the sum of the data in the j-th row of matrix T from the first column to the last column; is obtained from Equation (8); then the normalized training dataset is specifically divided into a training input set and a training label set ; the training input set X is as shown in Equation (10), and the training label set Y is as shown in Equation (11); ; ; ; ; ; ; S4-3: Re-segment the training input set and the training label set according to the depth layer The training input set and the training label set are respectively divided into , ; , which are respectively represented by formulas (12) and (13); ; ; It represents the data of the j-th row of matrix X, that is, the data of the j-th deep layer. Each time a window of size C is taken, and the window step size is 1; It represents the data of the j-th row of matrix Y, that is, the data of the j-th deep layer. Each time a window of size 1 is taken, and the window step size is 1.

4. The method for predicting the ocean sound speed distribution according to claim 1, wherein The said S6 includes: S6-1: Before model training, first determine the learning rate of each deep layer as and the number of iterations as , where j represents the corresponding deep layer, and j = 1, 2, …, J; S6-2: During model training, the mixed basis function KAN network model of each deep layer is trained separately. The mixed basis function KAN represents the mixed basis function KAN network of the j-th deep layer. The input and output of the model are expressed as: ; Among them represents the output matrix of the model, expressed as: ; After that, calculate the mean square error MSE between the output value of the model and the training label set: ; Then use the value of MSE to update the parameters of the model; S6-3: Implement the pruning strategy; S6-4: Model Validation: Use the data of the next cycle of the last time cycle data in the training input set as the input of the model for prediction, to predict the data of the next time resolution in the future, that is, the sequence data with a length of 1, and then slide the window to obtain new values as the input for predicting the next time resolution until the prediction of one time cycle is completed.

5. The method for predicting the ocean sound speed distribution according to claim 4, wherein The said S6-3 includes: S6-31: First, pre-train the model. At the beginning of model training, freeze the weight coefficients so that they cannot be trained, and only train the parameters in the basis functions. All weight coefficients are initialized to be , pre-train the model. At this time, the output of the hybrid basis function KAN layer is: ; S6-32: Perform pruning operation on the model. First, freeze the trainable parameters in all basis functions, then unfreeze the weight coefficients so that they can be trained. Then, after training, observe the weight coefficients of each basis function, find the largest coefficient. At this time, the basis function corresponding to this coefficient is the best basis function. Unfreeze its trainable parameters, set its weight coefficient to 1, and set other weight coefficients to 0; then continue training until the number of iterations is completed.

6. The method for predicting the ocean sound speed distribution according to claim 4, wherein The said S6-4 includes: The prediction result of the first depth layer is a time series of length C, as shown in formula (17): ; The finally predicted dataset is , as shown in formula (18): ; Perform denormalization processing on the predicted data. The denormalization formula is: ; Prediction result after denormalization : ; Among them, represents the value predicted at the i-th temporal resolution of the j-th depth layer; After the denormalization operation is completed, calculate the root mean square error between the actual value and the predicted value: ; wherein represents the RMSE of the i-th predicted time resolution; if the value of the RMSE is not less than the expected threshold, return to S6, increase the number of iterations of the training model, and continue to train the model until the model converges, that is, the RMSE is less than the expected threshold or there is no obvious change.

Citation Information

Patent Citations

  • Sound velocity distribution forecasting method based on long short-term memory neural network

    CN117076893A

  • Hypersonic flight vehicle trajectory prediction method based on LSTM-KAN

    CN118964980A

  • Underground water level prediction method based on KAN network architecture

    CN119151065A

  • Att-KAN neural network model-based sound velocity distribution forecasting method

    CN119918581A

  • Marine environmental element prediction method based on STEOF-LSTM

    JP7175415B1