Tide prediction method and system based on ISSA-AC-CNN-BiLSTM model
By using the ISSA-AC-CNN-BiLSTM model in tidal prediction and optimizing parameters with improved sparrow search algorithm, the problems of parameter selection uncertainty and inefficient search efficiency in existing neural network models in tidal height prediction are solved, and more efficient and accurate tidal prediction is achieved.
Patent Information
- Application Number
- CN202510457565.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing neural network models have problems such as parameter selection uncertainty, inefficient search efficiency and early maturity convergence in tidal height prediction, and it is difficult to effectively deal with the complexity and nonlinearity of tidal phenomena.
The tide prediction method based on the ISSA-AC-CNN-BiLSTM model is adopted to extract spatiotemporal features through the AC module, combine with the BiLSTM model to learn timing correlation, and optimize the model parameters using the improved sparrow search algorithm.
It improves the accuracy and efficiency of tide prediction, overcomes the problem of hyperparameter optimization in the training process of neural network model, and enhances the model's feature extraction and generalization capabilities.
Smart Images

Figure CN119986862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of application of deep learning models in tidal height prediction, and more particularly to a tidal prediction method and system based on an ISSA-AC-CNN-BiLSTM model. Background Art
[0002] Predicting tidal water levels has always been a key factor in ensuring the entry and exit of ships in ports, and plays a vital role in the safe navigation of ships. With the development and improvement of basic digital models, many neural network models have been applied to the study of accurately quantifying tidal heights. However, there are some defects in the application of neural network models to the study of accurately quantifying tidal heights, which mainly stem from the limitations of neural network models and the complexity of tidal phenomena. For example, neural network models have the limitations of uncertainty in parameter selection, low search efficiency and premature convergence, and noise if the training data is incomplete and inaccurate; the complexity of tidal phenomena includes the fact that tidal phenomena are affected by multiple factors, tidal phenomena have significant temporal and spatial variability, and extreme weather events (such as storm surges, typhoons, etc.) may have a significant impact on tides. Therefore, when dealing with such highly complex and nonlinear data, the neural network model increases the complexity and computational complexity of the model, and due to the improper parameter selection of the neural network model itself and the problem of low search efficiency and premature convergence, the use of neural network models may affect the accuracy of tidal height prediction.
[0003] To ensure the safe navigation of ships in the sea, timely and accurate tidal data is of vital importance to improve the safety of ships. Therefore, studying an efficient tidal prediction method has become an urgent problem to be solved. Summary of the invention
[0004] In view of the technical problem that the prediction accuracy is low in the process of predicting tidal height based on a neural network model in the prior art, the present invention provides a tidal prediction method and system based on an ISSA-AC-CNN-BiLSTM model.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is: a tidal prediction method based on the ISSA-AC-CNN-BiLSTM model, the method comprising: Step 1: Build the AC-CNN-BiLSTM model based on the AC module; Step 2: The ISSA-AC-CNN-BiLSTM model is constructed by combining the improved sparrow search algorithm SSA with the AC-CNN-BiLSTM model; Step 3: Based on the tidal data, the tidal water level is predicted through the constructed ISSA-AC-CNN-BiLSTM model to obtain the prediction results.
[0006] Furthermore, the AC-CNN-BiLSTM model in step 1 includes an input layer, a fusion feature extraction layer, a BiLSTM layer, a fully connected layer, and a prediction output layer.
[0007] Furthermore, the input layer obtains the time series subsequence by the sliding window. Each time the sliding window slides, a matrix sequence is obtained. The tidal water level element after the maximum moment of each sliding is used as the mapping result, that is, y t+k = x t+k,1 ,Will(( X t-L ,…, X t-1 , X t ), y t+k ) as a batch and put into the Attention-CNN parallel module; The fusion feature extraction layer is an AC module that uses multi-layer and multi-scale convolution kernels to extract fine-grained temporal and spatial features. sigmoid The function assigns weights and uses the Tensorflow framework multiply Function fusion branch features; The BiLSTM layer is composed of multi-layer convolution and BiLSTM network. In this layer, multi-layer convolution kernels are used to extract fusion features, and then the fusion features are input into the CNNS-BiLSTM network to learn the timing information between subsequences, and the overfitting of the model is prevented by the Dropout layer. The output of the BiLSTM layer is used as the input vector of the fully connected layer; finally, the Dense function outputs the tidal prediction value.
[0008] Furthermore, the AC module in step 1 captures the correlation information between the context of the tidal sample data through the Attention branch and the CNN branch, and uses the attention mechanism to extract the timing information in the tidal sequence to train the model; uses multi-layer and multi-scale convolution kernels to extract fine-grained temporal features, and then uses Softmax The function assigns contribution weights to each moment.
[0009] Furthermore, the weighted form of the attention mechanism adopts the weighted form of multiplication attention, and the formula involved in the weighted summation process of the multiplication attention is as follows: (1); (2); (3); (4); in, s t , s t-1 Respectively t , t-1 The hidden state of the space corresponding to each moment; h j Indicates j The hidden vector at the moment contains the information of the entire input subsequence, but focuses on the j a moment; T represents the input time step; α t,j represents the attention allocation coefficient, whose probability sum is 1; e t,j Indicates at time t and j Attention scores between W e express t The weight matrix corresponding to the -1 moment.
[0010] Furthermore, in the CNN branch, a one-dimensional convolutional neural network is used to extract high-level abstract feature information of the feature vector, namely, spatial features. The convolution kernel moves along a certain direction. Assuming that the length of the convolution kernel is s , each time moving one square, the sequence length is L , then the output sequence length is L - s +1, and fill the borders to keep the output length consistent with the input length; The convolution process is expressed as formula (5): (5); In the formula, C is the output feature of the convolutional layer, Z is the input feature, f(·) is the activation function, ☉ is the matrix inner product operation, W is the weight matrix corresponding to the convolution kernel, b is the bias term.
[0011] Furthermore, the temporal weight allocation structure of the Attention branch includes: firstly, the temporal data is transposed, and after the transposition, each row of the feature matrix is the same feature, and these features are arranged according to the time sequence; then, the temporal change vector ( x t-L,k , x t-L+1,k , x t,3 ,…, xt,k ) is used for temporal feature extraction by one-dimensional convolution, and convolution operation is performed row by row, and the attention mechanism is used to extract temporal features. Softmax The function calculates the weight matrix, which is expressed as formula (7), and transposes the weight matrix; finally, the weight matrix and the spatial feature matrix of the CNN branch are multiplied by corresponding elements, which are expressed as formulas (6) and (8) to give temporal feature information to spatial features at different times; (6); (7); (8);
[0012] in, C CNN is the spatial feature matrix of the CNN branch, α t,k For each row of features Softmax The output time weight matrix, C Mix is the spatiotemporal fusion feature. T For transposition, For the matrix Hadamard Operation, can be a vector or matrix containing all features, and represents only a specific eigenvalue in the vector or matrix, D represents the set of input data, W CNN is the weight matrix in the convolutional neural network (CNN), b CNN is the bias term in the convolutional neural network, W t For time t The weight matrix at time, b t is the bias term at time t.
[0013] Furthermore, the improved SSA algorithm in step 2 includes initializing the sparrow population by using Latin hypercube sampling LHS, and the specific steps of the LHS are as follows: S201: determining the number of samples and dimension of the vector space; S202: generating m non-overlapping partitions in each dimension so that each interval has the same probability; S203: selecting random data points in each partition to establish a sampling matrix; S204: randomly selecting a number in each column of the sampling matrix to form a vector.
[0014] Furthermore, the improved SSA algorithm in step 2 includes: adding a sine function dynamic adaptive weight to the sparrow finder position update formula to optimize the local exploration problem of the model. The finder position update formula of the improved SSA algorithm is (9) and (10): (9); (10); In formula (9), Indicated in t In the +1 iteration, i The sparrow's j The position update value of each dimension; Indicates i A sparrow in j The current position in dimensions; represents the attenuation factor, where i is the number of iterations, a is a constant, T is the total number of iterations. This decay factor is used to control the amplitude of position updates to gradually decrease as the number of iterations increases. ST Represents a threshold value used to compare R 2. Compare to determine the location update strategy; w is a weight or parameter used to adjust the moving step size of individual sparrows in the search space; R 2 is a random number in the range [0, 1], which is used to determine whether the sparrow individual adopts a local search strategy or a global search strategy in the current iteration; Q Used in R 2 ≥ ST Adjust the position of individual sparrows; L Represents a random step size or perturbation, which is used to introduce randomness in local search to prevent the algorithm from falling into the local optimum; In formula (10), represents the independent variable of the sine function, where t is the current iteration number, T is the total number of iterations. This expression ensures that w Changes periodically during iterations; b’ is a correction parameter used to adjust w The value range is set to 1 to ensure w The value of is in the interval [0,1].
[0015] Furthermore, the step 2 further comprises: using Rosenbrock Function simulation experiments are used to verify the convergence effect of the improved SSA algorithm. Rosenbrock The function expression is formula (11): (11).
[0016] Furthermore, step 2 further includes: using BiLSTM to train separately RMSE , MAE , R ² is the optimization range of the evaluation criteria and parameters.
[0017] The present application also provides a tide prediction system based on the ISSA-AC-CNN-BiLSTM model, the system comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the system is triggered to execute the above-mentioned tide prediction method based on the ISSA-AC-CNN-BiLSTM model.
[0018] Compared with the prior art, the beneficial effects of this application are: Through the novel AC-CNN-BiLSTM tide prediction model, AC (Attention-CNN fusion module) extracts spatiotemporal features from the time dimension and feature dimension respectively; then the spatiotemporal features are fused through Hadamard operation, thereby enhancing the ability to capture global information and time series information; finally, the CNN-BiLSTM model learns the time series association between the fused features to map the tide height. The present invention also uses the sparrow search algorithm (SSA) to optimize the parameters of the AC-CNN-BiLSTM model, and constructs the ISSA-AC-CNN-BiLSTM tide prediction model, which overcomes the difficulty that the hyperparameters need to be repeatedly experimented during the training process of the neural network model. The ISSA algorithm uses the Latin hypercube sampling method (LHS) to improve the population initialization method of the SSA algorithm, so that the random numbers are allocated in different regions and with equal probability in the solution domain; at the same time, the sine function adaptive weight is added to the sparrow finder position, so that the sparrow population changes the sparrow search range at different iteration stages to avoid falling into the local optimum. The ISSA-AC-CNN-BiLSTM model can obtain better feature extraction capabilities and prediction effects, and the prediction accuracy is also greatly improved. The model has high accuracy and strong generalization ability when predicting tide heights. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of the one-dimensional convolution process structure provided by an embodiment of the present invention; Figure 2 A schematic diagram of the AC model structure provided by an embodiment of the present invention; Figure 3 ISSA algorithm flow chart provided for the embodiment of the present invention; Figure 4 A comparison chart of Random initialization and Latin Hypercube sampling initialization provided in an embodiment of the present invention; Figure 5 A diagram showing the change of a sine function with the number of iterations provided in an embodiment of the present invention; Figure 6 A relationship diagram between the number of Hidden-nodes and Epochs provided in an embodiment of the present invention; Figure 7 This is a flow chart of the ISSA-AC-CNN-BiLSTM model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention, so as to more clearly understand the purposes, features and advantages of the present invention. It should be understood that the embodiments shown in the drawings are not limitations on the scope of the present invention, but are only intended to illustrate the essential spirit of the technical solutions of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention.
[0021] Unless the context requires otherwise, throughout the specification and claims, the word "comprise" and variations such as "include" and "have" should be construed in an open, inclusive sense, ie, should be interpreted as "including, but not limited to."
[0022] References throughout the specification to "one embodiment" or "some embodiments" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearance of "in one embodiment" or "in some embodiments" in various places throughout the specification need not all refer to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any manner in one or more embodiments.
[0023] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be noted that the term "or" is generally employed in its sense including "and / or" unless the context clearly dictates otherwise.
[0024] The implementation details of this embodiment are described in detail below with reference to the accompanying drawings. The following content is only provided for easy understanding of the implementation details and is not necessary for implementing this solution.
[0025] In the process of studying tidal data based on neural networks in the prior art, the training method of the attention mechanism is relatively simple. In tidal data with strong periodicity and more high-frequency data, although the data changes in each time interval are very small, there will be certain confusion in the prediction of the deep learning model. How to accurately locate the critical moment and fuse the characteristics of the data itself while giving special weights to the critical moment to avoid confusion in the model is more important. This application first built an AC-CNN-BiLSTM model, which is composed of an AC (Attention-CNN) fusion module and a CNN-BiLSTM model, in which the AC module solves the problem of data feature extraction at critical moments by fusing features of different data dimensions. Among them, CNN (Convolutional Neural Networks) is a convolutional neural network; BiLSTM (Bidirectional Long Short-term Memory) is a bidirectional long and short-term memory network model.
[0026] Secondly, in the process of training the neural network model, the model hyperparameters are often determined by observing the loss function curve. Relying on experience and repeated experiments, the parameter optimization algorithm is an effective solution for finding the optimal model. In response to the problem of optimizing the model hyperparameters, this application builds the ISSA-AC-CNN-BiLSTM model. First, the Latin Hypercube Sampling LHS (Latin Hypercube Sampling) method is used to initialize the average probability of the sparrow population of the Sparrow Search Algorithm SSA (Sparrow Search Algorithm). Then, according to the different positioning of the sparrow roles, the sine function adaptive weight is added to the discoverer position update, so that the sparrow population can dynamically search for food. Finally, the improved sparrow search algorithm (ISSA) is used to optimize the number of hidden nodes in the neural network to explore the optimal model.
[0027] In summary, this application combines the attention mechanism, the BiLSTM neural network model and the swarm intelligence algorithm to construct a new tidal prediction model. The construction of this new tidal prediction model is divided into two stages. In the first stage, the AC-CNN-BiLSTM model is first constructed, in which AC extracts spatiotemporal features from the time dimension and the feature dimension respectively; then the spatiotemporal features are fused through the Hadamard operation to enhance the ability to capture global information and timing information; finally, the CNN-BiLSTM model is used to learn the temporal association between the fused features to map the tidal height. In the second stage, the improved sparrow search algorithm (SSA) is used to optimize the parameters of the AC-CNN-BiLSTM model to construct the tidal prediction model ISSA-AC-CNN-BiLSTM.
[0028] The first stage of building a new tidal prediction model is to build an AC-CNN-BiLSTM model as follows: The AC-CNN-BiLSTM model consists of an input layer, a fusion feature extraction layer, a BiLSTM layer, a fully connected layer, and a prediction output layer.
[0029] Among them, (1) the input layer is to obtain the time series subsequence by sliding window. Each time the sliding window slides, a matrix sequence is obtained. The tidal water level element after the maximum moment of each sliding is used as the mapping result, that is, y t+k = x t+k,1 , Will(( X t-L ,…, X t-1 , X t ), y t+k ) as a batch and put into the Attention-CNN parallel module.
[0030] (2) The fusion feature extraction layer is an AC module that uses multi-layer and multi-scale convolution kernels to extract fine-grained temporal and spatial features. sigmoid The function assigns weights and uses the Tensorflow framework multiply Function fusion branch features.
[0031] (3) The BiLSTM layer is composed of multi-layer convolution and BiLSTM network. In this layer, multi-layer convolution kernels are used to extract the fusion features. The fusion features are then input into the CNNS-BiLSTM network to learn the timing information between subsequences. The Dropout layer is used to prevent the overfitting of the model. The output of the BiLSTM layer is used as the input vector of the fully connected layer. Finally, Dense The function outputs the tide prediction value.
[0032] During the training process of the AC-CNN-BiLSTM model, the above parameters are all adjusted through the model's error back propagation algorithm to adjust the weight parameters so that the deviation is minimized.
[0033] The AC module of the fusion feature extraction layer consists of: The AC module captures the correlation information between the context of tidal sample data through the Attention branch and the CNN branch, and uses the attention mechanism to extract the temporal information in the tidal sequence. Considering the characteristics of the attention mechanism and the data characteristics, a soft attention mechanism is used to train the model. Multi-layer and multi-scale convolution kernels are used to extract fine-grained temporal features, and then SoftmaxThe function assigns contribution weights to each moment, so that the model pays attention to the contribution of different moments to the model prediction during training, and highlights the periodic influence of tidal data. In addition, the weighted form of the attention mechanism is divided into different alignment methods. Since the feature dimension may be low, multiplicative attention is used. The mathematical representation of multiplicative attention involves formulas (1), (2), (3), and (4), where s t , s t-1 Respectively t , t-1 The hidden state of the space corresponding to each moment; h j Indicates j The hidden vector at the moment contains the information of the entire input subsequence, but focuses on the j a moment; e t,j Indicates at time t and j Attention scores between T represents the input time step; α t,j represents the attention allocation coefficient, whose probability sum is 1, W e express t The weight matrix corresponding to the -1 moment: (1); (2); (3); (4); In the CNN branch, a one-dimensional convolutional neural network is used to extract high-level abstract feature information of the feature vector, namely spatial features. The purpose is to enhance the spatial feature learning ability of the model, so as to deeply learn the hidden associations in the data. The convolution kernel moves along a certain direction. Assuming that the length of the convolution kernel is s , each time moving one square (step_size=1), the sequence length is L , then the output sequence length is L - s +1, and fill the borders to keep the output length consistent with the input length, such as Figure 1 As shown. The convolution process expression is as shown in (5), C is the output feature of the convolutional layer, Z is the input feature, f(·) is the activation function, ☉ is the matrix inner product operation, W is the weight matrix corresponding to the convolution kernel, b is the bias term.
[0034] (5) AC module such as Figure 2 As shown, the left side is the time series data of a single batch taken out, the horizontal axis represents features (Features), and the vertical axis represents the input step (Time_step). Figure 3 The upper branch is a CNN structure, which transforms the input feature vector ( x t,1 , x t,2 , x t,3 ,…, x t,n ) are respectively put into one-dimensional convolution for spatial feature extraction. At this time, the ordinate of the extracted spatial feature layer represents Time_step, and the abscissa represents Features.
[0035] Figure 2 The lower half of the branch is the Attention time series weight distribution structure. First, the time series data is transposed according to Figure 2 It can be seen that after transposition, each row of the feature matrix is the same feature. These features are arranged in chronological order, and then the time series change vector of each feature ( x t-L,k , x t-L+1,k , x t,3 ,…, x t,k ) One-dimensional convolution is used to extract time features. The vertical axis of the time feature represents Features, and the horizontal axis represents Time_step. Then the convolution operation is performed row by row, and the time feature is extracted through the Attention mechanism. Softmax The function calculates the weight matrix. In order to make the ordinate and abscissa of the weight matrix consistent with the feature layer of the CNN branch, the transposition operation is performed again. Next, the weight matrix of the Attention temporal weight allocation structure and the spatial feature matrix of the CNN branch are multiplied by the corresponding elements. The purpose of this is to give temporal feature information to spatial features at different times, so that moments with greater contributions are assigned greater weights, enhance the temporal attention of different spatial features, suppress the influence of spatial features at unimportant moments, and more effectively obtain the context information of the subsequence. The specific expressions are as follows (6), (7), and (8). C CNN is the spatial feature matrix of the CNN branch, α t,k For each row of features Softmax The output time weight matrix,C Mix is the spatiotemporal fusion feature. T For transposition, For the matrix Hadamard Operation, can be a vector or matrix containing all features, and represents only a specific eigenvalue in the vector or matrix, D represents the set of input data, W CNN is the weight matrix in the convolutional neural network (CNN), b CNN is the bias term in the convolutional neural network, W t For time t The weight matrix at time, b t is the bias term at time t.
[0036] (6); (7); (8); The above is the first stage of building a new tide prediction model. Based on convolutional neural networks and attention mechanisms, a fused convolution model AC-CNN-BiLSTM suitable for tide prediction tasks is proposed. Next, based on the sparrow search algorithm, an ISSA-AC-CNN-BiLSTM combined model is constructed, which is the second stage of building a new tide prediction model. This second stage aims at the hyperparameter optimization problem of the AC-CNN-BiLSTM model and improves the sparrow search algorithm SSA. The Latin hypercube sampling LHS method is used to initialize the sparrow population, so that the sparrow individuals are more evenly distributed in the solution domain; in addition, sinusoidal function adaptive weights are added to the process of sparrow individuals searching for things, so that sparrow individuals can search for food dynamically, avoiding sparrow individuals from falling into local optimality when searching for food. The ISSA algorithm process is as follows Figure 3 shown.
[0037] Population initialization is the initial state of the swarm intelligence optimization algorithm, which dominates the convergence speed and optimization accuracy of the optimization algorithm. When there are fewer samples or more optimization parameters, the distribution of random numbers will have greater randomness, and the traditional method of initializing the population cannot achieve good results. Therefore, this application uses Latin Hypercube Sampling (LHS) to initialize the population.
[0038] Taking m samples in an n-dimensional vector space as an example, the specific steps of LHS are as follows: (1) Determine the number of samples and the dimension of the vector space. (2) Generate m non-overlapping partitions in each dimension so that each interval has the same probability. (3) Select random data points in each partition to establish a sampling matrix. (4) Randomly select a number in each column of the sampling matrix to form a vector. In the sparrow search algorithm, the Latin hypercube sampling method is used to initialize the sparrow population. m represents the number of sparrow populations and can be manually customized. n represents the variable parameter dimension, which is determined by the number of parameters that the model needs to optimize. Figure 4 As shown in the figure, (a) is a sparrow search algorithm generated based on a random function. It can be found that the randomly generated population is not evenly distributed. In the initial population, there are frequent clustering and bunching phenomena, which will cause the sparrow search to fall into the local optimum. (b) is that LHS forms a population with a wider distribution range and more uniformity, and the probability of obtaining a solution with good diversity and convergence is higher.
[0039] In some embodiments, the sparrow search algorithm, like other common swarm intelligence optimization algorithms, has the problem of being easily trapped in the local optimum. In the later stage of the traditional SSA algorithm iteration, the positions of the three sparrows will be updated in a small range near the optimal point, and the position update in a small range is prone to stagnation. Based on this problem, the present application also proposes an improved SSA algorithm, namely the ISSA algorithm, which adds a sine function dynamic adaptive weight to the sparrow finder position update formula to optimize the local exploration problem of the model. The ISSA finder position update formula is (9) and (10).
[0040] (9); (10); In formula (9), Indicated in t In the +1 iteration, i The sparrow's j The position update value of each dimension; Indicates i A sparrow in j The current position in dimensions; represents the attenuation factor, where i is the number of iterations, a is a constant, T is the total number of iterations. This decay factor is used to control the amplitude of position updates to gradually decrease as the number of iterations increases. ST Represents a threshold value used to compare R 2. Compare to determine the location update strategy; wis a weight or parameter used to adjust the moving step size of individual sparrows in the search space; R 2 is a random number in the range [0, 1], which is used to determine whether the sparrow individual adopts a local search strategy or a global search strategy in the current iteration; Q Used in R 2 ≥ ST Adjust the position of individual sparrows; L Represents a random step size or perturbation, which is used to introduce randomness in local search to prevent the algorithm from falling into the local optimum; In formula (10), represents the independent variable of the sine function, where t is the current iteration number, T is the total number of iterations. This expression ensures that w Changes periodically during iterations; b’ is a correction parameter used to adjust w The value range is set to 1 to ensure w The value of is in the interval [0,1].
[0041] like Figure 5 It can be seen that w As the number of iterations is constantly changing, the sine function will w The range is controlled within the value range of [-1,0], and the parameters are corrected b’ adjust w This application will modify the parameter b’ Set to 1, w The range is limited to the interval [0,1], giving the discoverer a greater weight in the early stage of algorithm iteration w , which is conducive to global search. In the later stage of algorithm search, w Slow descent, ample time for local exploration, and due to w The value decreases slightly, and a relatively large weight can be obtained in the later stage of iteration, thereby speeding up the local exploration and, to a certain extent, speeding up the overall convergence of the algorithm.
[0042] In some embodiments, using Rosenbrock Function simulation experiments are carried out to verify the convergence effect of the improved ISSA algorithm. When it is a binary function, Rosenbrock The function expression is as shown in formula (11); (11).
[0043] This application is for RosenbrockThe function performs independent fitness convergence training. Within 500 iterations, SSA reaches convergence at 27 iterations, and the converged fitness value accuracy reaches 10e-7, while ISSA reaches convergence at 41 iterations, and the converged fitness value accuracy reaches 10e-23. The fitness convergence accuracy is much higher than that of SSA, which shows that the improved ISSA algorithm is much higher than the ordinary SSA algorithm in terms of convergence fitness accuracy.
[0044] Although it has been verified that the ISSA search algorithm has a higher optimization effect than the SSA search algorithm, considering the time efficiency and model accuracy of ISSA algorithm parameter optimization, a too large optimization range will lead to a long model training time, and a too small optimization range will result in a loss of model accuracy, so it is necessary to limit the optimization range of the parameters.
[0045] In the CNN-BiLSTM model, this application uses three BiLSTM layers to deeply extract time series sample features. The BiLSTM model contains many hyperparameters, among which the hyperparameters with the greatest impact on the accuracy of tide prediction are: the number of hidden units in the first layer G 1. Number of hidden units in the second layer G 2. Number of hidden units in the third layer G 3. The batch size of each training will have a great impact on the accuracy of the network model and the training speed. RMSE , MAE , R ² is the evaluation criterion for the number of hidden units G i , batch_size, these four parameters limit the optimization range. RMSE , MAE The smaller the result, the higher the accuracy of tidal level prediction. R The larger the value, the better the model is, that is, the better the model fit is.
[0046] like Figure 6 As shown in the figure, (a) is the relationship between the loss function curve and the hidden unit (Hidden_nodes). It can be seen that with the increase of training times and the increase of the number of Hidden_nodes, the loss function curve converges faster and the fitting degree becomes smoother. (b)(c)(d) are the relationship between the number of Hidden_nodes and RMSE , MAE , R ², it can be seen that as the number of Hidden_nodes increases, the histogram shows a V-shaped trend. RMSE , MAEWhen the number is 5 and 100, two evaluation indicators are low. R ² is higher, so these two Hidden_nodes values are better during the training process. However, in order to ensure the optimization speed of SSA and the training speed of the model, this application will abandon the range search when the number of Hidden_nodes is 100. In order to ensure that the model accurately searches for the position of the optimal parameters, the number of Hidden_nodes is set to a step-like shape, and the number of hidden units in the first layer is G The value range of 1 is set to [10,80], and the number of hidden units in the second layer G 2 is set to [30,70], the number of hidden units in the third layer G 3 is set to [40,60]. The batch_size cannot be too large or too small, so the most commonly used in actual projects is mini-batch, which is usually set to dozens or hundreds of batches, so the batch_size range is limited to [16,64].
[0047] The tide prediction model ISSA-AC-CNN-BiLSTM combines the improved SSA algorithm and the neural network model AC-CNN-BiLSTM for training. The training process first initializes the sparrow population, and then uses the CNN-BiLSTM model to calculate its fitness value. The fitness function is set to the root mean square error (RMSE) during the network model training process. Each training aims to find a position that makes the RMSE lower. Through the change of fitness value, the positions of sparrow finders, followers, and guards are updated until the training termination condition is met. Finally, the neural network model with the optimal parameters is used to predict the tide level, and the model is tested on the validation set.
[0048] like Figure 7 The figure shows the flow chart of the tide prediction model. Based on the ISSA-AC-CNN-BiLSTM model, the tide level is predicted on the original tide data processed by the data processing module to obtain the prediction results.
[0049] The present invention also provides a tidal prediction system based on the ISSA-AC-CNN-BiLSTM model, the system comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the system is triggered to execute the aforementioned tidal prediction method based on the ISSA-AC-CNN-BiLSTM model.
[0050] In summary, the ISSA-AC-CNN-BiLSTM model provided by the present invention is based on the effectiveness of the AC module, and an AC-CNN-BiLSTM model is built. The hyperparameters are optimized in combination with the sparrow search algorithm, and the basic sparrow search algorithm is improved. The initial population of sparrows is initialized with equal probability using Latin hypercube sampling, so that the sparrows are distributed more evenly. At the same time, the sine function adaptive weight formula is added. As the number of iterations increases, the sparrow individuals are dynamically searched at different iteration stages. Finally, the Rosenbrock function is used to verify the search ability of the improved sparrow search algorithm. On this basis, the ISSA-AC-CNN-BiLSTM model is constructed, which can obtain better feature extraction ability and prediction effect, and has higher accuracy and stronger generalization ability in tidal water level prediction.
[0051] It should be noted that the system of the present invention may also include one or more of the following components: a processor and a memory.
[0052] Optionally, the processor uses various interfaces and lines to connect various parts within the entire autonomous operation equipment, and executes various functions of the autonomous operation equipment and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processor (NPU). Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content that needs to be displayed on the touch display; and the NPU is used to implement artificial intelligence (AI) functions.
[0053] The memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area may store data created according to the use of the autonomous operation device, etc.
[0054] The present disclosure also provides a computer-readable storage medium storing a computer program, wherein the computer program is used to be executed by a processor to implement the positioning control method of the autonomous working equipment as described in the above embodiment.
[0055] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention or any person skilled in the art who is familiar with the present invention may easily think of changes or substitutions within the technical scope disclosed by the present invention, and shall be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A tidal prediction method based on ISSA-AC-CNN-BiLSTM model, characterized in that: The method includes: Step 1: Build the AC-CNN-BiLSTM model based on the AC module; Step 2: The ISSA-AC-CNN-BiLSTM model is constructed by combining the improved sparrow search algorithm SSA with the AC-CNN-BiLSTM model; Step 3: Based on the tidal data, the tidal water level is predicted through the constructed ISSA-AC-CNN-BiLSTM model to obtain the prediction results.
2. The method according to claim 1, characterized in that The AC-CNN-BiLSTM model in step 1 includes an input layer, a fusion feature extraction layer, a BiLSTM layer, a fully connected layer, and a prediction output layer; the input layer is a time series subsequence obtained by a sliding window, and each sliding of the sliding window will obtain a matrix sequence, and the tidal water level element after each maximum sliding moment is used as the mapping result, that is, y t+k = x t+k,1 , Will(( X t-L ,…, X t-1 , X t ), y t+k ) as a batch and put into the Attention-CNN parallel module; The fusion feature extraction layer is an AC module that uses multi-layer and multi-scale convolution kernels to extract fine-grained temporal and spatial features. sigmoid The function assigns weights and uses the Tensorflow framework multiply Function fusion branch features; The BiLSTM layer is composed of multi-layer convolution and BiLSTM network. In this layer, multi-layer convolution kernels are used to extract fusion features, and then the fusion features are input into the CNNS-BiLSTM network to learn the timing information between subsequences, and the overfitting of the model is prevented by the Dropout layer. The output of the BiLSTM layer is used as the input vector of the fully connected layer; finally, the Dense function outputs the tidal prediction value.
3. The method according to claim 1, characterized in that The AC module in step 1 captures the correlation information between the context of the tidal sample data through the Attention branch and the CNN branch, and uses the attention mechanism to extract the timing information in the tidal sequence to train the model; uses multi-layer and multi-scale convolution kernels to extract fine-grained temporal features, and then uses Softmax The function assigns contribution weights to each moment.
4. The method according to claim 3, characterized in that The weighted form of the attention mechanism adopts the weighted form of multiplication attention. The formula involved in the weighted summation process of multiplication attention is as follows: (1); (2); (3); (4); in, Respectively t , t-1 The hidden state of the space corresponding to each moment; h j Indicates j The hidden vector at the moment contains the information of the entire input subsequence, but focuses on the j a moment; T represents the input time step; α t,j represents the attention allocation coefficient, whose probability sum is 1; e t,j Indicates at time t and j Attention scores between W e express t The weight matrix corresponding to the -1 moment.
5. The method according to claim 4, characterized in that In the CNN branch, a one-dimensional convolutional neural network is used to extract high-level abstract feature information of the feature vector, namely, spatial features. The convolution kernel moves along a certain direction. Assuming that the length of the convolution kernel is s , each time moving one square, the sequence length is L , then the output sequence length is L - s +1, and fill the borders to keep the output length consistent with the input length; The convolution process is expressed as formula (5): (5); In the formula, C is the output feature of the convolutional layer, Z is the input feature, f(·) is the activation function, ☉ is the matrix inner product operation, W is the weight matrix corresponding to the convolution kernel, b is the bias term.
6. The method according to claim 5, characterized in that The temporal weight allocation structure of the Attention branch includes: firstly, the temporal data is transposed, and after the transposition, each row of the feature matrix is the same feature, and these features are arranged in chronological order; then, the temporal change vector ( x t-L,k , x t-L+1,k , x t,3 ,…, x t,k ) is used for one-dimensional convolution to extract temporal features, and convolution is performed row by row, and the attention mechanism is used to extract temporal features. Softmax The function calculates the weight matrix, which is expressed as formula (7), and transposes the weight matrix; finally, the weight matrix and the spatial feature matrix of the CNN branch are multiplied by corresponding elements, which are expressed as formulas (6) and (8) to give temporal feature information to spatial features at different times; (6); (7); (8); in, C CNN is the spatial feature matrix of the CNN branch, α t,k For each row of features Softmax The output time weight matrix, C Mix is the spatiotemporal fusion feature. T For transposition, For the matrix Hadamard Operations, can be a vector or matrix containing all features, and represents only a specific eigenvalue in the vector or matrix, D represents the set of input data, W CNN is the weight matrix in the convolutional neural network (CNN), b CNN is the bias term in the convolutional neural network, W t For time t The weight matrix at time, b t is the bias term at time t.
7. The method according to claim 6, characterized in that The improved SSA algorithm in step 2 includes initializing the sparrow population by using Latin hypercube sampling LHS, and the specific steps of LHS are as follows: S201: determining the number of samples and dimension of the vector space; S202: generating m non-overlapping partitions in each dimension so that each interval has the same probability; S203: selecting random data points in each partition to establish a sampling matrix; S204: randomly selecting a number in each column of the sampling matrix to form a vector.
8. The method according to claim 7, characterized in that The improved SSA algorithm in step 2 includes: adding a sine function dynamic adaptive weight to the sparrow finder position update formula to optimize the local exploration problem of the model. The finder position update formula of the improved SSA algorithm is (9) and (10): (9); (10); In formula (9), Indicated in t In the +1 iteration, i The first sparrow j The position update value of each dimension; Indicates i A sparrow in the j The current position in dimensions; represents the attenuation factor, where i is the number of iterations, a is a constant, T is the total number of iterations. This decay factor is used to control the amplitude of position updates to gradually decrease as the number of iterations increases. ST Represents a threshold value used to compare R 2. Compare to determine the location update strategy; w is a weight or parameter used to adjust the moving step size of individual sparrows in the search space; R 2 is a random number in the range [0, 1], which is used to determine whether the sparrow individual adopts a local search strategy or a global search strategy in the current iteration; Q Used in R 2 ≥ ST Adjust the position of individual sparrows; L Represents a random step size or perturbation, which is used to introduce randomness in local search to prevent the algorithm from falling into the local optimum; In formula (10), represents the independent variable of the sine function, where t is the current iteration number, T is the total number of iterations. This expression ensures that w Changes periodically during iterations; b’ is a correction parameter used to adjust w The value range is set to 1 to ensure w The value of is in the interval [0,1].
9. The method according to claim 1, characterized in that: The step 2 further comprises: using Rosenbrock Function simulation experiments are used to verify the convergence effect of the improved SSA algorithm. Rosenbrock The function expression is formula (11): (11); The step 2 further comprises: using BiLSTM to train separately RMSE , MAE , R ² is the optimization range of the evaluation criteria and parameters.
10. A tidal prediction system based on the ISSA-AC-CNN-BiLSTM model, the system comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein: When the computer program instructions are executed by the processor, the system is triggered to execute the tide prediction method based on the ISSA-AC-CNN-BiLSTM model described in any one of claims 1 to 9.
Citation Information
Patent Citations
CNN-BiLSTM-ATTENTION-based tide level prediction method and system
CN118656643A
Analyzing an inference of a machine learning predictor
US20250094811A1