A sea wave lightweight prediction method based on knowledge distillation
By constructing a teacher-student model with dual inputs of "large sea area + small sea area" and utilizing a knowledge distillation mechanism, the accuracy and stability issues of wave forecasting technology under computing power-constrained scenarios were resolved, achieving lightweight and real-time wave forecasting.
Patent Information
- Application Number
- CN202511339847.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing wave forecasting technologies face contradictions when implemented in engineering, such as the constraints of accuracy and computing power, timeliness and deployability, cross-scale information and local forecasting, and inconsistent definitions of lightweighting. These make it difficult to achieve real-time and stable wave forecasting in computing-limited scenarios such as shipborne and buoy-based systems.
We construct a teacher model and a student model with dual inputs of "large sea area + small sea area". Through a knowledge distillation mechanism, we transfer large-scale features and surge propagation capabilities to the student model, so that it can maintain prediction accuracy and stability under a lightweight structure and meet the real-time application requirements of edge devices.
It significantly improves the efficiency and accuracy of wave forecasting. The student model can inherit the cross-scale prediction performance of the teacher model without directly acquiring data from the vast ocean area, meeting the real-time operation requirements of shipborne, buoy and other equipment.
Smart Images

Figure CN120832958B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of marine environment prediction, in particular to a sea wave lightweight prediction method based on knowledge distillation. BACKGROUND
[0002] With the in-depth application of numerical models and artificial intelligence in the field of sea wave prediction, the typical practice of multi-source collaborative technology system in the global range includes the third generation wave model based on the spectral energy balance equation for regional / global scale prediction, combined with buoy, radar, satellite and other observation data for verification or correction, and the use of deep learning model for rapid prediction of key elements. However, the current intelligent sea wave prediction technology faces significant contradictions in engineering landing: first, the contradiction between precision and computing power. The existing method pays more attention to the improvement of index precision in the research level, and the model parameter scale and the calculation cost continue to grow, the training cost is high, the reasoning time is long, and it is difficult to meet the real-time prediction demand of shipborne, buoy, nearshore station and other scenes with limited computing power and energy consumption; second, the contradiction between timeliness and deployability. The sea operation, route planning and nearshore disaster prevention and warning require minute-level or even shorter time delay update frequency, but complex models often rely on special GPU or cloud resources, and are difficult to run stably in offline and weak network environment; third, the contradiction between cross-scale information and local prediction. The sea wave process has obvious multi-scale characteristics, and the swell can cross the basin and have a lasting impact on the local wave field. Many existing lightweight models for local prediction reduce complexity by cutting input range or reducing resolution, which easily causes the loss of external energy and transmission path information, and the prediction credibility is insufficient in complex boundaries and extreme sea conditions; fourth, the definition and evaluation of lightweight are not uniform, and the high demand of intelligent model for computing power in training and reasoning is ignored.
[0003] The prior art directly uses a small model, which is trained on local data by a shallow CNN, LSTM or spatio-temporal hybrid network to reduce single forward calculation; although such methods are faster than numerical mode inference, they are still not light enough compared to the needs of end-side deployment, and due to the lack of expression ability for large-scale background, they are not sufficient for generalization across sea areas and seasons; the complexity is reduced by downsampling, cropping or simplifying the feature set of the input field, but this sacrifices the information integrity of the spectral shape and propagation structure, and the error is significant in the low-frequency long-wave and multi-peak spectral scenarios; the inference is completed on the cloud side to avoid the algorithm power bottleneck on the end side, but in the weak network on the sea, offline and rigid demand for "on-site decision" scenarios. The existing methods pay too much attention to prediction accuracy and ignore the training cost and inference efficiency. The current research on sea wave prediction technology mainly focuses on improving prediction accuracy, and the model size is increasing, which requires a large amount of computing power to support the inference stage. Such methods can run in a scientific research environment, but they are difficult to deploy in real time on end-side devices with limited computing power such as ships and buoys, and cannot meet the demand for rapid prediction. Some existing lightweight marine environment prediction methods only use smaller neural network models, which reduce the computational load compared to traditional numerical models, but the network structure is not designed for resource-constrained scenarios on the end side, and the number of parameters, memory usage and energy consumption are still high, making it difficult to achieve the true goal of lightweight in practical applications. Existing lightweight solutions often reduce model complexity by cropping input ranges or reducing resolution, but this approach can easily lead to the loss of large-scale physical background information such as distant swell and cross-basin propagation, resulting in a significant decrease in prediction accuracy in multi-peak spectral and extreme sea conditions, making it difficult to ensure the stability and reliability of the prediction results. SUMMARY
[0004] To solve the above problems, the present application provides a sea wave lightweight prediction method based on knowledge distillation. To address the problem of existing techniques relying solely on small-scale neural networks for lightweight implementation but lacking accuracy and stability in extreme sea conditions, a "large sea area + small sea area" dual-input teacher model and a student model relying only on small sea area input data are constructed, and the knowledge distillation mechanism is used to transfer the expression ability of the teacher model in large-scale features, swell propagation and energy cross-zone transmission to the student model. The student model can maintain prediction accuracy and stability close to the teacher model while significantly reducing the number of parameters and inference time, thus achieving a balance between accuracy and efficiency and meeting the real-time application needs of end-side devices such as ships, buoys and near-shore disaster prevention and warning.
[0005] The sea wave lightweight prediction method based on knowledge distillation comprises:
[0006] Collecting historical wave height data of large and small sea areas, performing spatio-temporal alignment and normalization, and constructing training samples for the teacher model and the student model;
[0007] The teacher model and the student model are constructed, the teacher model receives the large sea area input data and the small sea area input data at the same time, and outputs the wave height prediction result of the target small sea area; the student model only receives the small sea area input data, and outputs the wave height prediction result of the small sea area;
[0008] A knowledge distillation mechanism is constructed, the wave height prediction result and the intermediate feature of the teacher model are used as soft labels, and the student model is trained through a comprehensive distillation loss function including a supervision loss function, a distillation loss function and a feature alignment loss function, so that the student model inherits the prediction ability of the teacher model under the lightweight structure.
[0009] Further, the teacher model receives the large sea area input data and the small sea area input data , and obtains the intermediate feature of the teacher model through feature encoder fusion :
[0010] ;
[0011] Among them, and respectively represent the feature encoders of the large sea area data and the small sea area data, represent feature splicing;
[0012] Then, the future wave height sequence of the small sea area predicted by the teacher model is obtained through the decoder :
[0013] .
[0014] Further, the student model only receives the small sea area input data , extracts features through the lightweight encoder , and obtains the intermediate feature of the student model :
[0015] ;
[0016] Then, the future wave height sequence of the small sea area predicted by the student model is generated through the decoder :
[0017] .
[0018] Further, in the training stage, the student model learns the knowledge of the teacher model through the comprehensive distillation loss function , and minimizes the error with the real label Y:
[0019] ;
[0020] wherein: is a supervised loss function for student prediction vs. ground truth; is a distillation loss function between student prediction and teacher prediction; is an intermediate feature alignment loss function; is a weighting coefficient.
[0021] Further, in the soft target distillation process, temperature scaling Softmax is adopted for teacher model predicted wave height and student model predicted wave height:
[0022] ;
[0023] wherein T is a distillation temperature parameter, and are normalized representations of teacher model predicted wave height and student model predicted wave height under temperature scaling, and are the i-th and j-th predicted wave height of the teacher model; and are the i-th and j-th predicted wave height of the student model;
[0024] Distillation loss function is:
[0025] .
[0026] Further, the following indicators are adopted to compare the prediction performance of the teacher model and the student model on the same test set:
[0027] Mean absolute error MAE:
[0028] ;
[0029] Root mean square error RMSE:
[0030] ;
[0031] Compression ratio CR:
[0032] ;
[0033] wherein N is the total number of test samples, is the true value of the i-th sample, is the predicted value of the i-th sample, denotes the parameter amount of the student model, denotes the parameter amount of the teacher model.
[0034] Compared with the prior art, the present application has the following beneficial effects:
[0035] The application significantly improves the efficiency of sea wave prediction and the ability to maintain accuracy under limited computing power conditions by constructing a teacher model with double input of "large sea area + small sea area" and a lightweight student model relying only on small sea area input data. Through the knowledge distillation mechanism, the student model can inherit the large-scale physical characteristics learned by the teacher model without directly accessing large sea area data, thereby ensuring stable prediction performance in cross-scale processes such as swell propagation and energy input. Compared with existing methods that rely only on small models, the application significantly reduces model computation while ensuring prediction accuracy, enabling real-time operation on shipboard, buoy and other end-side devices, meeting the needs of maritime navigation safety, operation window judgment and near-shore disaster warning and other application scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 FIG. 1 is a schematic diagram of the student-teacher model training process;
[0037] Figure 2 FIG. 4 is a comparison of prediction values and measured values of different models. DETAILED DESCRIPTION
[0038] The application will be described in detail below with specific embodiments. The following embodiments will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be noted that for those skilled in the art, without departing from the concept of the application, a number of changes and improvements can be made. These are within the scope of the present application.
[0039] The application provides a sea wave lightweight height prediction method based on knowledge distillation, and the method steps are as follows:
[0040] (1) Data preprocessing and sample construction.
[0041] Collect historical wave height data of large and small sea areas, complete spatiotemporal alignment and normalization processing, and form training samples for teacher and student models. Specifically:
[0042] 1. Data acquisition
[0043] Obtain large sea area wave height field and small sea area wave height field , with a time resolution of .
[0044] 2. Spatiotemporal alignment
[0045] Map the large sea area data to the small sea area grid by interpolation to obtain the mapped large sea area wave height field :
[0046] ;
[0047] where denotes the interpolation operator, is the small sea grid.
[0048] 3. Data standardization
[0049] Z-score normalization is applied to the input wave height.
[0050] 4. Sample construction
[0051] The sample is constructed in a sliding window manner, with the history length , the prediction step number , and the sample is:
[0052] ;
[0053] where is the small sea input data, is the large sea input data, is the target output, is the small sea wave height field from time point t-L+1, is the small sea wave height field from time point t-L+1 to time point t; is the small sea wave height field from time point t+1 to time point ; is the large sea wave height field from time point t-L+1 to time point t after mapping.
[0054] (2) Teacher-student model construction
[0055] The teacher model that receives both large sea and small sea input data, and the lightweight student model that only receives small sea input data are established, and the target small sea future wave height prediction results are output respectively. Specifically:
[0056] 1. Teacher model structure design.
[0057] The teacher model receives large sea input data and small sea input data , and performs joint modeling through a multi-layer encoder-decoder network.
[0058] The teacher model receives large sea input data and small sea input data , and obtains the intermediate feature of the teacher model through a feature encoder fusion:
[0059] ;
[0060] where and respectively represent the feature encoders of large and small sea area data, representing feature splicing, is the intermediate feature of the teacher model.
[0061] through the decoder get the small sea area future wave height sequence predicted by the student model step wave height sequence :
[0062] .
[0063] 2. Student model structure design.
[0064] The student model adopts a lightweight design in structure: the lightweight encoder of the student model Compared with the multi-layer stacked structure of the feature encoder of the teacher model, the number of convolution layers is reduced, and the number of convolution channels is compressed to less than half of that of the feature encoder of the teacher model; at the same time, redundant attention modules are removed, and only key local convolution and a layer of lightweight attention mechanism are reserved to ensure the ability to capture main spatio-temporal features.
[0065] The student model only receives small sea area input data , and extracts features through the lightweight encoder :
[0066] ;
[0067] Among them, is the intermediate feature of the student model.
[0068] through the decoder generate the small sea area future wave height sequence predicted by the student model :
[0069] .
[0070] 3. Teacher-student relationship.
[0071] The core purpose of the training stage is: through the introduction of multiple distillation constraints, the lightweight student model can still inherit the prediction ability formed by the teacher model under the fusion of large and small sea area features, while taking into account the calculation efficiency and prediction accuracy, under the condition of only accepting small sea area input data.
[0072] The intermediate feature of the teacher model is obtained by fusing the large sea area input data and the small sea area input data through the feature encoder, and contains comprehensive information across scales and regions, which cannot be directly obtained by the student model relying on small sea area input data alone.
[0073] In the training stage, the student model not only minimizes the error with the true label, but also learns the knowledge of the teacher model through the distillation loss function:
[0074] ;
[0075] wherein: is the comprehensive distillation loss function, is the supervision loss function of the student prediction and the true value; is the distillation loss function between the student prediction and the teacher prediction; is the intermediate feature alignment loss function for constraining the student hidden space to be consistent with the teacher; is the weighting coefficient.
[0076] (3) Knowledge distillation training.
[0077] Using the predicted wave height and the intermediate feature of the teacher model as a soft label, the student model is trained through the distillation loss function, so that the student model can still maintain high prediction accuracy under the lightweight structure. Specifically:
[0078] 1. Distillation training strategy.
[0079] In the training process, the student model updates the parameters by minimizing the comprehensive distillation loss function. The Adam optimizer is adopted, the initial learning rate , the weight decay coefficient , and the training round is .
[0080] In the soft target distillation process, the teacher model prediction wave height and the student model prediction wave height are temperature scaled using Softmax:
[0081] ;
[0082] wherein, is the distillation temperature parameter for smoothing the teacher prediction distribution, and are the i-th and j-th prediction wave heights of the teacher model; and are the i-th and j-th prediction wave heights of the student model.
[0083] and are the normalized representations of the teacher prediction wave height and the student prediction wave height under temperature scaling.
[0084] The corresponding distillation loss function is defined as:
[0085] .
[0086] 2. Model evaluation method.
[0087] The following indicators are used to compare the prediction performance of the teacher model and the student model on the same test set:
[0088] Mean Absolute Error (MAE):
[0089] ;
[0090] Root Mean Square Error (RMSE):
[0091] ;
[0092] Compression Ratio (CR):
[0093] ;
[0094] where N is the total number of test samples, is the true value of the i-th sample, is the predicted value of the i-th sample, represents the parameter amount of the student model, represents the parameter amount of the teacher model.
[0095] (4) Model evaluation and verification.
[0096] By comparing the prediction results of the teacher model and the student model, combining the accuracy and compression indicators, the effectiveness of the proposed method is comprehensively evaluated. Table 1 shows the comparison results of the simulation test, Figure 2 is the comparison of the predicted values and the measured values of different models.
[0097] Table 1 Comparison results of simulation test
[0098]
[0099] In the simulation experiment, the teacher model, the student model (without knowledge distillation), and the student model (with knowledge distillation) are constructed to predict sea waves, and the differences in accuracy and efficiency are compared. From the comparison results, it can be seen that the teacher model performs best in reasoning accuracy with an error of 0.20 m, but its parameter amount reaches 1.5 M, and the reasoning time (0.20 s) is significantly higher than that of the student model. Although the student model has a significant lightweight advantage with a parameter amount of only 0.4 M and a reasoning time of only 0.05 s, it has a larger error (0.33 m) without distillation. By introducing knowledge distillation, the student model reduces the error to 0.23 m under the premise of maintaining the same parameter amount and reasoning efficiency, close to the level of the teacher model, achieving a balance between accuracy and efficiency.
[0100] The specific embodiments of the present application are described above. It needs to be understood that the present application is not limited to the specific embodiments described above, and various changes or modifications can be made by those skilled in the art within the scope of the claims, which do not affect the essential content of the present application. The embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily without conflict.
Claims
1. A method for sea wave lightweight forecasting based on knowledge distillation, characterized in that, The application comprises the following steps: Collecting historical wave height data of large and small sea areas, performing spatio-temporal alignment and normalization processing, and constructing training samples for teacher and student models; Constructing teacher and student models, the teacher model simultaneously receives large and small sea area input data and outputs wave height prediction results of the target small sea area; the student model only receives small sea area input data and outputs wave height prediction results of the small sea area; The teacher model receives the large-area input data X L and the small-area input data X S , and obtains the intermediate feature H of the teacher model through the feature encoder T : wherein, Enc L (·) and Enc S (·) represent the feature encoders for large and small sea area data, respectively, denotes feature concatenation; The small sea area future τ step wave height sequence Y predicted by the teacher model is obtained through the decoder Dec(·) T : Y T = Dec(H T ) ; The student model only receives small sea area input data X S , and feature extraction is performed through a lightweight encoder Enc' S (·) to obtain intermediate features H S ' of the student model. H S ′ = Enc′ S (X S ); The student model predicted future sea surface height sequence Y is generated again by the decoder Dec'(·) S : Y S = Dec'(H S ') ; In the training phase, the student model learns the knowledge of the teacher model by minimizing the error with respect to the real labels Y through a comprehensive distillation loss function minimizing the error with respect to the real labels Y wherein: is a supervised loss function for student predictions vs. true values; is a distillation loss function between student predictions and teacher predictions; is an intermediate feature alignment loss function; and a, b, g are weighting coefficients. In the soft target distillation process, the temperature scaling Softmax is used for the wave height predicted by the teacher model and the wave height predicted by the student model: where T is a distillation temperature parameter, and is a normalized representation of the teacher model predicted wave height and the student model predicted wave height under temperature scaling, and is the i-th and j-th predicted wave height of the teacher model; and is the i-th and j-th predicted wave height of the student model; Distillation loss function is: A knowledge distillation mechanism is constructed, the wave height prediction results and intermediate features of the teacher model are used as soft labels, and the student model is trained through a comprehensive distillation loss function containing a supervision loss function, a distillation loss function and a feature alignment loss function, so that the student model inherits the prediction ability of the teacher model under a lightweight structure. 2.The knowledge distillation based sea wave lightweight prediction method according to claim 1, wherein, The following indicators are used to compare the prediction performance of the teacher model and the student model on the same test set: Mean absolute error MAE: Root mean square error RMSE: Lightweight ratio CR: where N is the total number of test samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample, and Params(Student) represents the number of parameters of the student model, and Params(Teacher) represents the number of parameters of the teacher model.
Citation Information
Patent Citations
Track target point prediction method based on knowledge distillation
CN116579423A
Lightweight image classification neural network architecture system based on knowledge distillation
CN119514594A