Polar region sea ice prediction model training method and device
By introducing a physical constraint loss function into the polar sea ice prediction model and combining it with deep learning, the problems of insufficient prediction accuracy and physical inconsistency in existing methods are solved, and higher accuracy and stability of polar sea ice prediction are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTH CHINA UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-08
AI Technical Summary
Existing polar sea ice prediction methods have significant limitations in terms of accuracy, generalization ability, and interpretability. In particular, physics-driven and data-driven methods are insufficient in prediction accuracy when dealing with complex and nonlinear environmental interactions, and deep learning methods lack physical consistency and model generalization ability.
By introducing multiple physical constraint loss functions into the polar sea ice prediction model, we ensure that the prediction results follow the laws of sea ice drift and buoyancy balance. Combined with a deep learning model, we improve the physical consistency and generalization ability of the model.
It improves the accuracy and stability of polar sea ice prediction, enhances the interpretability of the model and its sensitivity to spatial and temporal variations, reduces the risk of overfitting, and provides a more reliable data foundation.
Smart Images

Figure CN121997046A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a training method and apparatus for a polar sea ice prediction model. Background Technology
[0002] Against the backdrop of increasingly frequent global climate change and polar exploration activities, accurate prediction of polar sea ice conditions is crucial for navigation planning, resource exploration, and safety assurance. Predictions of sea ice concentration (SIC) and sea ice thickness (SIT) are core indicators for assessing ice conditions and supporting decision-making; sea ice concentration is also known as sea ice density.
[0003] Currently, polar sea ice prediction methods mainly rely on two types of technologies: physics-driven methods and data-driven methods. Physics-driven methods rely on models of ocean and atmospheric physical processes to simulate sea ice changes by solving physical equations. Data-driven methods predict polar sea ice by analyzing historical data to identify trends and periodic changes. Data-driven methods include statistical models, traditional machine learning, and deep learning.
[0004] However, physics-driven methods are limited by the accuracy of data acquisition and the adaptability of models, and cannot effectively handle complex and nonlinear environmental interactions, resulting in low prediction accuracy. Statistical models and traditional machine learning methods in data-driven approaches struggle to capture the nonlinear relationships of sea ice changes and tend to overlook important features, potentially limiting model performance. Deep learning methods in data-driven approaches, while capable of automatically extracting features and processing high-dimensional data, are often considered "black boxes" due to insufficient model generalization ability and poor interpretability. In particular, they fail to fully utilize the physical correlations between variables in multivariate time series data processing, leading to insufficient prediction accuracy. Summary of the Invention
[0005] This application provides a training method and apparatus for a polar sea ice prediction model to improve the prediction accuracy of polar sea ice prediction, as well as enhance the model's generalization ability, interpretability, and physical consistency.
[0006] In a first aspect, embodiments of this application provide a training method for a polar sea ice prediction model, the method comprising: Obtain a training sample set. Each training sample in the training sample set includes: historical sea ice state data for the historical period of the sample and actual sea ice state data for the predicted period of the sample. The polar sea ice prediction model to be trained is iteratively trained based on the training sample set to obtain the target polar sea ice prediction model. During one iteration of training, the following operations are performed: Feature extraction is performed on the historical sea ice state data of the selected target training samples to obtain initial sample sea ice features. Based on the initial sample sea ice features, the sample sea ice prediction results of the target training samples are determined. The sample sea ice prediction results include: predicted sea ice concentration and predicted sea ice thickness. Multiple loss functions are used to calculate the loss value of the sample sea ice prediction results, determine the corresponding loss value of the sample sea ice prediction results, and adjust the model parameters of the polar sea ice prediction model to be trained based on the loss value. Among them, the multiple loss functions include: a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function is used to measure whether the change of predicted sea ice concentration conforms to the law of sea ice drift in the actual environment, and the second physical constraint loss function is used to measure whether the predicted sea ice thickness satisfies the physical law of buoyancy balance.
[0007] In one optional embodiment, the loss value of the sample sea ice prediction result is calculated using multiple loss functions to determine the corresponding loss value of the sample sea ice prediction result, including: Based on multiple loss functions, calculate multiple sub-loss values corresponding to the sea ice prediction results of the samples; The loss value is obtained by summing the multiple sub-loss values.
[0008] In an optional embodiment, the predicted sea ice concentration includes: the predicted sea ice concentration value of each predicted sample point; the predicted sea ice thickness includes: the predicted sea ice thickness value of each predicted sample point; and the sample sea ice prediction result further includes: the predicted surface snow thickness, which includes: the predicted sea ice thickness value of each predicted sample point. Based on multiple loss functions, several sub-loss values corresponding to the sea ice prediction results of the samples are calculated, including: Based on the first rate of change of the predicted sea ice concentration value with time, the second rate of change with velocity, and convergence and divergence of each predicted sample point, the predicted concentration constraint value of each predicted sample point is determined. Based on the first physical constraint loss function, the error between the predicted concentration constraint value of each predicted sample point and the corresponding real concentration constraint value is calculated to obtain the first sub-loss value. The real concentration constraint value is determined based on the real sea ice state data of the target training sample. Based on the predicted sea ice thickness and predicted surface snow thickness values of each predicted sample point, and combined with seawater density, sea ice density, and snow density, the predicted free plate height of each predicted sample point is determined. Based on the second physical constraint loss function, the error between the predicted free plate height of each predicted sample point and the corresponding real free plate height is calculated to obtain the second sub-loss value. The real free plate height is determined based on the real sea ice state data of the target training sample.
[0009] In one optional embodiment, the multiple loss functions further include: a mean squared error loss function; the method further includes: Based on the mean squared error loss function, the error between the sea ice prediction results of the sample and the real sea ice state data of the target training sample is calculated to obtain the third sub-loss value.
[0010] In an optional embodiment, the sea ice prediction results for the sample further include: predicted eastward velocity and predicted northward velocity, wherein the predicted eastward velocity includes the predicted eastward velocity value for each predicted sample point, and the predicted northward velocity includes the predicted northward velocity value for each predicted sample point. Before determining the predicted concentration constraint value for each predicted sea ice concentration point based on the first rate of change of the predicted sea ice concentration over time, the second rate of change of the predicted sea ice concentration over velocity, and convergence and divergence, the following steps are also included: For each predicted sample point, perform the following operations: Based on the predicted sea ice concentration value of a predicted sample point, the predicted sea ice concentration value of the previous moment of a predicted sample point, and the time step, determine the first rate of change of the predicted sea ice concentration value of a predicted sample point over time. Based on the predicted sea ice concentration values of each of the neighboring predicted sample points in the east, west, north, and south directions of a predicted sample point, the longitude partial derivative of the predicted sea ice concentration value of a predicted sample point in the longitude direction and the latitude partial derivative of the predicted sea ice concentration value in the latitude direction are determined. Based on the predicted eastward velocity value and the predicted northward velocity value of a predicted sample point, as well as the longitude partial derivative and the latitude partial derivative of the concentration, the second rate of change of the predicted sea ice concentration value of a predicted sample point with velocity is determined. Based on the predicted eastward velocity values of each adjacent predicted sample point to the east and west, and the predicted northward velocity values of each adjacent predicted sample point to the north and south, the longitudinal partial derivative of the predicted eastward velocity value of a predicted sample point in the longitude direction and the latitudinal partial derivative of the predicted northward velocity value in the latitude direction are determined. Based on the longitudinal and latitudinal partial derivatives of the velocity, the velocity field divergence of a predicted sample point is determined. Based on the predicted sea ice concentration value and the velocity field divergence of a predicted sample point, the convergence and divergence of a predicted sample point are determined.
[0011] In one optional embodiment, based on the predicted eastward and predicted northward velocity values of a predicted sample point, as well as the longitude and latitude partial derivatives of the concentration, a second rate of change of the predicted sea ice concentration value of a predicted sample point with respect to velocity is determined, including: Based on the predicted eastward velocity value and the longitude partial derivative of the concentration at a predicted sample point, the first rate of change component of the predicted sea ice concentration value at the predicted sample point in the longitude direction with velocity is determined. Based on the predicted northward velocity value and the latitudinal partial derivative of the concentration at a predicted sample point, the second rate of change component of the predicted sea ice concentration value at a predicted sample point in the latitudinal direction with velocity is determined. Based on the first and second rate of change components, a second rate of change is determined for the predicted sea ice concentration value of a predicted sample point as a function of velocity.
[0012] In one optional embodiment, the historical sea ice state data of the target training sample includes: historical sea ice concentration value, historical sea ice thickness value, historical eastward velocity value, historical northward velocity value, and historical surface snow thickness value for each historical sample point; Then, feature extraction is performed on the historical sea ice state data of the selected target training samples to obtain the initial sample sea ice features, including: The data required for physical constraints are extracted from the historical sea ice state data of the target training samples. The data required for physical constraints include: the historical sea ice concentration values of each historical sample point corresponding to its east, west, north, and south adjacent historical sample points; the historical sea ice concentration value of each historical sample point corresponding to the previous moment; the historical eastward velocity values of each historical sample point corresponding to its east, west, and south adjacent historical sample points; and the historical northward velocity values of each historical sample point corresponding to its north, south, and south adjacent historical sample points. The data required for physical constraints are concatenated with the historical sea ice state data of the target training samples to form an enhanced input vector; The initial sample sea ice features are generated by fusing spatial and temporal information into the enhanced input vector through an embedding layer.
[0013] In one optional embodiment, determining the sea ice prediction result of the target training sample based on the initial sample sea ice features includes: The initial sample sea ice features are input into a sequence of predictor modules containing n predictor modules, and the output results corresponding to each of the n predictor modules are generated step by step. The output result corresponding to the nth predictor module is used as the sample sea ice prediction result. Where n is an integer greater than 1, each predictor module outputs a prediction result component and residual features based on its input data. The residual features serve as the input data for the next predictor module, and the initial sample sea ice features serve as the input data for the first predictor module. The output result of the nth predictor module is obtained by subtracting the output result of the (n-1)th predictor module from the prediction result component of the nth predictor module. The output result of the first predictor module is the prediction result component of the first predictor module.
[0014] In one alternative embodiment, each predictor module outputs prediction result components and residual features based on its input data, including: An attention mechanism is applied to the input data corresponding to a predictor module to generate attention features; The attention features are randomly discarded to obtain the discarded features; Subtract the discarded features from the input data corresponding to a predictor module to obtain the first intermediate feature; The first intermediate feature is processed by the first branch to obtain a residual feature output by a predictor module, and the first intermediate feature is processed by the second branch to obtain a prediction result component output by a predictor module.
[0015] Secondly, embodiments of this application also provide a training device for a polar sea ice prediction model, the device comprising: The acquisition module is used to acquire the training sample set. Each training sample in the training sample set includes: historical sea ice state data for the historical period of the sample and actual sea ice state data for the predicted period of the sample. The training module is used to iteratively train the polar sea ice prediction model to be trained based on the training sample set to obtain the target polar sea ice prediction model. During one iteration of training, the following operations are performed: Feature extraction is performed on the historical sea ice state data of the selected target training samples to obtain initial sample sea ice features. Based on the initial sample sea ice features, the sample sea ice prediction results of the target training samples are determined. The sample sea ice prediction results include: predicted sea ice concentration and predicted sea ice thickness. Multiple loss functions are used to calculate the loss value of the sample sea ice prediction results, determine the corresponding loss value of the sample sea ice prediction results, and adjust the model parameters of the polar sea ice prediction model to be trained based on the loss value. Among them, the multiple loss functions include: a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function is used to measure whether the change of predicted sea ice concentration conforms to the law of sea ice drift in the actual environment, and the second physical constraint loss function is used to measure whether the predicted sea ice thickness satisfies the physical law of buoyancy balance.
[0016] In an optional embodiment, when calculating the loss value of the sample sea ice prediction result using multiple loss functions and determining the loss value corresponding to the sample sea ice prediction result, the training module is further used to: Based on multiple loss functions, calculate multiple sub-loss values corresponding to the sea ice prediction results of the samples; The loss value is obtained by summing the multiple sub-loss values.
[0017] In an optional embodiment, the predicted sea ice concentration includes: the predicted sea ice concentration value of each predicted sample point; the predicted sea ice thickness includes: the predicted sea ice thickness value of each predicted sample point; and the sample sea ice prediction result further includes: the predicted surface snow thickness, which includes: the predicted sea ice thickness value of each predicted sample point. When calculating multiple sub-loss values corresponding to the sea ice prediction results of the samples based on multiple loss functions, the training module is also used for: Based on the first rate of change of the predicted sea ice concentration value with time, the second rate of change with velocity, and convergence and divergence of each predicted sample point, the predicted concentration constraint value of each predicted sample point is determined. Based on the first physical constraint loss function, the error between the predicted concentration constraint value of each predicted sample point and the corresponding real concentration constraint value is calculated to obtain the first sub-loss value. The real concentration constraint value is determined based on the real sea ice state data of the target training sample. Based on the predicted sea ice thickness and predicted surface snow thickness values of each predicted sample point, and combined with seawater density, sea ice density, and snow density, the predicted free plate height of each predicted sample point is determined. Based on the second physical constraint loss function, the error between the predicted free plate height of each predicted sample point and the corresponding real free plate height is calculated to obtain the second sub-loss value. The real free plate height is determined based on the real sea ice state data of the target training sample.
[0018] In one optional embodiment, the multiple loss functions further include: a mean squared error loss function; the training module is also used for: Based on the mean squared error loss function, the error between the sea ice prediction results of the sample and the real sea ice state data of the target training sample is calculated to obtain the third sub-loss value.
[0019] In an optional embodiment, the sea ice prediction results for the sample further include: predicted eastward velocity and predicted northward velocity, wherein the predicted eastward velocity includes the predicted eastward velocity value for each predicted sample point, and the predicted northward velocity includes the predicted northward velocity value for each predicted sample point. Before determining the predicted concentration constraint value for each predicted sea ice concentration point based on the first rate of change of the predicted sea ice concentration over time, the second rate of change of the predicted sea ice concentration over velocity, and convergence and divergence, the training module is also used for: For each predicted sample point, perform the following operations: Based on the predicted sea ice concentration value of a predicted sample point, the predicted sea ice concentration value of the previous moment of a predicted sample point, and the time step, determine the first rate of change of the predicted sea ice concentration value of a predicted sample point over time. Based on the predicted sea ice concentration values of each of the neighboring predicted sample points in the east, west, north, and south directions of a predicted sample point, the longitude partial derivative of the predicted sea ice concentration value of a predicted sample point in the longitude direction and the latitude partial derivative of the predicted sea ice concentration value in the latitude direction are determined. Based on the predicted eastward velocity value and the predicted northward velocity value of a predicted sample point, as well as the longitude partial derivative and the latitude partial derivative of the concentration, the second rate of change of the predicted sea ice concentration value of a predicted sample point with velocity is determined. Based on the predicted eastward velocity values of each adjacent predicted sample point to the east and west, and the predicted northward velocity values of each adjacent predicted sample point to the north and south, the longitudinal partial derivative of the predicted eastward velocity value of a predicted sample point in the longitude direction and the latitudinal partial derivative of the predicted northward velocity value in the latitude direction are determined. Based on the longitudinal and latitudinal partial derivatives of the velocity, the velocity field divergence of a predicted sample point is determined. Based on the predicted sea ice concentration value and the velocity field divergence of a predicted sample point, the convergence and divergence of a predicted sample point are determined.
[0020] In an optional embodiment, when determining a second rate of change of the predicted sea ice concentration value as a function of velocity for a predicted sample point based on the predicted eastward velocity value and the predicted northward velocity value, as well as the concentration longitude partial derivative and the concentration latitude partial derivative, the training module is further configured to: Based on the predicted eastward velocity value and the longitude partial derivative of the concentration at a predicted sample point, the first rate of change component of the predicted sea ice concentration value at the predicted sample point in the longitude direction with velocity is determined. Based on the predicted northward velocity value and the latitudinal partial derivative of the concentration at a predicted sample point, the second rate of change component of the predicted sea ice concentration value at a predicted sample point in the latitudinal direction with velocity is determined. Based on the first and second rate of change components, a second rate of change is determined for the predicted sea ice concentration value of a predicted sample point as a function of velocity.
[0021] In one optional embodiment, the historical sea ice state data of the target training sample includes: historical sea ice concentration value, historical sea ice thickness value, historical eastward velocity value, historical northward velocity value, and historical surface snow thickness value for each historical sample point; When extracting features from the historical sea ice state data of the selected target training samples to obtain the initial sample sea ice features, the training module is also used for: The data required for physical constraints are extracted from the historical sea ice state data of the target training samples. The data required for physical constraints include: the historical sea ice concentration values of each historical sample point corresponding to its east, west, north, and south adjacent historical sample points; the historical sea ice concentration value of each historical sample point corresponding to the previous moment; the historical eastward velocity values of each historical sample point corresponding to its east, west, and south adjacent historical sample points; and the historical northward velocity values of each historical sample point corresponding to its north, south, and south adjacent historical sample points. The data required for physical constraints are concatenated with the historical sea ice state data of the target training samples to form an enhanced input vector; The initial sample sea ice features are generated by fusing spatial and temporal information into the enhanced input vector through an embedding layer.
[0022] In an optional embodiment, when determining the sea ice prediction result of the target training sample based on the initial sample sea ice features, the training module is further configured to: The initial sample sea ice features are input into a sequence of predictor modules containing n predictor modules, and the output results corresponding to each of the n predictor modules are generated step by step. The output result corresponding to the nth predictor module is used as the sample sea ice prediction result. Where n is an integer greater than 1, each predictor module outputs a prediction result component and residual features based on its input data. The residual features serve as the input data for the next predictor module, and the initial sample sea ice features serve as the input data for the first predictor module. The output result of the nth predictor module is obtained by subtracting the output result of the (n-1)th predictor module from the prediction result component of the nth predictor module. The output result of the first predictor module is the prediction result component of the first predictor module.
[0023] In an optional embodiment, the training module is further configured to: An attention mechanism is applied to the input data corresponding to a predictor module to generate attention features; The attention features are randomly discarded to obtain the discarded features; Subtract the discarded features from the input data corresponding to a predictor module to obtain the first intermediate feature; The first intermediate feature is processed by the first branch to obtain a residual feature output by a predictor module, and the first intermediate feature is processed by the second branch to obtain a prediction result component output by a predictor module.
[0024] Thirdly, embodiments of this application also provide an electronic device, including: Processor; and Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the training method for the polar sea ice prediction model as described in the first aspect.
[0025] Fourthly, embodiments of this application also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the training method for the polar sea ice prediction model as described in the first aspect.
[0026] Fifthly, this application provides a computer program product that, when invoked by a computer, causes the computer to execute the training method steps of the polar sea ice prediction model as described in the first aspect.
[0027] The beneficial effects of this application are as follows: In the training method of the polar sea ice prediction model provided in the embodiments of this application, a training sample set is obtained. Each training sample in the training sample set includes: historical sea ice state data for the historical period of the sample and actual sea ice state data for the prediction period of the sample. The polar sea ice prediction model to be trained is iteratively trained based on the training sample set to obtain the target polar sea ice prediction model. During each iteration, the following operations are performed: feature extraction is performed on the historical sea ice state data of the selected target training samples to obtain initial sample sea ice features. Based on these initial features, the sample sea ice prediction results for the target training samples are determined. These results include predicted sea ice concentration and predicted sea ice thickness. Multiple loss functions are used to calculate the loss values for the predicted sea ice concentrations, and the corresponding loss values are determined. Based on these loss values, the model parameters of the polar sea ice prediction model to be trained are adjusted. The multiple loss functions include a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function measures whether the change in predicted sea ice concentration conforms to the laws of sea ice drift in the actual environment, and the second physical constraint loss function measures whether the predicted sea ice thickness satisfies the physical law of buoyancy balance. By integrating physical laws into the deep learning model, the model's predictions are ensured to adhere to these laws (not only ensuring that predicted sea ice concentration changes conform to the laws of sea ice drift in the actual environment, but also ensuring that predicted sea ice thickness satisfies the physical law of buoyancy balance). This makes the model output physically plausible, overcomes the physical inconsistencies caused by traditional methods, and enhances the model's sensitivity to spatial and temporal variations. Simultaneously, multiple physical constraint loss functions serve as regularization terms, reducing the risk of overfitting the model to training data and resulting in more stable performance in the complex polar environment. In summary, this improves the prediction accuracy of polar sea ice forecasting and enhances the model's generalization ability, interpretability, and physical consistency.
[0028] Furthermore, other features and advantages of this application will be set forth in the following description and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described herein are used to provide a further understanding of this application, constitute a part of this application, and do not constitute an improper limitation of this application. In the accompanying drawings: Figure 1 This is a schematic diagram of an optional system architecture applicable to the embodiments of this application.
[0030] Figure 2 This is a schematic diagram illustrating the implementation process of a training method for a polar sea ice prediction model provided in an embodiment of this application.
[0031] Figure 3 This is a schematic diagram illustrating the formation of an enhanced input vector, provided as an embodiment of this application.
[0032] Figure 4 This is a schematic diagram of the structure of a dual-stream and subtraction mechanism module provided in an embodiment of this application.
[0033] Figure 5 This is a schematic diagram of the structure of a predictor module provided in an embodiment of this application.
[0034] Figure 6 This is a schematic diagram of the structure of a training device for a polar sea ice prediction model provided in an embodiment of this application.
[0035] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0036] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0037] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0038] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0039] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0040] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0041] The design concept of the embodiments of this application is briefly introduced below: Against the backdrop of increasingly frequent global climate change and polar exploration activities, accurate prediction of polar sea ice conditions is crucial for navigation planning, resource exploration, and safety assurance. Predictions of sea ice concentration and thickness are core indicators for assessing ice conditions and supporting decision-making.
[0042] Currently, polar sea ice prediction methods mainly rely on two types of technologies: physics-driven methods and data-driven methods. Physics-driven methods rely on models of ocean and atmospheric physical processes to simulate sea ice changes by solving physical equations. Data-driven methods predict polar sea ice by analyzing historical data to identify trends and periodic changes. Data-driven methods include statistical models, traditional machine learning, and deep learning.
[0043] However, physics-driven methods are limited by the accuracy of data acquisition and the adaptability of models, failing to effectively handle complex and nonlinear environmental interactions, resulting in low prediction accuracy. Statistical models and traditional machine learning methods within data-driven approaches struggle to capture the nonlinear relationships in sea ice changes, easily overlooking important features and potentially limiting model performance. Deep learning methods within data-driven approaches, while capable of automatically extracting features and processing high-dimensional data, are often considered "black boxes," exhibiting insufficient generalization ability and poor interpretability. Particularly in multivariate time-series data processing, they fail to fully utilize the physical correlations between variables, leading to insufficient prediction accuracy. Furthermore, these methods are susceptible to overfitting, especially when co-predicting multiple sea ice variables (such as SIC and SIT). Large sample size variations and high feature extraction difficulty can exacerbate overfitting risks by blindly increasing model depth, while ignoring relevant variables (such as sea ice drift speed and temperature) reduces the model's ability to understand environmental factors.
[0044] Overall, existing methods have significant limitations in terms of accuracy, generalization ability, and interpretability, which restricts the application potential of short-term polar sea ice prediction in scenarios such as route planning and resource exploration.
[0045] In view of this, this application proposes a training method for a polar sea ice prediction model, which may specifically include: acquiring a training sample set, wherein each training sample in the training sample set includes: historical sea ice state data for a historical period and actual sea ice state data for a prediction period; iteratively training the polar sea ice prediction model to be trained based on the training sample set to obtain a target polar sea ice prediction model, wherein, in one iteration of training, the following operations are performed: feature extraction is performed on the historical sea ice state data of the selected target training samples to obtain initial sample sea ice features, and based on the initial sample sea ice features, ... The sea ice prediction results for the target training samples are determined, including predicted sea ice concentration and predicted sea ice thickness. Multiple loss functions are used to calculate the loss values for the sea ice prediction results, and the model parameters of the polar sea ice prediction model to be trained are adjusted based on these loss values. The multiple loss functions include a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function measures whether the change in predicted sea ice concentration conforms to the law of sea ice drift in the actual environment, and the second physical constraint loss function measures whether the predicted sea ice thickness satisfies the physical law of buoyancy balance.
[0046] By employing the above method, multiple physical constraint loss functions are constructed based on the physical relationship between input variables and prediction targets (predicted sea ice concentration and predicted sea ice thickness). Physical laws are integrated into the deep learning model, ensuring that the model's output predictions follow physical laws (ensuring not only that predicted sea ice concentration changes conform to the laws of sea ice drift in the actual environment, but also that predicted sea ice thickness satisfies the physical law of buoyancy balance). This makes the model output physically reasonable, overcoming the physical inconsistency problems caused by traditional methods and enhancing the model's sensitivity to spatial and temporal changes. Simultaneously, multiple physical constraint loss functions serve as regularization terms, reducing the risk of overfitting the model to training data and resulting in more stable performance in the complex polar environment. In summary, this application provides a training method for a polar sea ice prediction model, improving the prediction accuracy of polar sea ice prediction and enhancing the model's generalization ability, interpretability, and physical consistency. It not only overcomes the limitations of traditional methods but also provides a more reliable data foundation for polar shipping route planning, possessing significant practical value and potential for widespread application.
[0047] In particular, the preferred embodiments of this application will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments of this application and the features in the embodiments can be combined with each other without conflict.
[0048] See Figure 1 The diagram illustrates an optional system architecture applicable to an embodiment of this application. This system architecture may include a data acquisition device 101 and a server 102. The data acquisition device 101 and the server 102 can interact via a communication network, where the communication network can employ wireless communication and wired communication methods. For example, the data acquisition device 101 can access the network and communicate with the server 102 via cellular mobile communication technology. This cellular mobile communication technology may include, for example, 5G (5th generation mobile networks) or next-generation mobile communication technology. Optionally, the data acquisition devices (101a, 101b) can access the network and communicate with the server 102 via short-range wireless communication. This short-range wireless communication method may include, for example, Wi-Fi (wireless fidelity) technology.
[0049] This application embodiment does not impose any limitation on the number of devices involved in the above system architecture. For example, the above system architecture may include more data acquisition devices, or it may include other devices. Figure 1 As shown, only the data acquisition device 101 and server 102 are described as examples. The following is a brief introduction to each of the above devices and their respective functions.
[0050] Data acquisition device 101 is a device used to acquire sea ice state data. For example, data acquisition device 101 may include, but is not limited to, satellite sensors, ground observation equipment, and lidar.
[0051] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0052] It is worth mentioning that the server 102 in this embodiment can iteratively train the polar sea ice prediction model to be trained based on the obtained training sample set to obtain the target polar sea ice prediction model. In one iteration of training, the following operations are performed: feature extraction is performed on the historical sea ice state data of the selected target training sample to obtain the initial sample sea ice features, and the sample sea ice prediction result of the target training sample is determined based on the initial sample sea ice features; the loss value is calculated on the sample sea ice prediction result through multiple loss functions to determine the loss value corresponding to the sample sea ice prediction result, and the model parameters of the polar sea ice prediction model to be trained are adjusted based on the loss value.
[0053] Optionally, a pre-trained target polar sea ice prediction model can be deployed on server 102. In this way, after receiving a sea ice prediction request for a target prediction period and obtaining sea ice state data for historical periods of the target, server 102 can input the sea ice state data for historical periods of the target into the target polar sea ice prediction model to determine the target sea ice prediction result for the target prediction period.
[0054] For example, the aforementioned target sea ice prediction results may include: target predicted sea ice concentration, target predicted sea ice thickness, target predicted surface snow thickness, target predicted eastward velocity, and target predicted northward velocity.
[0055] The training method of the polar sea ice prediction model provided by the exemplary embodiments of this application will be described below in conjunction with the above system architecture and with reference to the accompanying drawings. It should be noted that the above system architecture is only shown for the purpose of understanding the spirit and principles of this application, and the embodiments of this application are not limited in any way.
[0056] See Figure 2The diagram shown illustrates the implementation flow of a training method for a polar sea ice prediction model provided in this application. Taking a server as an example, the specific implementation flow of this method is as follows: S20: Obtain the training sample set.
[0057] Each training sample in the training sample set includes: historical sea ice state data for the sample's historical period and actual sea ice state data for the sample's prediction period. Historical sea ice state data includes, but is not limited to: historical sea ice concentration, historical sea ice thickness, historical eastward velocity, historical northward velocity, historical surface snow thickness, historical sea ice surface temperature, and historical seabed pressure. Specifically, historical sea ice concentration includes the historical sea ice concentration value for each historical sample point; historical sea ice thickness includes the historical sea ice thickness value for each historical sample point; historical eastward velocity includes the historical eastward velocity value for each historical sample point; historical northward velocity includes the historical northward velocity value for each historical sample point; historical surface snow thickness includes the historical surface snow thickness value for each historical sample point; historical sea ice surface temperature includes the historical sea ice surface temperature value for each historical sample point; and historical seabed pressure includes the historical seabed pressure value for each historical sample point. The actual sea ice state data includes: actual sea ice concentration, actual sea ice thickness, actual eastward velocity, actual northward velocity, and actual surface snow thickness. Among them, actual sea ice concentration includes: actual sea ice concentration values for each predicted sample point; actual sea ice thickness includes: actual sea ice thickness values for each predicted sample point; actual eastward velocity includes: actual eastward velocity values for each predicted sample point; actual northward velocity includes: actual northward velocity values for each predicted sample point; and actual surface snow thickness includes: actual surface snow thickness values for each predicted sample point.
[0058] Furthermore, it is worth noting that the sample points in this application embodiment refer to grid points (m, n, t) at a certain longitude, latitude, and time. m refers to the longitude index, n refers to the latitude index, and t refers to the time index. Each sample point contains the physical quantity value at that grid point. Historical sample points refer to sample points within the historical sample period, and predicted sample points refer to sample points within the predicted sample period.
[0059] In this embodiment of the application, eastward velocity refers to sea ice eastward velocity (SIEV), northward velocity refers to sea ice northward velocity (SINV), surface snow thickness refers to surface snow thickness (SST), sea ice surface temperature refers to sea ice surface temperature (SIST), and seawater pressure at sea floor refers to sea water pressure at sea floor (SWPSF).
[0060] In this embodiment of the application, sea ice state data is collected by a data acquisition device to construct a training sample set that can be used to train a target polar sea ice prediction model.
[0061] Optionally, in this embodiment of the application, the collected sea ice state data is preprocessed, including but not limited to: sorting, singular value processing, and missing value processing.
[0062] Optionally, in this embodiment, before training the polar sea ice prediction model to be trained using the obtained training sample set, the input variables are first screened and reconstructed. Specifically, deep learning algorithms are used to systematically process and analyze the features of multivariate time-series data related to polar sea ice, selecting high-quality variables that contribute significantly to prediction accuracy, and completing variable reconstruction and representation optimization, thereby building efficient data support for subsequent model training. This step overcomes the blindness in variable selection in traditional methods, ensuring that the model focuses on high-contribution factors.
[0063] In this embodiment, the input variables selected through experimental testing of the model include: sea ice concentration, sea ice thickness, eastward velocity, northward velocity, surface snow thickness, sea ice surface temperature, and seabed pressure. These variables provide multi-dimensional environmental information, enabling the model to capture various factors affecting sea ice properties. Among them, there are explicit physical equations between the variables SIC, SIT, SIEV, and SINV, which can map the relationships between variables. Therefore, these four variables form a basic group as the core physical foundation. SST affects the thermal insulation effect of sea ice, and a thick snow layer can reduce the melting rate of the ice layer; SIST directly affects the melting or freezing state of sea ice and, together with the presence of snow, regulates the overall density of the ice; SWPSF can affect the stability and melting rate of bottom sea ice. Although not all variables have explicit physical equations, this invention uses deep learning algorithms to capture the dynamic relationships between variables, thereby achieving more accurate predictions. In addition, through this learning step, the model improves its generalization ability in several aspects and improves its performance on unseen data.
[0064] S21: Iteratively train the polar sea ice prediction model to be trained based on the training sample set to obtain the target polar sea ice prediction model. In one iteration of training, the following steps are included: S210 to S211.
[0065] In this embodiment, the polar sea ice prediction model includes: a data augmentation module, an embedding layer module, and a dual-stream and subtraction mechanism module. Each module is described below: (1) The data augmentation module is used to: extract the data required for physical constraints from the historical sea ice state data of the target training sample, and concatenate the data required for physical constraints with the historical sea ice state data of the target training sample to form an augmented input vector.
[0066] (2) The embedding layer module is used to: fuse and map spatial and temporal information of the enhanced input vector to generate initial sample sea ice features.
[0067] (3) The dual-stream and subtraction mechanism module includes: n predictor modules and a subtractor. The dual-stream and subtraction mechanism module is used to perform feature analysis on the sea ice features of the initial sample and generate the sea ice prediction results of the target training sample.
[0068] In this embodiment of the application, in order to obtain a target polar sea ice prediction model, the polar sea ice prediction model to be trained is iteratively trained based on the training sample set until convergence.
[0069] The convergence conditions include, but are not limited to: the loss value and accuracy remain stable or change very little during training.
[0070] Furthermore, the detailed process of each round of iterative training can be found in S210-S211, as follows: S210: Extract features from the historical sea ice state data of the selected target training samples to obtain initial sample sea ice features, and determine the sample sea ice prediction results of the target training samples based on the initial sample sea ice features.
[0071] The sea ice prediction results for the sample include: predicted sea ice concentration and predicted sea ice thickness. The sea ice prediction results also include: predicted eastward velocity, predicted northward velocity, predicted surface snow thickness, and previously predicted sea ice concentration. Specifically, the predicted sea ice concentration includes the predicted sea ice concentration value for each sample point; the predicted sea ice thickness includes the predicted sea ice thickness value for each sample point; the predicted eastward velocity includes the predicted eastward velocity value for each sample point; the predicted northward velocity includes the predicted northward velocity value for each sample point; the predicted surface snow thickness includes the predicted surface snow thickness value for each sample point; and the previously predicted sea ice concentration includes the predicted sea ice concentration value for each sample point at the previous time step.
[0072] Optionally, in this embodiment of the application, a possible implementation is provided for extracting features from the historical sea ice state data of the selected target training samples to obtain the initial sample sea ice features, specifically by performing the following operations: S2100: Extract the data required for physical constraints from the historical sea ice state data of the target training samples.
[0073] The data required for physical constraints include: the historical sea ice concentration values of each historical sample point corresponding to its east, west, north, and south adjacent historical sample points; the historical sea ice concentration value of each historical sample point corresponding to its previous moment; the historical eastward velocity values of each historical sample point corresponding to its east, west, north, and south adjacent historical sample points; and the historical northward velocity values of each historical sample point corresponding to its north, south, and south adjacent historical sample points.
[0074] Among them, the adjacent historical sample points that are adjacent to a historical sample point in the east, west, north, and south are: in the grid, the eastern adjacent historical sample point to the east of a historical sample point, the western adjacent historical sample point to the west of a historical sample point, the southern adjacent historical sample point to the south of a historical sample point, and the northern adjacent historical sample point to the north of a historical sample point.
[0075] S2101: The data required for physical constraints are concatenated with the historical sea ice state data of the target training samples to form an enhanced input vector.
[0076] For example, see Figure 3 The diagram shown is a schematic of forming an enhanced input vector in an embodiment of this application. The historical sea ice concentration, historical sea ice thickness, historical eastward velocity, historical northward velocity, historical surface snow thickness, historical sea ice surface temperature, and historical seabed pressure in the historical sea ice state data are denoted as SIC, SIT, SIEV, SINV, SST, SIST, and SWPSF, respectively. First, the data of SIC in the four directions of east, south, west, and north and the previous time are extracted from the two-dimensional grid map. The extracted data are denoted as W_SIC, N_SIC, E_SIC, S_SIC, and T-1_SIC. If the current time is the start time, the data of the previous time is filled with the data of the current time. Since sea ice velocity itself is directional, the eastward velocity (SIEV) extracts data in the east and west directions from the two-dimensional grid, and the northward velocity (SINV) extracts data in the south and north directions from the two-dimensional grid. These data are denoted as W_SIEV, E_SIEV, N_SINV and S_SINV, respectively. The extracted data required for physical constraints are concatenated with the original data (SIC, SIT, SIEV, SINV, SST, SIST, SWPSF) to form an enhanced input vector, which serves as the basic input unit for training the network model.
[0077] S2102: The initial sample sea ice features are generated by fusing spatial and temporal information of the enhanced input vector through the embedding layer.
[0078] In this embodiment, the model's embedding layer consists of a fully connected neural network (FCNN), which embeds and maps the spatial and temporal information of the enhanced input vector. The temporal encoding includes year, month, and day information to capture seasonal, periodic, and other temporal patterns. Then, the temporal and spatial data processed by FCNN are added together to obtain the initial sample sea ice features.
[0079] In this embodiment of the application, the formula for calculating the sea ice characteristics of the initial sample can be expressed as:
[0080] Among them, Embed data Perform spatial encoding, Embed time Perform time encoding.
[0081] In this way, the data (neighborhood values and historical values) required for physical constraints are introduced for feature enhancement. Neighborhood data provides spatial context, and historical data provides temporal context, which helps the model learn the spatiotemporal evolution law. The embedding layer integrates spatiotemporal information, so that the input features contain both the original observation values and physical relationship hints, thereby improving feature quality and thus improving prediction accuracy.
[0082] Optionally, in this embodiment of the application, when determining the sea ice prediction result of the target training sample based on the sea ice features of the initial sample, a dual-stream and subtraction mechanism module is used to perform subtraction operation on the sea ice features of the initial sample to obtain the sea ice prediction result of the target training sample.
[0083] Specifically, the initial sample sea ice features are input into a sequence of predictor modules containing n predictor modules, and the output results corresponding to each of the n predictor modules are generated step by step. The output result corresponding to the nth predictor module is used as the sample sea ice prediction result.
[0084] Where n is an integer greater than 1, each predictor module outputs a prediction result component and residual features based on its input data. The residual features serve as the input data for the next predictor module, and the initial sample sea ice features serve as the input data for the first predictor module. The output result of the nth predictor module is obtained by subtracting the output result of the (n-1)th predictor module from the prediction result component of the nth predictor module. The output result of the first predictor module is the prediction result component of the first predictor module (that is, the initial value of the output result is 0). The prediction result component is the initial direct prediction of the target sea ice physical quantities (such as SIC, SIT, SIEV, SINV, SST) by the current predictor module based on its input data. It reflects the spatiotemporal patterns learned by the current predictor module from the input data. The residual feature is the residual feature representation after processing by the current predictor module. It carries the remaining information, nonlinear correlations, and complex patterns in the input data of the current predictor module that have not been fully learned or interpreted by the current predictor module. It will be used as the input data of the next predictor module to drive subsequent residual learning and prediction correction.
[0085] In this embodiment, the dual-stream and subtraction mechanism module consists of two data streams: one is the input stream that decomposes multiple residuals through subtraction, and the other is the output stream that learns the residuals of the input stream stepwise. These output streams then pass through the predictor module (Block) to generate prediction result components and residual features. Each predictor module (Block) subtracts the output of the previous predictor module (Block) from its prediction result component. Specifically, if there is only one predictor in the predictor module sequence, the sample sea ice prediction result (Output) is represented as the predictor's prediction result component minus the initial value of the output (0), i.e., O1 = P1 - 0, where P1 is the prediction result component of the first predictor module. If there are two predictors in the predictor module sequence, the sample sea ice prediction result (Output) is represented as O2 = P2 - O1, where P2 is the prediction result component of the second predictor module, and O1 is the output of the first predictor module. If there are more predictors in the predictor module sequence, the sample sea ice prediction result can be expressed similarly as: O n =P n -O n-1 .
[0086] For example, see Figure 4The diagram shows the structure of the dual-stream and subtraction mechanism module in this embodiment. The initial sample sea ice features are input into the first predictor module (Block) in the predictor module sequence, resulting in the prediction result component P1 and residual feature X1 output by the first predictor module (Block). The initial value 0 is subtracted from the prediction result component of the first predictor module (Block), resulting in the output result O1 = P1 - 0. Then, the residual feature X1 is input into the second predictor module (Block) in the predictor module sequence, resulting in the prediction result component P2 and residual feature X2 output by the second predictor module (Block). The output result of the first predictor module (Block) is subtracted from the prediction result component of the second predictor module (Block), resulting in the output result O2 = P2 - O1. This process is repeated for each prediction result component. n-1 The input is fed into the nth predictor module Block in the predictor module sequence to obtain the prediction result component P output by the nth predictor module Block. n Subtracting the output of the (n-1)th predictor module from the prediction result of the nth predictor module block yields an output of O for the nth predictor module block. n =P n -O n-1 and O n Output is the prediction result of the sample sea ice.
[0087] Traditional Transformer models inevitably face the risk of overfitting. Overfitting occurs when a model performs well on training data but performs poorly on unseen samples due to its high complexity, leading to a significant difference between low training error and high test error. Current time series prediction models, especially very deep networks, typically contain millions of parameters. Although skip connections facilitate training deeper networks by mitigating the vanishing gradient problem, the sheer number of parameters significantly increases model complexity, making them prone to overfitting when trained on highly volatile time series datasets. Therefore, a dual-stream and subtraction mechanism module is employed. The physical meaning of the dual-stream and subtraction mechanism module is an implicit decomposition of the input and output streams, thereby reducing model complexity and mitigating the risk of overfitting. Information aggregation is changed from addition to subtraction to resist overfitting and improve model robustness. Subsequent predictor modules learn the residuals (subtle patterns) of preceding predictor modules rather than simply stacking features, reducing the risk of overfitting in complex models. Furthermore, each predictor module focuses on correcting the output of the preceding predictor module, achieving progressive refinement of the prediction results and improving accuracy.
[0088] Optionally, in this embodiment of the application, the process steps for each predictor module to output prediction result components and residual features based on its input data are as follows: SA1: Apply an attention mechanism to the input data corresponding to a predictor module to generate attention features.
[0089] In this embodiment of the application, attention features are generated by weighting the importance of different parts of the output data through an attention mechanism.
[0090] SA2: Randomly discard the attention features to obtain the discarded features.
[0091] In this embodiment of the application, a dropout layer is used to randomly set some units of the attention feature to 0 in order to prevent overfitting.
[0092] SA3: Subtract the discarded features from the input data of a predictor module to obtain the first intermediate feature.
[0093] In this embodiment of the application, a subtraction operation is performed on the input data and the discarded features to obtain a first intermediate feature, which is equal to the input data minus the discarded features.
[0094] In this way, the first intermediate feature obtained extracts the difference information between the input data and the attention features, ensuring that subsequent branches utilize both attention weights and original feature details, retaining ignored boundaries or weak signals, and preventing feature omission.
[0095] SA4: Perform first branch processing on the first intermediate feature to obtain a residual feature output by a predictor module, and perform second branch processing on the first intermediate feature to obtain a prediction result component output by a predictor module.
[0096] In this embodiment, the first branch processing includes: normalizing the first intermediate feature to obtain a normalized feature; performing a first convolution transformation on the normalized feature to obtain a first convolution feature; applying the Gelu activation function to the first convolution feature to obtain an activated feature; performing a first random discarding process on the activated feature to obtain a first discarded feature; performing a second convolution transformation on the first discarded feature to obtain a second convolution feature; subtracting the normalized feature from the second convolution feature to obtain a second intermediate feature; performing a third convolution transformation on the second intermediate feature and generating a first gating weight using the Sigmoid activation function; multiplying the first gating weight with the result of the fourth convolution transformation of the second intermediate feature to obtain a gated output feature; and finally normalizing the gated output feature to obtain a residual feature output by a predictor module.
[0097] In this way, introducing a dropout layer in the first branch processing forces the model to not rely too much on specific features, thus improving robustness.
[0098] Among them, convolutional layers are used to capture the spatial hierarchical structure in the data, Gelu activation function is used to add non-linear features, softmoid activation function compresses values between 0 and 1 to eliminate the negative impact of the attention layer, so that the input can flow smoothly to the feedforward layer, and gating mechanism is used for each neural module to autonomously adjust the speed of information transmission.
[0099] In this embodiment of the application, the second branch processing includes: concatenating the attention-weighted features with the second convolutional features in the first branch processing to obtain concatenated features; performing a fifth convolutional transformation on the concatenated features and generating a second gate weight through the Sigmoid activation function; and multiplying the second gate weight with the result of the sixth convolutional transformation of the concatenated features to obtain a prediction result component output by a predictor module. For example, see Figure 5 The diagram shown is a schematic diagram of the predictor module in an embodiment of this application. After the input data corresponding to a predictor module is input to the predictor module, the predictor module outputs the prediction result components and residual features corresponding to the predictor module.
[0100] In this way, through two processing branches, the first processing branch focuses on extracting deep implicit features from the input data that are not captured by the attention mechanism, carrying complex patterns such as spatial structure and long-term dependencies, and passing them to the subsequent predictor module as information carriers. The second processing branch directly generates the preliminary prediction value of the current predictor module for the sea ice state, focusing on the real-time estimation of interpretable physical quantities, avoiding the conflicting tasks of information transmission and result prediction by a single feature representation, and improving the model's expression efficiency.
[0101] S211: Calculate the loss value of the sample sea ice prediction results using multiple loss functions, determine the corresponding loss value of the sample sea ice prediction results, and adjust the model parameters of the polar sea ice prediction model to be trained based on the loss value.
[0102] The loss functions include a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function measures whether the predicted sea ice concentration change conforms to the law of sea ice drift in the actual environment, while the second physical constraint loss function measures whether the predicted sea ice thickness satisfies the physical law of buoyancy equilibrium. The first physical constraint loss function is constructed based on the Ice Concentration Budget Method (ICBM), which represents the physical relationship between changes in sea ice concentration and sea ice drift velocity. The second physical constraint loss function is constructed based on the Hydrostatic Equilibrium Method (HEM).
[0103] Optionally, in this embodiment of the application, a possible implementation is provided for calculating the loss value of the sample sea ice prediction result using multiple loss functions to determine the loss value corresponding to the sample sea ice prediction result, specifically by performing the following operations: S2110: Calculate multiple sub-loss values corresponding to the sea ice prediction results of the sample based on multiple loss functions.
[0104] In this embodiment of the application, the multiple loss functions include: a first physical constraint loss function and a second physical constraint function, and the multiple loss functions may also include: a mean squared error loss function.
[0105] For example, multiple loss functions include: a first physical constraint loss function and a second physical constraint function; or, for example, multiple loss functions include: a first physical constraint loss function, a second physical constraint function, and a mean squared error loss function.
[0106] Optionally, in this embodiment of the application, a loss function is used to calculate the sub-loss value corresponding to the loss function, including the following three methods: Method 1: Based on the first rate of change of the predicted sea ice concentration value with time, the second rate of change with velocity, and convergence and divergence of each predicted sample point, determine the predicted concentration constraint value of each predicted sample point, and calculate the error between the predicted concentration constraint value and the corresponding real concentration constraint value of each predicted sample point based on the first physical constraint loss function, to obtain the first sub-loss value.
[0107] The true concentration constraint value is determined based on the real sea ice state data of the target training samples. Specifically, based on the real sea ice state data, the third rate of change of the predicted sea ice concentration value over time, the fourth rate of change of the predicted sea ice concentration value over velocity, and the convergence and divergence are calculated for each predicted sample point. Then, based on the third rate of change of the predicted sea ice concentration value over time, the fourth rate of change of the predicted sea ice concentration value over velocity, and the convergence and divergence, the true concentration constraint value for each predicted sample point is determined.
[0108] In this embodiment of the application, the mathematical expression of the first physical constraint loss function is:
[0109] Where N is the number of each predicted sample point, and each predicted sample point is represented by i, satisfying i=1,2,3…,N, y(ICBM) i It is the true concentration constraint value of the i-th predicted sample point. It is the predicted concentration constraint value for the i-th predicted sample point.
[0110] In this embodiment of the application, based on the first rate of change of the predicted sea ice concentration value with time, the second rate of change with velocity, and convergence and divergence of each predicted sample point, the predicted concentration constraint value of each predicted sample point is determined. The mathematical expression for the predicted concentration constraint value of a predicted sample point is as follows:
[0111] in, C i / t represents the first rate of change of the predicted sea ice concentration value at the i-th predicted sample point over time. Let be the second rate of change of the predicted sea ice concentration value for the i-th predicted sample point as a function of velocity. Let be the convergence and divergence of the i-th predicted sample point.
[0112] In this embodiment, the mathematical expression for the predicted concentration constraint value of a predicted sample point is constructed using the variable relationship mapped by the formula of the ICBM method. The formula of the ICBM method is:
[0113] Where C represents sea ice concentration and U represents sea ice drift velocity. C / t represents the rate of change of sea ice concentration over time. This is the gradient operator, and `residual` is the residual term. The residual term mainly covers the effects caused by thermodynamic processes (such as melting and freezing) and mechanical redistribution. The residual term has a small impact on the rate of change of sea ice concentration over time. Therefore, this residual term is ignored, and the focus is on the impact of sea ice drift velocity on the rate of change. A mathematical expression for the predicted concentration constraint value of the predicted sample points is constructed based on the formula of the ICBM method.
[0114] Optionally, in this embodiment, the time derivative of the predicted sea ice concentration value of a predicted sample point is calculated using the finite difference method to obtain the first rate of change. The finite difference method includes backward difference, forward difference, and central difference. In this embodiment, the backward difference method is preferentially chosen to calculate the time derivative of the predicted sea ice concentration value of a predicted sample point to obtain the first rate of change. Specifically, based on the predicted sea ice concentration value of a predicted sample point, the predicted sea ice concentration value of the previous time step of the predicted sample point, and the time step, the first rate of change of the predicted sea ice concentration value of the predicted sample point over time is determined. The predicted sea ice concentration value of the previous time step of the predicted sample point can be obtained from the predicted sea ice data; that is, the predicted sea ice concentration value of the predicted sample point at the previous adjacent time step is used as the predicted sea ice concentration value of the previous time step of the predicted sample point.
[0115] In this embodiment of the application, the mathematical expression for the first rate of change of the predicted sea ice concentration value of a predicted sample point over time is:
[0116] Among them, C i (t) represents the predicted sea ice concentration value for the i-th predicted sample point, C i (t-△t) represents the predicted sea ice concentration value of the i-th predicted sample point at the previous moment, where △t is the time step, i.e., the difference between the time corresponding to the predicted sample point and the previous moment. The time step can be 1 day. For example, if the time corresponding to the i-th predicted sample point is Tuesday, then the time before the i-th predicted sample point is Monday. If the time corresponding to the i-th predicted sample point is the start time, and there is no data for the time before the i-th predicted sample point, then the predicted sea ice concentration value of the time before the i-th predicted sample point is the predicted sea ice concentration value of the i-th predicted sample point.
[0117] Optionally, in this embodiment of the application, based on the predicted sea ice concentration values of each of the adjacent predicted sample points that are adjacent to a predicted sample point in the east, west, north, and south, the longitude partial derivative of the predicted sea ice concentration value of a predicted sample point in the longitude direction and the latitude partial derivative of the predicted sea ice concentration value in the latitude direction are determined. Based on the predicted eastward velocity value and the predicted northward velocity value of a predicted sample point, as well as the longitude partial derivative and the latitude partial derivative of the concentration, a second rate of change of the predicted sea ice concentration value of a predicted sample point with respect to velocity is determined.
[0118] Among them, the adjacent predicted sample points that are adjacent to a predicted sample point in the east, west, north, and south are: in the grid, the east adjacent predicted sample point to the east of a predicted sample point, the west adjacent predicted sample point to the west of a predicted sample point, the south adjacent predicted sample point to the south of a predicted sample point, and the north adjacent predicted sample point to the north of a predicted sample point.
[0119] Optionally, in this embodiment of the application, when determining the longitude partial derivative of the predicted sea ice concentration of a predicted sample point in the longitude direction and the latitude partial derivative of the predicted sea ice concentration in the latitude direction based on the predicted sea ice concentration values of each of the adjacent predicted sample points in the east, west, north, and south directions of a predicted sample point, the following operations are performed: Based on the difference in predicted sea ice concentration between the eastern and western adjacent predicted sample points of a predicted sample point, and the difference in longitude between the eastern and western adjacent predicted sample points, the longitude partial derivative of the predicted sea ice concentration of a predicted sample point in the longitude direction is determined; Based on the difference in predicted sea ice concentration between the northern and southern adjacent predicted sample points of a predicted sample point, and the difference in latitude between the northern and southern adjacent predicted sample points, the latitude partial derivative of the predicted sea ice concentration of a predicted sample point in the latitude direction is determined.
[0120] In this embodiment of the application, the mathematical expression for the partial derivative of the predicted sea ice concentration value of a predicted sample point in the longitude direction is:
[0121] Where C(m+1,n,t) is the predicted sea ice concentration value of the eastern adjacent predicted sample point of the i-th predicted sample point, C(m-1,n,t) is the predicted sea ice concentration value of the western adjacent predicted sample point of the i-th predicted sample point, and λ[m+1]-λ[m-1] is the longitude difference between the eastern and western adjacent predicted sample points. In three-dimensional space, the concentration field, eastward velocity field, and northward velocity field are represented as C(m,n,t), U, and U, respectively. λ (m,n,t) and U (m,n,t).
[0122] In this embodiment of the application, the mathematical expression for the latitudinal partial derivative of the predicted sea ice concentration value of a predicted sample point in the latitudinal direction is:
[0123] Where C(m,n+1,t) is the predicted sea ice concentration value of the northern neighboring predicted sample point of the i-th predicted sample point, and C(m,n-1,t) is the predicted sea ice concentration value of the southern neighboring predicted sample point of the i-th predicted sample point. [n+1]- [n-1] represents the latitude difference between the northern and southern adjacent predicted sample points.
[0124] Optionally, in this embodiment of the application, when determining the second rate of change of the predicted sea ice concentration value of a predicted sample point with respect to velocity based on the predicted eastward velocity value and the predicted northward velocity value, as well as the concentration longitude partial derivative and the concentration latitude partial derivative of a predicted sample point, the following operations are performed: based on the predicted eastward velocity value and the concentration longitude partial derivative of a predicted sample point, determine the first rate of change component of the predicted sea ice concentration value of a predicted sample point with respect to velocity in the longitude direction; based on the predicted northward velocity value and the concentration latitude partial derivative of a predicted sample point, determine the second rate of change component of the predicted sea ice concentration value of a predicted sample point with respect to velocity in the latitude direction; based on the first rate of change component and the second rate of change component, determine the second rate of change of the predicted sea ice concentration value of a predicted sample point with respect to velocity.
[0125] In this embodiment of the application, the mathematical expression for the second rate of change of the predicted sea ice concentration value of a predicted sample point with respect to velocity is:
[0126] in, Ci / λ is the longitude partial derivative of the concentration at the i-th predicted sample point. C i / U is the latitudinal partial derivative of the concentration at the i-th predicted sample point. i λ Let U be the predicted eastward velocity value for the i-th predicted sample point. i Let be the predicted northward velocity value for the i-th predicted sample point.
[0127] In this way, the advection effect of sea ice velocity on sea ice concentration is orthogonally decomposed along the longitude and latitude directions, preventing the excessive influence of errors in a single direction on the second rate of change, ensuring the accuracy of the second rate of change, and thus ensuring the accuracy of the first physical constraint loss function. This enables the model to submit the accuracy of the loss value calculation and better capture the anisotropic motion characteristics.
[0128] Optionally, in this embodiment of the application, based on the predicted eastward velocity values of each adjacent predicted sample point that is east-west adjacent to a predicted sample point, and the predicted northward velocity values of each adjacent predicted sample point that is north-south adjacent to a predicted sample point, the longitudinal partial derivative of the predicted eastward velocity value of a predicted sample point in the longitude direction and the latitudinal partial derivative of the predicted northward velocity value in the latitude direction are determined. Based on the longitudinal and latitudinal partial derivatives of the velocity, the velocity field divergence of a predicted sample point is determined. Based on the predicted sea ice concentration value and the velocity field divergence of a predicted sample point, the convergence and divergence of a predicted sample point are determined.
[0129] Among them, the adjacent predicted sample points that are east-west adjacent to a predicted sample point are: in the grid, the east-adjacent predicted sample point to the east of a predicted sample point and the west-adjacent predicted sample point to the west of a predicted sample point; the adjacent predicted sample points that are north-south adjacent to a predicted sample point are: in the grid, the south-adjacent predicted sample point to the south of a predicted sample point and the north-adjacent predicted sample point to the north of a predicted sample point.
[0130] In this embodiment, based on the predicted eastward velocity values of each adjacent predicted sample point that is east-west adjacent to a predicted sample point, the longitudinal partial derivative of the predicted eastward velocity value of a predicted sample point in the longitude direction is determined. Based on the predicted northward velocity values of each adjacent predicted sample point that is north-south adjacent to a predicted sample point, the latitudinal partial derivative of the predicted northward velocity value of a predicted sample point in the latitude direction is determined. Then, based on the longitudinal and latitudinal partial derivatives of the velocity, the velocity field divergence of a predicted sample point is determined. After obtaining the velocity field divergence of a predicted sample point, the product of the predicted sea ice concentration value and the velocity field divergence of a predicted sample point is calculated to obtain the convergence and divergence of a predicted sample point.
[0131] In this embodiment of the application, the mathematical expression for the velocity field divergence of a predicted sample point is:
[0132] in, U λ (m,n,t) / λ is the partial derivative of the velocity longitude of the i-th predicted sample point. U (m,n,t) / Let U be the latitude partial derivative of the velocity of the i-th predicted sample point. λ (m+1,n,t) represents the predicted eastward velocity value of the eastern neighboring predicted sample point of the i-th predicted sample point, U λ (m-1,n,t) represents the predicted eastward velocity value of the western neighboring predicted sample point of the i-th predicted sample point, U (m,n+1,t) represents the predicted northward velocity value of the northern neighboring predicted sample point of the i-th predicted sample point, U (m,n-1,t) represents the predicted southward velocity value of the southern neighboring predicted sample point of the i-th predicted sample point, and λ[m+1]-λ[m-1] represents the longitude difference between the eastern and western neighboring predicted sample points. [n+1]- [n-1] represents the latitude difference between the northern and southern adjacent predicted sample points.
[0133] In this way, by calculating spatial partial derivatives (e.g., concentration gradient, velocity gradient) using the finite difference method, the ICBM method's formula is transformed into the first physical constraint loss function. This ensures the accuracy of the first physical constraint term calculation and lays the technical foundation for the effectiveness of the first physical constraint mechanism, increasing the model's interpretability based on the laws of sea ice drift. Furthermore, it transforms deep learning from purely data-driven to physics-guided data-driven learning. During training, the model not only needs to fit the data but also must obey physical laws, significantly improving the physical rationality and interpretability of the prediction results.
[0134] Method 2: Based on the predicted sea ice thickness and predicted surface snow thickness values of each predicted sample point, combined with seawater density, sea ice density and snow density, determine the predicted free plate height of each predicted sample point, and calculate the error between the predicted free plate height and the corresponding real free plate height of each predicted sample point based on the second physical constraint loss function, to obtain the second sub-loss value.
[0135] The true freeboard height is determined based on real sea ice state data from the target training samples. Specifically, the true freeboard height for each predicted sample point is determined by combining the true sea ice thickness and true surface snow thickness values with seawater density, sea ice density, and snow density.
[0136] In this embodiment of the application, the mathematical expression of the second physical constraint loss function is:
[0137] Where N is the number of each predicted sample point, y(h f ) i It is the true free plate height of the i-th predicted sample point. It is the predicted free plate height of the i-th predicted sample point.
[0138] In this embodiment, based on the predicted sea ice thickness and predicted surface snow thickness of a predicted sample point, and combined with seawater density, sea ice density, and snow cover density, the predicted free plate height of a predicted sample point is determined. The mathematical expression for the predicted free plate height of a predicted sample point is:
[0139] Among them, h i h is the predicted sea ice thickness value for the i-th predicted sample point. s Let ρ be the predicted surface snow thickness value for the i-th predicted sample point. w ρ is the density of seawater. i ρ is the density of sea ice. s Let be the density of snow cover. Among them, the density of seawater, the density of sea ice, and the density of snow cover are constants.
[0140] In this embodiment, the mathematical expression for the predicted free plate height of a predicted sample point is constructed using the variable relationship mapped by the formula of the HEM method. The formula of the HEM method is:
[0141] Among them, h f The freeboard height refers to the height from sea level to the top of sea ice and snow. A mathematical expression for predicting the freeboard height of a sample point is constructed based on the formula of the HEM method.
[0142] In this way, the formula of the HEM method is transformed into a second physical constraint loss function, ensuring the accuracy of the calculation of the second physical constraint term and laying a technical foundation for the effectiveness of the second physical constraint mechanism. This increases the interpretability of the model based on the laws of buoyancy balance. Furthermore, it transforms deep learning from purely data-driven to physics-guided data-driven learning. During training, the model not only needs to fit the data but also must obey physical laws, significantly improving the physical rationality and interpretability of the prediction results.
[0143] Method 3: Based on the mean squared error loss function, calculate the error between the sea ice prediction results of the sample and the real sea ice state data of the target training sample to obtain the third sub-loss value.
[0144] In this embodiment of the application, based on the mean squared error loss function, the error between the sea ice prediction result and the actual sea ice state data of each predicted sample point is calculated to obtain the third sub-loss value.
[0145] In this embodiment of the application, the mathematical expression for the mean squared error loss function is:
[0146] Where N is the number of each predicted sample point, y i It is the actual sea ice state data of the i-th predicted sample point. It is the sea ice prediction result for the i-th prediction sample point.
[0147] Furthermore, it is worth noting that in this embodiment, the mean squared error of each output variable predicted by the model can be calculated separately, and then the mean squared errors of each output variable can be weighted and averaged to obtain the third sub-loss value. The sea ice prediction result includes all output variables predicted by the model.
[0148] For example, assuming the output variables predicted by the model include: predicted sea ice concentration, predicted sea ice thickness, predicted eastward velocity, predicted northward velocity, and predicted surface snow thickness, the mean squared error of the predicted sea ice concentration is calculated (i.e., the error between the predicted sea ice concentration value and the actual predicted sea ice concentration value for each predicted sample point is calculated) to obtain the mean squared error of the predicted sea ice concentration. The mean squared error of the predicted sea ice thickness is calculated, the mean squared error of the predicted eastward velocity is calculated, the mean squared error of the predicted northward velocity is calculated, and the mean squared error of the predicted surface snow thickness is calculated. Finally, the mean squared errors of each output variable are weighted and averaged to obtain the third sub-loss value.
[0149] S2111: Summing multiple sub-loss values yields the total loss value.
[0150] In this embodiment of the application, multiple sub-loss values are summed according to the coefficients corresponding to each of the multiple loss functions to obtain the loss value. After obtaining the loss value, the model parameters are adjusted based on the loss value.
[0151] For example, taking multiple loss functions including: the first physical constraint loss function, the second physical constraint loss function, and the mean squared error loss function, the mathematical expression for the loss value is as follows: Loss=MSE+α×ICBM_MSE+β×HEM_MSE Where α is the coefficient of the first physical constraint loss function, ICBM_MSE is the first sub-loss value corresponding to the first physical constraint loss function, β is the coefficient of the second physical constraint loss function, HEM_MSE is the second sub-loss value corresponding to the second physical constraint loss function, and MSE is the third sub-loss value corresponding to the mean squared error loss function. α and β are both variable parameters.
[0152] In this way, the model parameters of the polar sea ice prediction model are adjusted by combining the three losses optimized using gradient descent in the above formula, and this process is repeated until convergence. The loss corresponding to the first physical constraint loss function ensures that the model's prediction is not only based on data-driven methods but also follows the basic physical laws of sea ice drift. The loss corresponding to the second physical constraint loss function ensures that the model's prediction follows the physical laws of buoyancy balance. The loss corresponding to the mean squared error loss function ensures the error between the sample predicted value and the sample true value. By combining multiple loss terms, data-driven and physical constraints are combined, ensuring that the model output follows physical laws (e.g., sea ice density prediction method and buoyancy balance method), enhancing the model's sensitivity to spatial and temporal changes, and improving the prediction's generalization ability and interpretability. In addition, balancing multiple loss terms can prevent the optimization process from getting trapped in local minima and accelerate convergence.
[0153] Furthermore, after obtaining the trained target polar sea ice prediction model, sea ice prediction is performed using the target polar sea ice prediction model. Specifically, a sea ice prediction request for the target prediction period is received, then the sea ice state data for the target historical period is obtained, and the sea ice state data for the target historical period is input into the target polar sea ice prediction model to obtain the target polar sea ice prediction result output by the target polar sea ice prediction model.
[0154] Furthermore, in this embodiment, based on the result requirement information, target display results that meet the result requirement information are filtered from the target sea ice prediction results. The target display results can then be further displayed and / or fed back to the target object. The result requirement information can be obtained by parsing from the sea ice prediction request or can be pre-configured; this embodiment does not impose any limitations on this.
[0155] For example, assuming the target prediction period is the next 7 days and the required information is SIC and SIT, the sea ice state data of the previous 2 weeks is obtained and input into the target polar sea ice prediction model to obtain the target sea ice prediction results output by the target polar sea ice prediction model. The target sea ice prediction results include: the predicted sea ice concentration, the predicted sea ice thickness, the predicted surface snow thickness, the predicted eastward velocity, and the predicted northward velocity for the next 7 days. Then, based on the required information, the target display results that meet the required information are selected from the target sea ice prediction results. The target display results include: the predicted sea ice concentration and the predicted sea ice thickness for the next 7 days.
[0156] In this way, using a target polar sea ice prediction model for polar sea ice prediction can improve the computational efficiency and prediction accuracy of polar sea ice prediction, and is applicable to polar scenarios. Furthermore, dynamically determining the target display results based on the required information can enhance the user experience.
[0157] Based on the above embodiments, it can be seen that the training method of the polar sea ice prediction model provided by this application has the following beneficial effects: (1) It has higher prediction accuracy and generalization ability: Existing technologies often fail to make full use of the physical correlation between variables when processing multivariate time series data, resulting in insufficient accuracy. This application significantly improves the accuracy of the joint prediction of SIC and SIT by screening high-quality variables, reconstructing and using physical constraint loss functions. Experimental results show that this invention can capture the complex relationship between variables, which is better than the traditional "black box" deep learning model and surpasses the limitations of physical driving methods in generalization performance. (2) It enhances interpretability and physical consistency: Existing deep learning schemes are mostly "black box" and lack physical interpretability, while physical driving methods, although focusing on process simulation, have poor adaptability. This application innovatively integrates physical laws (such as ICBM and HEM methods) into the loss function, making the model output more consistent with the real physical process, improving interpretability and scientific rationality. Compared with similar schemes such as gray box models, this invention is more optimized for multivariate scenarios of polar sea ice. (3) Better resistance to overfitting and robustness: Existing time series prediction models are susceptible to overfitting, especially in multivariate prediction with mismatched sample sizes. This application introduces a dual-stream and subtraction mechanism, changing information aggregation from addition to subtraction, effectively mitigating the risk of overfitting and improving the stability of the model in complex polar environments. This is superior to solutions that rely solely on CNN or LSTM, which may exacerbate overfitting due to blindly increasing depth. (4) More efficient data processing and application potential: Existing methods are limited in data acquisition and feature extraction. This application provides efficient data support through the systematic processing of deep learning algorithms, overcoming the linear assumptions of statistical models and the data dependence problem of physical models. At the same time, this invention is applicable to short-term sea ice prediction, supports real-time route planning, and ensures navigation safety and economic benefits. Compared with existing solutions, it has stronger practicality and promotional value. Overall, this application overcomes the core shortcomings of existing technologies through the deep integration of physical features and deep learning, achieving a comprehensive improvement in prediction accuracy, generalization ability, interpretability, and robustness, and has significant technological progress and application advantages.
[0158] Furthermore, based on the same technical concept, embodiments of this application provide a training device for a polar sea ice prediction model, which is used to implement the above-described method flow of embodiments of this application. For example, see [link to relevant documentation]. Figure 6 As shown, the training device 600 for the polar sea ice prediction model may include: an acquisition module 601 and a training module 602, wherein: The acquisition module 601 is used to acquire a training sample set. Each training sample in the training sample set includes: historical sea ice state data for the historical period of the sample and actual sea ice state data for the predicted period of the sample. Training module 602 is used to iteratively train the polar sea ice prediction model to be trained based on the training sample set to obtain the target polar sea ice prediction model. During one iteration of training, the following operations are performed: Feature extraction is performed on the historical sea ice state data of the selected target training samples to obtain initial sample sea ice features. Based on the initial sample sea ice features, the sample sea ice prediction results of the target training samples are determined. The sample sea ice prediction results include: predicted sea ice concentration and predicted sea ice thickness. Multiple loss functions are used to calculate the loss value of the sample sea ice prediction results, determine the corresponding loss value of the sample sea ice prediction results, and adjust the model parameters of the polar sea ice prediction model to be trained based on the loss value. Among them, the multiple loss functions include: a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function is used to measure whether the change of predicted sea ice concentration conforms to the law of sea ice drift in the actual environment, and the second physical constraint loss function is used to measure whether the predicted sea ice thickness satisfies the physical law of buoyancy balance.
[0159] In an optional embodiment, when calculating the loss value of the sample sea ice prediction result using multiple loss functions and determining the loss value corresponding to the sample sea ice prediction result, the training module 602 is further configured to: Based on multiple loss functions, calculate multiple sub-loss values corresponding to the sea ice prediction results of the samples; The loss value is obtained by summing the multiple sub-loss values.
[0160] In an optional embodiment, the predicted sea ice concentration includes: the predicted sea ice concentration value of each predicted sample point; the predicted sea ice thickness includes: the predicted sea ice thickness value of each predicted sample point; and the sample sea ice prediction result further includes: the predicted surface snow thickness, which includes: the predicted sea ice thickness value of each predicted sample point. When calculating multiple sub-loss values corresponding to the sea ice prediction results of the samples based on multiple loss functions, the training module 602 is also used for: Based on the first rate of change of the predicted sea ice concentration value with time, the second rate of change with velocity, and convergence and divergence of each predicted sample point, the predicted concentration constraint value of each predicted sample point is determined. Based on the first physical constraint loss function, the error between the predicted concentration constraint value of each predicted sample point and the corresponding real concentration constraint value is calculated to obtain the first sub-loss value. The real concentration constraint value is determined based on the real sea ice state data of the target training sample. Based on the predicted sea ice thickness and predicted surface snow thickness values of each predicted sample point, and combined with seawater density, sea ice density, and snow density, the predicted free plate height of each predicted sample point is determined. Based on the second physical constraint loss function, the error between the predicted free plate height of each predicted sample point and the corresponding real free plate height is calculated to obtain the second sub-loss value. The real free plate height is determined based on the real sea ice state data of the target training sample.
[0161] In an optional embodiment, the multiple loss functions further include: a mean squared error loss function; the training module 602 is also used for: Based on the mean squared error loss function, the error between the sea ice prediction results of the sample and the real sea ice state data of the target training sample is calculated to obtain the third sub-loss value.
[0162] In an optional embodiment, the sea ice prediction results for the sample further include: predicted eastward velocity and predicted northward velocity, wherein the predicted eastward velocity includes the predicted eastward velocity value for each predicted sample point, and the predicted northward velocity includes the predicted northward velocity value for each predicted sample point. Before determining the predicted concentration constraint value for each predicted sample point based on the first rate of change of the predicted sea ice concentration value over time, the second rate of change of the predicted sea ice concentration value over velocity, and convergence and divergence, the training module 602 is also used for: For each predicted sample point, perform the following operations: Based on the predicted sea ice concentration value of a predicted sample point, the predicted sea ice concentration value of the previous moment of a predicted sample point, and the time step, determine the first rate of change of the predicted sea ice concentration value of a predicted sample point over time. Based on the predicted sea ice concentration values of each of the neighboring predicted sample points in the east, west, north, and south directions of a predicted sample point, the longitude partial derivative of the predicted sea ice concentration value of a predicted sample point in the longitude direction and the latitude partial derivative of the predicted sea ice concentration value in the latitude direction are determined. Based on the predicted eastward velocity value and the predicted northward velocity value of a predicted sample point, as well as the longitude partial derivative and the latitude partial derivative of the concentration, the second rate of change of the predicted sea ice concentration value of a predicted sample point with velocity is determined. Based on the predicted eastward velocity values of each adjacent predicted sample point to the east and west, and the predicted northward velocity values of each adjacent predicted sample point to the north and south, the longitudinal partial derivative of the predicted eastward velocity value of a predicted sample point in the longitude direction and the latitudinal partial derivative of the predicted northward velocity value in the latitude direction are determined. Based on the longitudinal and latitudinal partial derivatives of the velocity, the velocity field divergence of a predicted sample point is determined. Based on the predicted sea ice concentration value and the velocity field divergence of a predicted sample point, the convergence and divergence of a predicted sample point are determined.
[0163] In an optional embodiment, when determining a second rate of change of the predicted sea ice concentration value as a function of velocity for a predicted sample point based on the predicted eastward velocity value and the predicted northward velocity value, as well as the concentration longitude partial derivative and the concentration latitude partial derivative, the training module 602 is further configured to: Based on the predicted eastward velocity value and the longitude partial derivative of the concentration at a predicted sample point, the first rate of change component of the predicted sea ice concentration value at the predicted sample point in the longitude direction with velocity is determined. Based on the predicted northward velocity value and the latitudinal partial derivative of the concentration at a predicted sample point, the second rate of change component of the predicted sea ice concentration value at a predicted sample point in the latitudinal direction with velocity is determined. Based on the first and second rate of change components, a second rate of change is determined for the predicted sea ice concentration value of a predicted sample point as a function of velocity.
[0164] In one optional embodiment, the historical sea ice state data of the target training sample includes: historical sea ice concentration value, historical sea ice thickness value, historical eastward velocity value, historical northward velocity value, and historical surface snow thickness value for each historical sample point; When extracting features from the historical sea ice state data of the selected target training samples to obtain the initial sample sea ice features, the training module 602 is also used for: The data required for physical constraints are extracted from the historical sea ice state data of the target training samples. The data required for physical constraints include: the historical sea ice concentration values of each historical sample point corresponding to its east, west, north, and south adjacent historical sample points; the historical sea ice concentration value of each historical sample point corresponding to the previous moment; the historical eastward velocity values of each historical sample point corresponding to its east, west, and south adjacent historical sample points; and the historical northward velocity values of each historical sample point corresponding to its north, south, and south adjacent historical sample points. The data required for physical constraints are concatenated with the historical sea ice state data of the target training samples to form an enhanced input vector; The initial sample sea ice features are generated by fusing spatial and temporal information into the enhanced input vector through an embedding layer.
[0165] In an optional embodiment, when determining the sea ice prediction result of the target training sample based on the initial sample sea ice features, the training module 602 is further configured to: The initial sample sea ice features are input into a sequence of predictor modules containing n predictor modules, and the output results corresponding to each of the n predictor modules are generated step by step. The output result corresponding to the nth predictor module is used as the sample sea ice prediction result. Where n is an integer greater than 1, each predictor module outputs a prediction result component and residual features based on its input data. The residual features serve as the input data for the next predictor module, and the initial sample sea ice features serve as the input data for the first predictor module. The output result of the nth predictor module is obtained by subtracting the output result of the (n-1)th predictor module from the prediction result component of the nth predictor module. The output result of the first predictor module is the prediction result component of the first predictor module.
[0166] In an optional embodiment, the training module 602 is further configured to: An attention mechanism is applied to the input data corresponding to a predictor module to generate attention features; The attention features are randomly discarded to obtain the discarded features; Subtract the discarded features from the input data corresponding to a predictor module to obtain the first intermediate feature; The first intermediate feature is processed by the first branch to obtain a residual feature output by a predictor module, and the first intermediate feature is processed by the second branch to obtain a prediction result component output by a predictor module.
[0167] Based on the description of the method and apparatus embodiments above, an exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method according to an embodiment of the present invention.
[0168] This application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.
[0169] This application also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.
[0170] See Figure 7The diagram shown below illustrates the structure of an electronic device 700 that can serve as a server or client in this application, and is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0171] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0172] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 707 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth devices, WiFi devices, worldwide interoperability for microwave access (WiMax) devices, cellular communication devices, and / or the like.
[0173] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above. For example, in some embodiments, the training method of the polar sea ice prediction model described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. In some embodiments, the computing unit 701 can be configured to perform the training method of the polar sea ice prediction model described above by any other suitable means (e.g., by means of firmware).
[0174] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0175] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM) or flash memory, optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0176] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device, PLD) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0177] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0178] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0179] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0180] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of the invention. Therefore, any equivalent variations made in accordance with the claims of this invention are still within the scope of this application.
Claims
1. A training method for a polar sea ice prediction model, characterized in that, include: Obtain a training sample set, wherein each training sample in the training sample set includes: historical sea ice state data for the historical period of the sample and actual sea ice state data for the predicted period of the sample. Based on the training sample set, the polar sea ice prediction model to be trained is iteratively trained to obtain the target polar sea ice prediction model. During each iteration of training, the following operations are performed: Feature extraction is performed on the historical sea ice state data of the selected target training sample to obtain initial sample sea ice features, and based on the initial sample sea ice features, the sample sea ice prediction result of the target training sample is determined, wherein the sample sea ice prediction result includes: predicted sea ice concentration and predicted sea ice thickness. The loss values of the sample sea ice prediction results are calculated using multiple loss functions to determine the corresponding loss values. Based on these loss values, the model parameters of the polar sea ice prediction model to be trained are adjusted. The multiple loss functions include a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function measures whether the change in the predicted sea ice concentration conforms to the law of sea ice drift in the actual environment, and the second physical constraint loss function measures whether the predicted sea ice thickness satisfies the physical law of buoyancy balance.
2. The method as described in claim 1, characterized in that, The step of calculating the loss value of the sample sea ice prediction result using multiple loss functions to determine the loss value corresponding to the sample sea ice prediction result includes: Based on the multiple loss functions, calculate multiple sub-loss values corresponding to the sea ice prediction results of the samples; The loss value is obtained by summing the multiple sub-loss values.
3. The method as described in claim 2, characterized in that, The predicted sea ice concentration includes: the predicted sea ice concentration value of each predicted sample point; the predicted sea ice thickness includes: the predicted sea ice thickness value of each predicted sample point; the sample sea ice prediction result also includes: the predicted surface snow thickness; the predicted surface snow thickness includes: the predicted sea ice thickness value of each predicted sample point. The calculation of multiple sub-loss values corresponding to the sample sea ice prediction results based on the multiple loss functions includes: Based on the first rate of change of the predicted sea ice concentration value over time, the second rate of change of the predicted sea ice concentration value over velocity, and the convergence and divergence of each predicted sample point, the predicted concentration constraint value of each predicted sample point is determined. Based on the first physical constraint loss function, the error between the predicted concentration constraint value and the corresponding real concentration constraint value of each predicted sample point is calculated to obtain the first sub-loss value. The real concentration constraint value is determined based on the real sea ice state data of the target training sample. Based on the predicted sea ice thickness and predicted surface snow thickness values of each predicted sample point, and combined with seawater density, sea ice density, and snow density, the predicted free plate height of each predicted sample point is determined. Based on the second physical constraint loss function, the error between the predicted free plate height of each predicted sample point and the corresponding real free plate height is calculated to obtain the second sub-loss value. The real free plate height is determined based on the real sea ice state data of the target training sample.
4. The method as described in claim 3, characterized in that, The plurality of loss functions further includes: a mean squared error loss function; the method further includes: Based on the mean squared error loss function, the error between the sea ice prediction result of the sample and the real sea ice state data of the target training sample is calculated to obtain the third sub-loss value.
5. The method as described in claim 3, characterized in that, The sea ice prediction results for the sample also include: predicted eastward velocity and predicted northward velocity. The predicted eastward velocity includes the predicted eastward velocity value for each predicted sample point, and the predicted northward velocity includes the predicted northward velocity value for each predicted sample point. Before determining the predicted concentration constraint value for each predicted sample point based on the first rate of change of the predicted sea ice concentration value over time, the second rate of change of the predicted sea ice concentration value over velocity, and convergence and divergence, the method further includes: For each of the predicted sample points, perform the following operations: Based on the predicted sea ice concentration value of a predicted sample point, the predicted sea ice concentration value of the previous moment of the predicted sample point, and the time step, a first rate of change of the predicted sea ice concentration value of the predicted sample point over time is determined. Based on the predicted sea ice concentration values of each adjacent predicted sample point in the east, west, north, and south directions of the predicted sample point, the longitude partial derivative of the predicted sea ice concentration value of the predicted sample point in the longitude direction and the latitude partial derivative of the predicted sea ice concentration value in the latitude direction are determined. Based on the predicted eastward velocity value and the predicted northward velocity value of the predicted sample point, as well as the longitude partial derivative of the concentration and the latitude partial derivative of the concentration, a second rate of change of the predicted sea ice concentration value of the predicted sample point with velocity is determined. Based on the predicted eastward velocity values of each adjacent predicted sample point to the east and west, and the predicted northward velocity values of each adjacent predicted sample point to the north and south, the longitude partial derivative of the predicted eastward velocity value and the latitude partial derivative of the predicted northward velocity value of the predicted sample point are determined. Based on the longitude and latitude partial derivatives, the velocity field divergence of the predicted sample point is determined. Based on the predicted sea ice concentration value and the velocity field divergence, the convergence and divergence of the predicted sample point are determined.
6. The method as described in claim 5, characterized in that, The determination of the second rate of change of the predicted sea ice concentration value as a function of velocity for a given predicted sea ice sample point, based on the predicted eastward and northward velocity values of the predicted sample point, and the longitude and latitudinal partial derivatives of the concentration, includes: Based on the predicted eastward velocity value of the predicted sample point and the longitude partial derivative of the concentration, the first rate of change component of the predicted sea ice concentration value of the predicted sample point in the longitude direction with velocity is determined. Based on the predicted northward velocity value of the predicted sample point and the latitudinal partial derivative of the concentration, the second rate of change component of the predicted sea ice concentration value of the predicted sample point in the latitudinal direction with velocity is determined; Based on the first rate of change component and the second rate of change component, a second rate of change of the predicted sea ice concentration value of the predicted sample point as a function of velocity is determined.
7. The method as described in claim 1, characterized in that, The historical sea ice state data of the target training samples include: historical sea ice concentration, historical sea ice thickness, historical eastward velocity, historical northward velocity, and historical surface snow thickness for each historical sample point; The step of extracting features from the historical sea ice state data of the selected target training samples to obtain the initial sample sea ice features includes: The data required for physical constraints are extracted from the historical sea ice state data of the target training sample. The data required for physical constraints includes: the historical sea ice concentration values of each historical sample point corresponding to its east, west, north, and south adjacent historical sample points; the historical sea ice concentration value of each historical sample point corresponding to its previous moment; the historical eastward velocity values of each historical sample point corresponding to its east, west, and south adjacent historical sample points; and the historical northward velocity values of each historical sample point corresponding to its north, south, and south adjacent historical sample points. The data required for the physical constraints are concatenated with the historical sea ice state data of the target training sample to form an enhanced input vector; The initial sample sea ice features are generated by fusing spatial and temporal information into the enhanced input vector through an embedding layer.
8. The method as described in claim 1, characterized in that, The step of determining the sea ice prediction result of the target training sample based on the initial sample sea ice features includes: The initial sample sea ice features are input into a sequence of predictor modules containing n predictor modules, and the output results corresponding to each of the n predictor modules are generated step by step. The output result corresponding to the nth predictor module is used as the sample sea ice prediction result. Where n is an integer greater than 1, each predictor module outputs a prediction result component and residual features based on its input data. The residual features serve as the input data for the next predictor module, and the initial sample sea ice features serve as the input data for the first predictor module. The output result corresponding to the nth predictor module is obtained by subtracting the output result corresponding to the (n-1)th predictor module from the prediction result component of the nth predictor module. The output result corresponding to the first predictor module is the prediction result component of the first predictor module.
9. The method as described in claim 8, characterized in that, Each predictor module outputs prediction result components and residual features based on its input data, including: An attention mechanism is applied to the input data corresponding to a predictor module to generate attention features; The attention features are randomly discarded to obtain the discarded features; Subtract the discarded features from the input data corresponding to the predictor module to obtain the first intermediate feature; The first intermediate feature is processed by a first branch to obtain the residual feature output by the predictor module, and the first intermediate feature is processed by a second branch to obtain the prediction result component output by the predictor module.
10. A training device for a polar sea ice prediction model, characterized in that, include: The acquisition module is used to acquire a training sample set, wherein each training sample in the training sample set includes: historical sea ice state data for the historical period of the sample and actual sea ice state data for the predicted period of the sample. The training module is used to iteratively train the polar sea ice prediction model to be trained based on the training sample set to obtain the target polar sea ice prediction model. During one iteration of training, the following operations are performed: Feature extraction is performed on the historical sea ice state data of the selected target training sample to obtain initial sample sea ice features, and based on the initial sample sea ice features, the sample sea ice prediction result of the target training sample is determined, wherein the sample sea ice prediction result includes: predicted sea ice concentration and predicted sea ice thickness. The loss values of the sample sea ice prediction results are calculated using multiple loss functions to determine the corresponding loss values. Based on these loss values, the model parameters of the polar sea ice prediction model to be trained are adjusted. The multiple loss functions include a first physical constraint loss function and a second physical constraint loss function. The first physical constraint loss function measures whether the change in the predicted sea ice concentration conforms to the law of sea ice drift in the actual environment, and the second physical constraint loss function measures whether the predicted sea ice thickness satisfies the physical law of buoyancy balance.