A fast reconstruction method of high-resolution ocean thermocline depth based on random forest
Through a random forest-based method, combined with sea surface data and satellite remote sensing data, a thermospring depth reconstruction model is constructed, which solves the problem of low spatial and temporal resolution of ocean thermospring depth data in the existing technology, and achieves efficient and quasi-real-time thermospring depth reconstruction.
Patent Information
- Application Number
- CN202411264727.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The prior art is difficult to achieve rapid reconstruction of ocean thermoclip depth data with high spatiotemporal resolution, especially in the study of physical properties within the ocean, which cannot be observed in real time.
A random forest-based method is adopted to construct a thermoclip depth reconstruction model through the fusion of sea surface data and satellite remote sensing data, and train and evaluate the monthly average thermoclip-sea surface parameter matching data set to achieve rapid reconstruction of thermoclip depth.
The spatial and temporal resolution of ocean thermoclimb depth data is improved, the quasi-real-time acquisition of thermoclimb depth is achieved, and the understanding and prediction ability of ocean temperature fields is enhanced.
Smart Images

Figure CN119168095B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of physical oceanography, marine engineering, artificial intelligence, etc., and specifically relates to a high-resolution ocean thermocline depth rapid reconstruction method based on random forest, which is suitable for rapid reconstruction and quasi-real-time acquisition of ocean thermocline depth using sea surface data. Background Art
[0002] Thermocline refers to a water layer where the seawater temperature changes sharply in the vertical direction and its vertical gradient reaches a certain critical value. It is an important physical property indicator reflecting the ocean temperature field. The existence of the thermocline strengthens the upper ocean stratification. This thermal structure has an important impact on the upper ocean circulation. At the same time, the strengthening of the upper stratification and upwelling is also conducive to the uplift of the thermocline. In addition, the average thermocline changes in the central tropical Pacific (including the equatorial and non-equatorial regions) are a good precursor to the evolution of ENSO. Therefore, studying the changes in the thermocline in the tropical Pacific will help to fully understand the changing characteristics and mechanisms of ENSO. At the same time, the thermocline has an extremely important impact on the activities of underwater ships, the propagation of light and sound waves in the ocean, and marine fishing. For a long time, the study of thermocline has attracted the attention of many scientific researchers.
[0003] Since the thermocline is distributed underwater, it cannot be directly observed by ocean measurement instruments and needs to be calculated based on measured temperature profile data. The measured ocean profile data mainly comes from ship navigation, station observations, submersibles, underwater gliders and buoys, etc., and the resolution in time and space is still insufficient. Satellite altimeter remote sensing data has high temporal and spatial resolution, but it is impossible to obtain information below the ocean surface for studying the structure and changing laws of the thermocline inside the ocean.
[0004] Since remote sensing methods are limited to the ocean surface, the physical properties of the ocean interior can only be observed through ships, drifting buoys, etc., and real-time observation is not possible. In order to solve the problem of low spatiotemporal resolution of ocean thermocline data, it is necessary to combine the advantages of field observation data and satellite remote sensing products. To this end, artificial intelligence methods can be used to find the regression relationship between sea surface elements and thermocline depth, thereby realizing multi-source ocean data fusion and high-resolution rapid reconstruction of ocean thermocline depth.
[0005] Among artificial intelligence algorithms, the random forest algorithm is widely used in the inversion of ocean physical parameters and can achieve good results due to its high accuracy, strong resistance to overfitting, and simple parameter setting. Since the random forest algorithm has good accuracy and stability, it can be efficiently trained on large-scale data and can process high-dimensional data without feature selection. It performs well in handling classification and regression problems, can handle unbalanced data sets, and can output the importance of features, which is conducive to feature engineering and model interpretation. Summary of the invention
[0006] The purpose of the present invention is to overcome the shortcomings of the prior art and propose a high-resolution ocean thermocline depth rapid reconstruction method based on random forests, which uses sea surface data to achieve high-resolution ocean thermocline depth rapid reconstruction, and provide more accurate data support for ENSO phenomenon mechanism research and underwater ship activities.
[0007] The technical solution adopted by the present invention is a high-resolution ocean thermocline depth rapid reconstruction method based on random forest, which is divided into the following steps:
[0008] S1: Collect the measured temperature-salinity profile data of the sea area to be studied, pre-process the obtained data, and obtain the temperature-salinity profile data set;
[0009] S2: The temperature-salinity profile data set obtained in S1 is used to select the temperature-salinity data of the surface layer of each temperature-salinity profile as the corresponding sea surface temperature (SST) and sea surface salinity (SSS) data, and the vertical gradient method is used to calculate the thermocline depth of each temperature-salinity profile.
[0010] S3: Collect the sea level anomaly (SLA) grid data of the sea area to be studied, and use the linear interpolation method based on triangular grid to interpolate the SLA grid data to the coordinate points of the profile containing the thermocline.
[0011] S4: Construct a monthly average thermocline-sea surface parameter matching dataset, as follows:
[0012] S4.1: Perform spatial and temporal matching and standardization on the data obtained from S2 and S3, aggregate them by month, and obtain the matching dataset X in the same spatial and temporal space:
[0013]
[0014] Among them, each element in the data set X is recorded as x i,j , x i,j represents the jth element of the ith sample, and the longitude and latitude of all samples are the same; i = [1, N] represents the 1st to Nth samples, a sample is a row in the matrix X, each row contains 7 elements, j = [1, 7] represents the 1st to 7th elements, corresponding to the 7 columns of the matrix X: the first 6 columns are the 6 elements of the input data: month, longitude, latitude, SST, SSS, SLA, and the last column of label data is the thermocline depth data of the thermohaline profile, which will be used as the output of the reconstruction model; data from January to December 1993-2017 are extracted from the dataset X as the training set, and data from January to December 2018-2020 are extracted as the test set.
[0015] S4.2: Normalize the input data. Select the MinMax normalization method and normalize the elements with j = [1, 6] to the interval [0, 1]. The normalization method is as follows:
[0016]
[0017] Among them, X j represents the j-th column element of the dataset X, and Respectively represent the maximum and minimum values of the elements in the jth column of the dataset X; for the monthly data and longitude and latitude data, since the training set and the test set are consistent in monthly distribution and the sea area to be studied, their maximum and minimum values are used as the corresponding For the sea surface parameters SST, SSS, and SLA, in order to ensure their interval consistency, the maximum and minimum values of the parameters in the theoretical range are selected as the corresponding
[0018] S5: Construct a thermocline depth reconstruction model based on the random forest method, and train and evaluate the model.
[0019] S5.1 Constructing a thermocline depth reconstruction model based on random forest method
[0020] S5.1.1 uses month, longitude, latitude, SSS, SST, SLA and six other variables as inputs to the thermocline depth reconstruction model, and the thermocline depth as the model output. The number of decision trees n_estimators and the maximum split depth max_depth of the random forest method are preliminarily set based on the amount of training set data and the number of variables in the input data.
[0021] S5.1.2 Randomly extract 1 training sample from the training set, put the extracted sample back into the training set and randomly extract it again, repeat n times to obtain n training samples.
[0022] S5.1.3 Randomly select a training sample from n training samples, use the Gini index to calculate the information gain of the training sample data, and use the information gain as an indicator to screen the features of this training sample until the best feature is found as the split feature of the root node of the decision tree, split the root node into two child nodes, and finally assign the training sample to the child nodes; repeat the above steps to split the child nodes until the maximum split depth max_depth is reached, and a decision tree is constructed.
[0023] S5.1.4 Repeat step S5.1.3 to construct n_estimators decision trees to form a random forest, and obtain a thermocline depth reconstruction model based on random forest.
[0024] S5.2 Training Model
[0025] The training set constructed in S4.1 is input into the model to train the model.
[0026] S5.3 Evaluation Model
[0027] The test set constructed in S4.1 is input into the thermocline depth reconstruction model trained in S5.2. For each test instance i in the test set, the model outputs the thermocline depth result T i With thermocline depth label data P i For comparison, the relative error mre is used as the evaluation indicator to evaluate the model reconstruction effect:
[0028]
[0029] S5.4 Parameter adjustment
[0030] According to the relative error mre of S5.3, the random grid search method is used to adjust the number of decision trees n_estimators and the maximum split depth max_depth of each tree in the model, and then S5.3 is repeated to evaluate the model again. The model evaluation results before and after the adjustment are compared, and S5.2-S5.4 are repeated until the optimal model parameters are obtained, including the number of decision trees n_estimators and the maximum split depth max_depth of each tree, thereby obtaining the optimal thermocline depth reconstruction model.
[0031] S6: Thermocline depth reconstruction
[0032] The six variables of SSS, SST, SLA, month, longitude and latitude of the area to be reconstructed are input into the optimal thermocline depth reconstruction model obtained by S5 to obtain the thermocline depth reconstruction result of the area.
[0033] Furthermore, S1 collects the measured temperature and salinity profile data within the sea area and time range to be studied, and the specific process of preprocessing the obtained data is as follows:
[0034] S1.1 Eliminate data points with a depth of less than 20 meters or with less than 4 observation layers. Too few observation data will increase the error of the calculation results.
[0035] S1.2 interpolate the analysis data of each grid point to the standard layer with a sampling interval of 1 meter;
[0036] S1.3 Check the interpolated data again to see if there are any abnormal values, such as temperatures exceeding the normal range (271-306K);
[0037] S1.4 Manually remove the outliers detected in S1.3.
[0038] Furthermore, the specific process of calculating the thermocline depth in S2 is as follows:
[0039] S2.1 For the temperature profile data in the temperature-salinity dataset, calculate the temperature gradient of each depth layer Δt represents the temperature difference between the upper and lower interfaces of each layer, and Δd represents the thickness of the layer;
[0040] S2.2 According to the minimum standard value of thermocline strength specified in the national standard "Marine Survey Specification Part 7: Marine Survey Data Exchange", in deep water areas with a water depth greater than 200 meters, the minimum temperature gradient of the thermocline is 0.05℃ / m; in shallow water areas with a water depth less than 200 meters, the minimum temperature gradient of the thermocline is 0.2℃ / m. Based on the above conditions, the obtained temperature gradient value τ is judged layer by layer to determine whether it meets the thermocline standard, and the depth layer that meets the standard is defined as the thermocline.
[0041] S2.3 Analyze all the strata defined as thermoclines in a profile, merge two continuous layers into one thermocline segment, and for two discontinuous thermocline segments, when the upper boundary depth of the lower layer is less than 50 meters, if the interval between the lower boundary depth of the upper layer and the upper boundary depth of the lower layer is less than 10 meters; or when the upper boundary depth of the lower layer is greater than 50 meters, the interval between the two is less than 30 meters, then the water layer between the upper boundary of the upper layer and the lower boundary of the lower layer is taken as a new strata for gradient calculation. If the new strata still meet the thermocline judgment criteria, the new strata are defined as thermoclines; if the new strata do not meet the criteria, compare the gradients of the original two strata, and select the one with the larger gradient as the final selected thermocline. If the upper boundary depth of the finally selected thermocline is less than 50 meters, its thickness is required to be not less than 10 meters, and if the upper boundary depth is greater than 50 meters, its thickness is required to be not less than 20 meters. If the final thermocline does not meet the requirements, it is judged that there is no thermocline in the profile. The final upper boundary depth of the thermocline is recorded as the thermocline depth of the profile.
[0042] Furthermore, the measured ocean profile data in S1 are all from the product EN 4.2.2, which can be downloaded from https: / / www.metoffice.gov.uk / hadobs / en4 / download-en4-2-2.html.
[0043] Furthermore, the SLA data in S3 comes from Aviso’s satellite remote sensing data, which can be downloaded from: https: / / marine.copernicus.eu / .
[0044] The present invention has the following beneficial effects:
[0045] 1. The present invention provides an accurate and efficient thermocline depth reconstruction method. With the help of satellite remote sensing data with high temporal and spatial coverage, the thermocline depth can be quickly calculated, the temporal and spatial resolution of thermocline depth calculation can be improved, and the understanding and prediction capabilities of ocean temperature fields can be further enhanced;
[0046] 2. Based on the random forest algorithm, the present invention can effectively process large-scale data sets and achieve accurate prediction of the thermocline. Compared with traditional methods, the present invention can more effectively process complex data of sea surface elements and show good performance and adaptability in thermocline depth prediction;
[0047] 3. The present invention can be applied to different ocean observation data and remote sensing data, and has wide applicability and promotion value. At the same time, it has high prediction accuracy for different sea areas and strong portability. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 : Implementation flow chart of the method of the present invention;
[0049] Figure 2 : Comparison chart of reconstruction results of the first 600 test samples. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] Figure 1 The implementation flow chart of the method of the present invention is given. The present invention is a method for reconstructing the depth of the thermocline based on random forest, comprising the following steps:
[0052] S1: Collect the monthly average measured temperature profile data (EN4.2.2 data, https: / / www.metoffice.gov.uk / hadobs / en4 / download-en4-2-2.html) in the sea area to be studied (99°E-160°W, 60°S, 66°N) and time range (January 1993-December 2020), and preprocess the obtained data. The specific process is as follows:
[0053] S1.1 Eliminate profile data with a depth of less than 20 meters or less than 4 observation layers. Too few observation data will increase the error of the calculation results. To reduce the amount of calculation, cut off the depth of 0-1000 meters for all profile data;
[0054] S1.2 The analysis data of each grid point is interpolated to the standard layer with a sampling depth interval of 1 meter using the cubic spline interpolation method. The basic calculation method of the cubic spline interpolation method is as follows:
[0055] In each interpolation interval [x i ,x {i+1} ], the interpolation function can be expressed as:
[0056]
[0057] Where S i(x) is the interpolation function of the i-th segment, a i ,b i ,c i ,d i are the unknown coefficients. These coefficients need to satisfy the following conditions:
[0058] Interpolation condition: S i(xi) =y i , That is, the interpolation function is required to pass through the given data points.
[0059] Smoothing condition: S′ {i-1} (x i ) = S′ i (x i ), S″ {i-1} (x i )=S″ i (x i ), that is, the first-order derivative and the second-order derivative of adjacent interpolation functions at the junction are required to be equal to ensure the smoothness of the interpolation curve.
[0060] By using the least squares method, we can solve a i ,b i ,c i ,d i The value of makes the sum of square errors between the interpolation function y(x) at the known data point and the actual value minimum, and ensures the smoothness of the interpolation curve by satisfying the interpolation condition and smoothing condition. Finally, the estimated value of the data at the corresponding standard depth is obtained through the total interpolation function S;
[0061] S1.3 Check the interpolated data again to see if there are any abnormal values, such as temperatures exceeding the normal range (271-306K);
[0062] S1.4 Manually remove a small amount of obviously abnormal data.
[0063] S2: The temperature-salinity profile data set obtained in S1 is used to select the temperature-salinity data of the surface layer of each temperature-salinity profile as its corresponding sea surface temperature (SST) and sea surface salinity (SSS) data, and the thermocline depth of each temperature-salinity profile is calculated using the vertical gradient method:
[0064] S2.1 For the temperature data in the temperature-salinity dataset, calculate the temperature gradient for each depth layer Δt represents the temperature difference between the upper and lower interfaces of each layer, and Δd represents the thickness of the layer;
[0065] S2.2 According to the minimum standard value of thermocline strength stipulated in the "Ocean Survey Specifications" and the "Technical Regulations for Exclusive Economic Zone and Continental Shelf Surveys in my country", in deep waters with a water depth greater than 200 meters, the minimum temperature gradient of the thermocline is 0.05℃ / m; in shallow waters with a water depth less than 200 meters, the minimum temperature gradient of the thermocline is 0.2℃ / m. Based on the above conditions, the obtained gradient value τ is judged layer by layer to determine whether it meets the thermocline standard, and the depth layer that meets the standard is defined as the thermocline.
[0066] S2.3 Diagnostic analysis of thermoclines. Analyze all thermoclines defined in a temperature profile and merge two consecutive layers into one thermocline segment. For two discontinuous thermocline segments, when the upper boundary depth of the lower thermocline is less than 50 meters, if the interval between the lower boundary depth of the upper thermocline and the upper boundary depth of the lower thermocline is less than 10 meters; or when the upper boundary depth of the lower thermocline is greater than 50 meters, the interval between the two is less than 30 meters, then the water layer between the upper boundary of the upper thermocline and the lower boundary of the lower thermocline is taken as a new stratification for gradient calculation. If the new stratification still meets the thermocline judgment standard, the new stratification is defined as a thermocline; if the new stratification does not meet the standard, compare the gradients of the original two stratifications, and select the one with the larger gradient as the final selected thermocline. If the upper boundary depth of the final selected thermocline is less than 50 meters, its thickness is required to be no less than 10 meters; if the upper boundary depth is greater than 50 meters, its thickness is required to be no less than 20 meters. If the final thermocline does not meet the requirements, it is judged that there is no thermocline in the profile. The final upper boundary depth of the thermocline is recorded as the thermocline depth of the section.
[0067] S3: Collect the sea level anomaly (SLA) data (Aviso satellite remote sensing data, https: / / marine.copernicus.eu / ) in the sea area and time range to be studied, and use the linear interpolation method based on triangular grid to interpolate the 0.25°×0.25° SLA grid data to the coordinate point of the profile containing the thermocline calculated in S2.3, so that the SLA data corresponds to the SST and SSS data in time and space. The specific method is as follows:
[0068] Draw every three adjacent data points into a triangle to form a triangle network, and interpolate inside each triangle. Inside each triangle, the coordinates of the point to be interpolated are expressed as (x, y):
[0069] (x,y)=a×(x1,y1)+b×(x2,y2)+c×(x3,y3)
[0070] Where (x1, y1), (x2, y2), (x3, y3) are the vertex coordinates of the triangle, and a, b, c are non-negative weight coefficients that satisfy the following conditions:
[0071] a+b+c=1.
[0072] Calculate a, b, c, and calculate the approximate value f(x, y) of the data at the point (x, y) based on the data values f(x1, y1), f(x2, y2), and f(x3, y3) at the triangle vertices:
[0073] f(x,y)≈a×f(x1,y1)+b×f(x2,y2)+c×f(x3,y3)
[0074] S4: Constructing a monthly mean thermocline-sea surface parameter matching dataset
[0075] S4.1: Perform spatiotemporal matching and standardization on the data obtained from S2 and S3, and aggregate them by month to obtain a matching dataset X in the same spatiotemporal space. The dataset X is as follows:
[0076]
[0077] Among them, each sample in the data set X is denoted as x i,jIt is represented as the jth element of the ith sample, and the longitude and latitude of all samples are consistent; i = [1, N] represents the 1st to Nth samples, one sample is a row in the matrix X, each row contains 7 elements, j = [1, 7] represents the 1st to 7th elements, corresponding to the 7 columns of the matrix X. The first 6 columns are the 6 elements of the input data: month, longitude, latitude, SST, SSS, SLA, and the last column of label data is the thermocline depth data of the thermohaline profile; the data from January to December 1993 to 2017 are extracted from the data set as the training set, and the data from January to December 2018 to 2020 are extracted as the test set.
[0078] S4.2: Normalize the input data. Select the MinMax normalization method and normalize the elements with j = [1, 6] to the interval [0, 1]. The specific normalization method is as follows:
[0079]
[0080] Among them, for the monthly data and longitude and latitude data, since the training set and the test set have the same monthly distribution and research sea area, their maximum and minimum values are used as the corresponding For the sea surface parameters SST, SSS, and SLA, in order to ensure their interval consistency, the maximum and minimum values of the theoretical range of the parameter are selected as the corresponding Get the maximum and minimum values of each variable, where the month is (1,12), the longitude is (99°,200°), the latitude is (-60°,60°), the SST is (-2,38), the SSS is (0,40), and the SAL is (-1.5,1)
[0081] S5: Construct a thermocline depth reconstruction model based on the random forest method, and train and evaluate the model.
[0082] S5.1 Constructing a thermocline depth reconstruction model
[0083] S5.1.1 takes the six sea surface variables of SSS, SST, SLA, month, longitude and latitude as the input of the random forest method and the thermocline depth as the output, and initially sets the number of decision trees n_estimators and the maximum split depth max_depth of the random forest method to [200,50].
[0084] S5.1.2 Randomly extract 1 training sample from the training set, put the selected sample back into the training set and randomly extract it again, repeat n times to obtain n training samples.
[0085] S5.1.3 Randomly select a training sample from n training samples, use the Gini index to calculate the information gain of the training sample data, and select the features of this training sample based on the information gain. Find the best feature as the split feature of the root node of the decision tree, and split the root node into two child nodes. Assign the training samples to the child nodes, and repeat the above steps to split the child nodes until the maximum split depth is reached, and a decision tree is constructed.
[0086] S5.1.4 Repeat step S5.1.3 to construct multiple decision trees within the limit of the number of decision trees n_estimators to form a thermocline depth reconstruction model.
[0087] S5.2 Training Model
[0088] The data from 1993 to 2017 were used as the training set to input the thermocline depth reconstruction model constructed in 5.1 to train the model.
[0089] S5.3 Evaluation Model
[0090] The data from 2018 to 2020 are used as the test set to input the random forest model trained in S5.2. For each test instance i in the test set, the model outputs the thermocline depth result T i With thermocline depth label data P i For comparison, the relative error mre is used as the evaluation indicator to evaluate the model reconstruction effect:
[0091]
[0092] S5.4 Parameter adjustment
[0093] Set the number of decision trees n_estimators to [200:1:500], the maximum split depth of each tree max_depth to [50:1:100], and based on the model evaluation results of S5.3, repeat S5.3 for different combinations of the number of decision trees n_estimators and the maximum split depth of each tree max_depth in the model to evaluate the model again using the grid search method. . Through the grid search method, the values of different combinations of the number of decision trees n_estimators and the maximum split depth of each tree are divided into discrete grids to traverse all possible combinations. For all combinations, the random forest model is trained on the training set, and finally the optimal hyperparameter combination max_depth = 67, n_estimators = 217 is obtained.
[0094] S6: Thermocline depth reconstruction
[0095] The six variables of SSS, SST, SLA, month, longitude and latitude of the area to be reconstructed are input into the optimal thermocline depth reconstruction model obtained by S5 to obtain the thermocline depth reconstruction result of the area.
[0096] Figure 2 The comparison between the reconstruction results and the actual values of the first 600 samples in the test set is given. The reconstruction results are consistent with the actual values, the model fitting coefficient is 0.86, and the average relative error is 0.2.
[0097] The present invention has achieved obvious implementation effects in typical embodiments. The model for reconstructing the thermocline depth based on random forest is accurate, efficient, widely applicable, has high temporal and spatial resolution, and is suitable for real-time calculation of the thermocline depth using sea surface data.
Claims
1. A high-resolution ocean thermocline depth rapid reconstruction method based on random forest, characterized in that: The method is divided into the following steps: S1: Collect the measured temperature-salinity profile data of the sea area to be studied, pre-process the obtained data, and obtain the temperature-salinity profile data set; S2: The temperature-salinity profile data set obtained in S1 is used to select the temperature-salinity data of the surface layer of each temperature-salinity profile as the corresponding sea surface temperature SST and sea surface salinity SSS data, and the vertical gradient method is used to calculate the thermocline depth of each temperature-salinity profile; S3: Collect the SLA grid data of sea surface height anomaly in the sea area to be studied, and use the linear interpolation method based on triangular grid to interpolate the SLA grid data to the coordinate points of the profile containing the thermocline; S4: Construct a monthly average thermocline-sea surface parameter matching dataset, as follows: S4.1: Perform spatial and temporal matching and standardization on the data obtained from S2 and S3, aggregate them by month, and obtain the matching dataset X in the same spatial and temporal space: Among them, each element in the data set X is recorded as x i,j , x i,j represents the jth element of the ith sample, and the longitude and latitude of all samples are consistent; i = [1, N] represents the 1st to Nth samples, one sample is a row in the matrix X, each row contains 7 elements, j = [1, 7] represents the 1st to 7th elements, corresponding to the 7 columns of the matrix X: the first 6 columns are the 6 elements of the input data: month, longitude, latitude, SST, SSS, SLA, and the last column of label data is the thermocline depth data of the thermohaline profile, which will be used as the output of the reconstruction model; extract the data from January to December 1993-2017 from the dataset X as the training set, and the data from January to December 2018-2020 as the test set; S4.2: Normalize the input data. Select the MinMax normalization method and normalize the elements with j = [1, 6] to the interval [0, 1]. The normalization method is as follows: Among them, X j represents the j-th column element of the dataset X, and Represent the maximum and minimum values of the elements in the jth column of the dataset X respectively; for the monthly data and longitude and latitude data, since the monthly distribution and the sea area to be studied are consistent between the training set and the test set, their maximum and minimum values are taken as the corresponding For the sea surface parameters SST, SSS, and SLA, in order to ensure their interval consistency, the maximum and minimum values of the parameters in the theoretical range are selected as the corresponding S5: Construct a thermocline depth reconstruction model based on the random forest method, and train and evaluate the model: S5.1 Constructing a thermocline depth reconstruction model based on random forest method S5.1.1 Use month, longitude, latitude, SSS, SST, SLA as the input of the thermocline depth reconstruction model, and the thermocline depth as the model output. Preliminary settings are made for the number of decision trees n_estimators and the maximum split depth max_depth of the random forest method based on the amount of training set data and the number of variables in the input data. S5.1.2 Randomly extract a training sample from the training set, put the extracted sample back into the training set and randomly extract it again, repeat n times to obtain n training samples; S5.1.3 Randomly select a training sample from n training samples, use the Gini index to calculate the information gain of the training sample data, and use the information gain as an indicator to screen the features of this training sample until the best feature is found as the split feature of the root node of the decision tree, split the root node into two child nodes, and finally assign the training sample to the child nodes; repeat the above steps to split the child nodes until the maximum split depth max_depth is reached, and a decision tree is constructed; S5.1.4 Repeat step S5.1.3 to construct n_estimators decision trees to form a random forest, and obtain a thermocline depth reconstruction model based on random forest; S5.2 Training Model Input the training set constructed in S4.1 into the model to train the model; S5.3 Evaluation Model The test set constructed in S4.1 is input into the thermocline depth reconstruction model trained in S5.
2. For each test instance i in the test set, the model outputs the thermocline depth result T i With thermocline depth label data P i For comparison, the relative error mre is used as the evaluation indicator to evaluate the model reconstruction effect: S5.4 Parameter adjustment According to the relative error mre of S5.3, the number of decision trees n_estimators and the maximum splitting depth max_depth of each tree in the model are adjusted by random grid search method, and then S5.3 is repeated to evaluate the model again. The model evaluation results before and after the adjustment are compared, and S5.2-S5.4 are repeated until the best model parameters are obtained, including the number of decision trees n_estimators and the maximum splitting depth max_depth of each tree, so as to obtain the optimal thermocline depth reconstruction model; S6: Thermocline depth reconstruction The six variables of SSS, SST, SLA, month, longitude and latitude of the area to be reconstructed are input into the optimal thermocline depth reconstruction model obtained by S5 to obtain the thermocline depth reconstruction result of the area.
2. The high-resolution ocean thermocline depth rapid reconstruction method based on random forest according to claim 1 is characterized by: In S1, the measured temperature and salinity profile data within the sea area and time range to be studied are collected, and the specific process of preprocessing the obtained data is as follows: S1.1 Eliminate data points with a depth of less than 20 meters or with less than 4 observation layers; S1.2 interpolate the analysis data of each grid point to the standard layer with a sampling interval of 1 meter; S1.3 Check the interpolated data again to see if there are any abnormal values; S1.4 Manually remove the outliers detected in S1.
3.
3. The high-resolution ocean thermocline depth rapid reconstruction method based on random forest according to claim 1 is characterized in that: The specific process of calculating the thermocline depth in S2 is as follows: S2.1 For the temperature profile data in the temperature-salinity dataset, calculate the temperature gradient of each depth layer Δt represents the temperature difference between the upper and lower interfaces of each layer, and Δd represents the thickness of the layer; S2.2 According to the minimum standard value of thermocline strength specified in the national standard "Marine Survey Specification Part 7: Marine Survey Data Exchange", in deep water areas with a water depth greater than 200 meters, the minimum temperature gradient of the thermocline is 0.05℃ / m; in shallow water areas with a water depth less than 200 meters, the minimum temperature gradient of the thermocline is 0.2℃ / m; according to the above conditions, the obtained temperature gradient value τ is judged layer by layer whether it meets the thermocline standard, and the depth layer that meets the standard is defined as the thermocline; S2.3 Analyze all the strata defined as thermoclines in a profile, merge two consecutive layers into one thermocline segment, and for two discontinuous thermocline segments, when the upper boundary depth of the lower layer is less than 50 meters, if the interval between the lower boundary depth of the upper layer and the upper boundary depth of the lower layer is less than 10 meters; or when the upper boundary depth of the lower layer is greater than 50 meters, and the interval between the two is less than 30 meters, then the water layer between the upper boundary of the upper layer and the lower boundary of the lower layer is taken as a new strata for gradient calculation; if the new strata still meet the thermocline judgment criteria, then the new strata are defined as thermoclines; if the new strata do not meet the criteria, compare the gradients of the original two strata, and select the one with the larger gradient as the final selected thermocline; if the upper boundary depth of the final selected thermocline is less than 50 meters, its thickness is required to be not less than 10 meters, and if the upper boundary depth is greater than 50 meters, its thickness is required to be not less than 20 meters; if the final thermocline does not meet the requirements, then it is judged that there is no thermocline in the profile; the final upper boundary depth of the thermocline is recorded as the thermocline depth of the profile.
Citation Information
Patent Citations
High-resolution ocean water temperature calculation method based on hierarchical clustering and random forest
CN111242206A
Ocean thermohaline structure inversion method based on artificial intelligence
CN116822381A