Random Forest correction-based TPXO9-atlas-v5 global tide model forecasting method

By correcting the TPXO9-atlas-v5 global tidal model based on Random Forest, the problem of forecast error in traditional models in complex marine environments and nearshore areas is solved, and higher tidal forecast accuracy and computing efficiency are achieved.

CN120105099APending Publication Date: 2025-06-06BEIJING ZHONGAN INTELLIGENT INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510181893.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional tidal forecasting models have forecast errors in complex marine environments and nearshore areas, which are difficult to accurately reflect the influence of factors such as terrain and river runoff.

Method used

The TPXO9-atlas-v5 global tide model was modified by using a method based on Random Forest. By screening characteristic factors affecting tides, such as temperature, salinity and wind speed, a tide error regression model was constructed to correct the model forecast results.

Benefits of technology

It significantly improves the accuracy of nearshore tide forecasts, reduces the amount of calculation and time, provides more reliable tide data, and reduces the risks caused by inaccurate tide forecasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105099A_ABST
    Figure CN120105099A_ABST
Patent Text Reader

Abstract

The invention discloses a TPXO9-atlas-v5 global tide model forecasting method based on Random Forest correction, and relates to the technical field of tide forecasting, characteristic factors influencing local tide are screened out through historical data analysis in advance, and the characteristic factors comprise temperature T, salinity S and wind speed V; acquiring an actually measured data set of the corresponding characteristic factors influencing the local tide acquired at the same time and at the same position; arranging the measured data set to form a characteristic factor matrix, and inputting the characteristic factor matrix as a tide error regression model generated by training; and running the tide error regression model and carrying out calculation to obtain a tide error prediction result Rerror. According to the method, the output result of the TPXO9-atlas-v5 model is directly corrected according to the Random Forest method, a complex harmonic analysis process is avoided, the calculation amount is remarkably reduced, the calculation efficiency is improved, the time for obtaining an accurate tide forecasting result is greatly shortened, and timely and effective data support can be provided for related ocean activities more quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tidal forecasting, and in particular to a TPXO9-atlas-v5 global tidal model forecasting method based on Random Forest correction. Background Art

[0002] As a complex and long-standing natural phenomenon in a wide area, accurate tide forecasting plays a key role in many fields. In terms of marine transportation, accurate tide forecasting helps ships to rationally plan routes and sailing times, avoid dangerous currents and shoal areas, thereby significantly improving transportation efficiency and ensuring navigation safety. Taking large container ships as an example, if the best time to enter and leave the port can be selected based on accurate tidal information, waiting time can be effectively reduced and operating costs can be reduced. For offshore engineering, the accuracy of tide forecasting is directly related to the safety and cost of engineering construction and operation and maintenance. For example, when constructing large projects such as cross-sea bridges and offshore wind farms, the construction party needs to reasonably arrange the construction progress and operation methods according to the changes in tides, otherwise the water flow and water level changes caused by the tides may increase the difficulty of construction and even cause safety accidents. In the field of coastal disaster prevention and mitigation, timely and accurate tide forecasting can provide an important basis for early warning and prevention of disasters such as storm surges and tsunamis, so that residents and facilities in coastal areas can be more effectively protected and the losses caused by disasters can be greatly reduced.

[0003] Traditional tidal forecasting methods mostly rely on numerical models, which are mainly based on the principles of tidal dynamics, such as the equilibrium tide theory and the dynamic tide theory. However, the ocean environment is extremely complex, covering factors such as changing topography, different seawater temperature and salinity distribution, and complex ocean circulation; the ocean dynamic process is also varied, all of which poses a huge challenge to the accurate prediction of traditional numerical models. In practical applications, when faced with complex ocean environments and ever-changing ocean dynamics, traditional models often have forecast deviations, resulting in reduced accuracy.

[0004] At the same time, as one of the most advanced tidal models in the world, TPXO9-atlas-v5 has greatly improved the accuracy of tidal forecasts by fusing satellite altimeter data and tidal instrument observation data through data assimilation technology. However, in nearshore and complex terrain areas, the model still has large forecast errors. This is mainly because in the process of model construction, the setting of initial conditions is difficult to fully fit the actual situation, the processing of boundary conditions is not perfect, and the simulation ability of small-scale physical processes is insufficient. For example, the nearshore area is strongly affected by multiple factors such as land topography and estuary runoff, and these complex small-scale processes cannot be fully and accurately reflected in the model, which leads to a large deviation between the forecast results and the actual tidal conditions.

[0005] With the rapid development of machine learning and big data technology, data-driven model correction methods have gradually emerged and become a new direction for improving the accuracy of tidal forecasting. These methods can fully tap the potential information in massive data, capture the complex laws of tidal changes through learning and analyzing a large amount of historical data, and provide new ways and possibilities for optimizing tidal forecasting models.

[0006] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention

[0007] In view of the problems in the related art, the present invention proposes a TPXO9-atlas-v5 global tidal model forecasting method based on Random Forest correction to overcome the above-mentioned technical problems existing in the existing related technology.

[0008] The technical solution of the present invention is achieved in this way:

[0009] A TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction includes the following steps:

[0010] The characteristic factors that affect the local tides are screened out in advance through historical data analysis, where the characteristic factors include: temperature T, salinity S and wind speed V;

[0011] Obtain the measured data set collected at the same time and location for the corresponding characteristic factors affecting the local tide;

[0012] The measured data set is sorted to form a characteristic factor matrix As input to the tidal error regression model generated by training;

[0013] Run the tidal error regression model and perform calculations to obtain the tidal error prediction result R error ;

[0014] According to the measured data set, the time and location information contained in the measured data set is obtained;

[0015] Run the TPXO9-atlas-v5 model to obtain the corresponding local tidal model forecast results R model ;

[0016] The corresponding tidal error prediction result R error Compared with the local tidal model forecast results R model Add together to get the corresponding corrected tidal forecast value R predict , expressed as:

[0017] R predict =R model +R error .

[0018] Furthermore, the measured data set includes: a temperature data set T set , salinity dataset S set and wind speed dataset V set .

[0019] Furthermore, the tidal error regression model generated by the training comprises the following steps:

[0020] Obtain data on factors affecting nearshore tides and measured tidal data in advance, use the time and location information included in the measured tidal data as input, use the TPXO9-atlas-v5 global tidal model to forecast local nearshore tides, and obtain model forecast data for the corresponding time and location;

[0021] The acquired influencing factor data, measured tide data and model prediction data are used to construct a data set, which includes: constructing the influencing factor data into the Features of the new data set; at the same time, calculating the difference between the measured tide data and the model prediction data to obtain the tidal error data set, and using the tidal error data set as the label of the constructed new data set;

[0022] The Label of the new data set is used for model training, and the RF integrated training technology is used to generate a tidal error regression model.

[0023] Further, the method of using the tidal error dataset as the Label of the constructed new dataset includes: performing a dataset reconstruction operation on the Label of the new dataset, including the following steps:

[0024] The constructed data set was standardized according to the Z-Score technique;

[0025] By using the Polynomial Features technology to expand the features in the standardized data set, a feature data set is obtained;

[0026] The Random Forest technology in machine learning is used to construct a feature screening model to perform feature screening on the above-formed data set, which at least includes screening out the characteristics of the impact of local nearshore tides, and using the impact feature data set as the Features of the reconstructed new data set, and the corresponding tidal error data set as the Label of the reconstructed new data set.

[0027] Furthermore, the RF integrated training technology is used to generate the tidal error regression model, including: an integrated algorithm based on a decision tree, and Bootstrap aggregating is used in prediction, the average value of the results of multiple tidal error decision trees is taken as the prediction output of the final integrated tidal error regression model, and when constructing the nodes of each tidal error decision tree, a part of the features is randomly selected to find the best segmentation point.

[0028] Furthermore, the parameters of the tidal error regression model include: n_estimators=100, max_depth=10, min_samples_split=2, min_samples_leaf=1 and random_state=42.

[0029] Furthermore, it also includes: inputting Features in the test set, obtaining the test results, and calculating the mean square error MSE between the test results and the tidal error Label in the original test set, expressed as:

[0030]

[0031] Calculate the mean absolute error MAE, expressed as:

[0032]

[0033] Among them, n represents the number of times the feature is input into the test set, F(feature i ) represents the regression model test result corresponding to the i-th feature in the input test set, L i Indicates the Label value of the tidal error in the test set corresponding to the i-th feature.

[0034] Beneficial effects of the present invention:

[0035] The present invention directly corrects the output results of the TPXO9-atlas-v5 model according to the Random Forest method, avoiding the complex reconciliation analysis process, significantly reducing the amount of calculation, improving calculation efficiency, greatly shortening the time to obtain accurate tidal forecast results, and being able to provide timely and effective data support for ocean-related activities more quickly. At the same time, it fully considers the various factors that affect the tides in the nearshore area, comprehensively screens out the key characteristic factors, and deeply combines the high-precision TPXO-atlas-v5 forecast results. In addition, by constructing an accurate tidal error regression model, the forecast error of TPXO9-atlas-v5 in the nearshore area is effectively corrected, thereby significantly improving the accuracy of the nearshore tidal forecast data. In the nearshore complex terrain area, this advantage is particularly prominent, providing more reliable tidal data for marine transportation, offshore engineering construction and other activities, and reducing the risks caused by inaccurate tidal forecasts.

[0036] In addition, the Random Forest technology is used to explore the potential factors affecting the tides, making the revised forecast model more adaptable to different nearshore scenarios. Whether it is a bay with complex terrain or an estuary area that is greatly affected by river runoff, it can stably output accurate forecast results, effectively improving the robustness and scene migration capabilities of the model in different environments, expanding the application scope of the tidal forecast model, and can be quickly deployed to the actual ocean monitoring and forecasting system, facilitating integration with existing software and hardware facilities, reducing application costs and development cycles, and providing strong support for the informatization construction of ocean-related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Figure 1 The figure is a flow chart of a TPXO9-atlas-v5 global tidal model forecasting method based on Random Forest correction according to an embodiment of the present invention. Figure 1 ;

[0039] Figure 2 The figure is a flow chart of a TPXO9-atlas-v5 global tidal model forecasting method based on Random Forest correction according to an embodiment of the present invention. Figure 2 . DETAILED DESCRIPTION

[0040] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.

[0041] According to an embodiment of the present invention, a TPXO9-atlas-v5 global tidal model forecasting method based on Random Forest correction is provided.

[0042] like Figure 1-Figure 2 As shown, the TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to an embodiment of the present invention includes the following steps:

[0043] The characteristic factors that affect the local tides are screened out in advance through historical data analysis, where the characteristic factors include: temperature T, salinity S and wind speed V;

[0044] Obtain the measured data sets collected at the same time and location for the corresponding characteristic factors affecting the local tide, including: temperature data set T set (T 1 , T 2 , T 3 , ...), salinity dataset S set (S 1 , S 2 , S 3 , ...) and wind speed dataset V set (V 1 , V 2 , V 3 , …);

[0045] The measured data set is sorted to form a characteristic factor matrix As input to the tidal error regression model generated by training;

[0046] Run the tidal error regression model and perform calculations to obtain the tidal error prediction result R error ;

[0047] According to the measured data set, the time and location information contained in the measured data set is obtained;

[0048] Specifically, when applied, such as location information: Dongshan Port (Latitude: 23.45°, Longitude: 117.31°), time information: December 19, 2024, as input information of the TPXO9-atlas-v5 model.

[0049] Run the TPXO9-atlas-v5 model to obtain the corresponding local tidal model forecast results R model ;

[0050] The corresponding tidal error prediction result R error Compared with the local tidal model forecast results R model Add together to get the corresponding corrected tidal forecast value R predict , expressed as:

[0051] R predict =R model +R error .

[0052] This technical solution, for the tidal error regression model generated by the above training, includes the following steps:

[0053] Determine in advance the factors affecting the nearshore tides, including temperature T, salinity S, wind speed V, etc., and obtain the data of these factors affecting the nearshore tides; at the same time, obtain the measured tidal data, and use the TPXO9-atlas-v5 global tidal model to forecast the local nearshore tides based on the time and location information attached to the measured tidal data as input, so as to obtain the model forecast data for the corresponding time and location.

[0054] The acquired influencing factor data, measured tide data and model prediction data are used to construct a data set, which includes: constructing the influencing factor data into the Features of the new data set; at the same time, calculating the difference between the measured tide data and the model prediction data to obtain the tidal error data set, and using the tidal error data set as the label of the constructed new data set;

[0055] Specifically, when the technical solution is applied, it also includes a label of a new data set to perform a data set reconstruction operation, including the following steps:

[0056] First, the constructed data set is standardized using the Z-Score technique; second, the PolynomialFeatures technique is used to expand the features in the standardized data set, for example: 2 , T 2 *P, V*S*P, etc., thus obtaining a relatively rich feature data set. On this basis, the Random Forest technology in machine learning is used to build a feature screening model to perform feature screening on the above-formed data set. The features that are more important for the local nearshore tidal impact are screened out. Specifically, the top 30% of important features can be taken, and this important feature data set is used as the Features of the reconstructed new data set, and the corresponding tidal error data set is used as the Label of the reconstructed new data set.

[0057] The Label of the new data set is used for model training, which includes the following steps:

[0058] The reconstructed data set is split, and 80% of the data in the data set is randomly selected as the training set of the model, and the remaining 20% ​​of the data is used as the test set for model performance testing.

[0059] In the model training process, RF integrated training technology is used to generate a tidal error regression model.

[0060] Specifically, RF ensemble training technology is an ensemble algorithm based on decision trees. Each decision tree (model) is trained independently, and Bootstrap aggregating technology is used in prediction. The average of the results of multiple tidal error decision trees is taken as the prediction output of the final integrated tidal error regression model. Each decision tree is trained on different samples to obtain different tidal error decision trees. When constructing the nodes of each tidal error decision tree, a part of the features is randomly selected to find the best split point.

[0061] In addition, when training the model, the corresponding parameters are set as follows: n_estimators = 100 (the number of decision trees constructed), max_depth = 10 (the maximum depth of each tree), min_samples_split = 2 (the minimum number of samples required for node splitting), min_samples_leaf = 1 (the minimum number of samples required for leaf nodes), and random_state = 42 (used to control the seed of the random number generator).

[0062] Among them, during the model performance test, the test set is used to test the above integrated model, and the Features in the corresponding test set are input to obtain the test results. The mean square error (MSE) and mean absolute error (MAE) between the result and the tidal error Label in the original test set are calculated. The corresponding calculation formula is as follows:

[0063]

[0064] In the above formula, n represents the number of times the feature is input into the test set, F(featurei) represents the regression model test result corresponding to the i-th feature in the test set, and L iIndicates the Label value of the tidal error in the test set corresponding to the i-th feature. The smaller the calculated MSE, the better the tidal error regression model. The calculated mean absolute error can reflect the average deviation between the output value of the regression model and the actual true value. The above two indicators are combined to evaluate the prediction effect of the tidal error regression model. Repeat the above operation and adjust different parameters within a reasonable range until the calculated MSE and MAE loss of the model reach the optimal state in the training set test set, that is, the value is the smallest.

[0065] In summary, with the help of the above technical solution of the present invention, the following effects can be achieved:

[0066] The present invention directly corrects the output results of the TPXO9-atlas-v5 model according to the Random Forest method, avoiding the complex reconciliation analysis process, significantly reducing the amount of calculation, improving calculation efficiency, greatly shortening the time to obtain accurate tidal forecast results, and being able to provide timely and effective data support for ocean-related activities more quickly. At the same time, it fully considers the various factors that affect the tides in the nearshore area, comprehensively screens out the key characteristic factors, and deeply combines the high-precision TPXO-atlas-v5 forecast results. In addition, by constructing an accurate tidal error regression model, the forecast error of TPXO9-atlas-v5 in the nearshore area is effectively corrected, thereby significantly improving the accuracy of the nearshore tidal forecast data. In the nearshore complex terrain area, this advantage is particularly prominent, providing more reliable tidal data for marine transportation, offshore engineering construction and other activities, and reducing the risks caused by inaccurate tidal forecasts.

[0067] In addition, the Random Forest technology is used to explore the potential factors affecting the tides, making the revised forecast model more adaptable to different nearshore scenarios. Whether it is a bay with complex terrain or an estuary area that is greatly affected by river runoff, it can stably output accurate forecast results, effectively improving the robustness and scene migration capabilities of the model in different environments, expanding the application scope of the tidal forecast model, and can be quickly deployed to the actual ocean monitoring and forecasting system, facilitating integration with existing software and hardware facilities, reducing application costs and development cycles, and providing strong support for the informatization construction of ocean-related fields.

[0068] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. After considering the disclosure of the specification and the examples, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and the examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the claims.

[0069] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction, characterized in that: The following steps are involved: The characteristic factors that affect the local tides are screened out in advance through historical data analysis, where the characteristic factors include: temperature T, salinity S and wind speed V; Obtain the measured data set collected at the same time and location for the corresponding characteristic factors affecting the local tide; The measured data set is sorted to form a characteristic factor matrix As input to the tidal error regression model generated by training; Run the tidal error regression model and perform calculations to obtain the tidal error prediction result R error ; According to the measured data set, the time and location information contained in the measured data set is obtained; Run the TPXO9-atlas-v5 model to obtain the corresponding local tidal model forecast results R model ; The corresponding tidal error prediction result R error Compared with the local tidal model forecast results R model Add together to get the corresponding corrected tidal forecast value R predict , expressed as: R predict =R model +R error 。 2. The TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to claim 1, characterized in that: The measured data set includes: temperature data set T set , salinity dataset S set and wind speed dataset V set .

3. The TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to claim 1, characterized in that: The tidal error regression model generated by the training comprises the following steps: Obtain data on factors affecting nearshore tides and measured tidal data in advance, use the time and location information included in the measured tidal data as input, use the TPXO9-atlas-v5 global tidal model to forecast local nearshore tides, and obtain model forecast data for the corresponding time and location; The acquired influencing factor data, measured tide data and model prediction data are used to construct a data set, which includes: constructing the influencing factor data into the Features of the new data set; at the same time, calculating the difference between the measured tide data and the model prediction data to obtain the tidal error data set, and using the tidal error data set as the label of the constructed new data set; The Label of the new data set is used for model training, and the RF integrated training technology is used to generate a tidal error regression model.

4. The TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to claim 3 is characterized in that: The step of using the tidal error dataset as the Label of the constructed new dataset includes: performing a dataset reconstruction operation on the Label of the new dataset, including the following steps: The constructed data set was standardized according to the Z-Score technique; By using the PolynomialFeatures technology to expand the features in the standardized data set, a feature data set is obtained; The Random Forest technology in machine learning is used to construct a feature screening model to perform feature screening on the above-formed data set, which at least includes screening out the characteristics of the impact of local nearshore tides, and using the impact feature data set as the Features of the reconstructed new data set, and the corresponding tidal error data set as the Label of the reconstructed new data set.

5. The TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to claim 3, characterized in that: The RF integrated training technology is used to generate the tidal error regression model, including: an integrated algorithm based on a decision tree, and Bootstrap aggregating is used in prediction, the average value of the results of multiple tidal error decision trees is taken as the prediction output of the final integrated tidal error regression model, and when constructing the nodes of each tidal error decision tree, a part of the features is randomly selected to find the best segmentation point.

6. The TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to claim 5, characterized in that: The parameters of the tidal error regression model include: n_estimators=100, max_depth=10, min_samples_split=2, mmin_samples_leaf=1 and random_state=42.

7. The TPXO9-atlas-v5 global tidal model prediction method based on Random Forest correction according to claim 6, characterized in that: Also includes: Input the features in the test set, get the test results, and calculate the mean square error (MSE) between the test results and the tidal error Label in the original test set, expressed as: Calculate the mean absolute error MAE, expressed as: Among them, n represents the number of times the feature is input into the test set, F(feature i ) represents the regression model test result corresponding to the i-th feature in the input test set, L i Indicates the Label value of the tidal error in the test set corresponding to the i-th feature.