Tidal station tide level correction applicability evaluation method and system

By constructing an evaluation method for the applicability of tidal level correction at tidal stations, calculating eight performance characteristic indices, and combining physical mechanism constraints, the problem of evaluating the applicability of machine learning correction at tidal stations was solved, and objective and reliable tidal level prediction and feature acquisition were achieved.

CN121901680APending Publication Date: 2026-04-21CTG JIANGSU ENERGY INVESTMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies lack a unified criterion to quantitatively determine whether tidal stations need to use machine learning to correct tidal levels, resulting in a lack of a systematic evaluation system for different stations and conditions, and making it impossible to effectively avoid overfitting and physical distortion.

Method used

This paper provides a method for evaluating the applicability of tidal level correction at tidal stations. By constructing physical error sequences and machine learning error sequences, eight performance characteristic indices are calculated. Combined with the constraints of tidal physical mechanisms, the applicability evaluation conclusions are output and a results report is generated.

Benefits of technology

A standardized, multi-indicator evaluation system was established to objectively determine whether machine learning correction is recommended, providing a reliable basis to improve the accuracy of tide prediction and the reliability of feature acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901680A_ABST
    Figure CN121901680A_ABST
Patent Text Reader

Abstract

The invention discloses a tide level correction applicability evaluation method and system for a tide station, and belongs to the technical field of machine learning processing under tide harmonic analysis, and the method comprises the steps: carrying out the unified time reference alignment and preprocessing of an actually measured tide level time sequence; respectively constructing a physical error sequence and a machine learning error sequence; obtaining a physical side performance characteristic index and a machine learning side performance characteristic index; according to the tide level correction applicability evaluation method and system of the tide station, eight novel performance indexes are respectively calculated by constructing an error sequence of an actually measured tide level, a physical tide level and a machine learning tide level, so that a quantitative judgment result whether machine learning correction is recommended or not is given under the constraint of a physical mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine learning processing technology under tidal harmonic analysis, specifically relating to a method and system for evaluating the applicability of tidal level correction at tidal stations. Background Technology

[0002] Tides, as a natural phenomenon of ocean currents, are traditionally analyzed using Classical Harmonic Analysis (CHA). This involves solving for the amplitude and phase of tidal constituents using least-squares regression, given a specific astronomical frequency. This method has long been widely used for tidal level forecasting in coastal and offshore areas. In most near-stationary tidal environments, the vast majority of water level variance can be explained by a limited number of tidal constituents. However, in areas with estuarine tides, internal tides, strong storm surges, or significant non-tidal forcing effects such as runoff and sea ice, tidal processes exhibit markedly non-stationary characteristics. The stationarity assumption of CHA deviates systematically from the actual dynamic processes, leading to significantly increased backcalar and forecasting errors, making it difficult to accurately characterize the time-varying tidal range, phase, and mean water level.

[0003] For non-steady tides, various improved physical methods exist, such as wavelet or short-time harmonic analysis based on complex demodulation and short-time windows to obtain the time-varying amplitude and phase of tidal constituents. Tools developed based on these ideas, such as NS_TIDE, EHA, and S_TIDE, incorporate non-stationary factors like runoff, open ocean tidal range, or independent point interpolation into the harmonic framework, improving the fitting ability for processes like estuarine tides and internal tides to some extent. However, these methods often rely on specific dynamic factors (such as high-quality runoff data) or have limitations in independent point selection and generalization to complex environments, and are not universally applicable to all stations and conditions. Meanwhile, machine learning methods, such as machine learning and deep learning, are rapidly being applied in hydrological and oceanographic forecasting. A common approach is to use the physical model output and temporal features as input, employing algorithms such as random forests, gradient boosting, AdaBoost, and LSTM to fit the residuals between measured and physical results, thereby improving the accuracy of short-term or medium- to long-term tide level predictions. These methods can characterize complex nonlinear relationships, but they usually rely on long, high-quality training sequences and lack clear physical meaning. Once the training samples are insufficient, the observation error is large, or the model is extrapolated to an out-of-sample scenario, the model may exhibit a mathematical illusion that the error is reduced but in fact destroys the physical structure.

[0004] Currently, when introducing machine learning to correct tidal levels, most methods only use a few global indicators such as RMSE, MAE, and the coefficient of determination to simply compare the "physical model" and the "physical + machine learning model," lacking a unified framework for systematically evaluating the reliability of the physical model and the true improvement of machine learning from multiple dimensions. Under different stations and observation conditions, there is currently no universally applicable criterion to answer under what conditions machine learning should be used to correct physical tidal levels, and under what conditions it is not advisable to introduce machine learning to avoid overfitting and physical distortion. Current technical solutions typically only use a few indicators such as RMSE, MAE, and R² to simply compare the improvement effect, without establishing a systematic and generalizable evaluation system for the applicability of machine learning models for tidal stations, and lacking a unified criterion to quantitatively determine whether machine learning correction is necessary. Therefore, it is necessary to develop a new method and system for evaluating the applicability of tidal level correction at tidal stations to solve the existing problems. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for evaluating the applicability of tidal level correction at tidal stations, in order to solve the problem of lacking a unified criterion for quantitatively determining whether it is necessary to use machine learning correction.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for evaluating the applicability of tidal level correction at tidal stations, comprising: Tidal data input: Input the measured tidal level time series of the target tidal station, and perform unified time reference alignment and preprocessing on the measured tidal levels; For index calculation, physical error sequences and machine learning error sequences are constructed on the validation segment, and four performance characteristic indices are calculated on both the physical and machine learning sides. The four performance characteristic indices on the physical side are, in order: Physical Fit Adequacy Index (PFA), Tidal Range Coverage Index (TRC), Residual Normality Index (RNI), and Physical Stability Index (PSI). The four performance characteristic indices on the machine learning side are, in order: Relative Gain Index (RGI), Consistency Improvement Index (CII), Significance Credibility Index (SCI), and Comprehensive Improvement Reliability Index (CRI). The mechanism and performance characteristic indicators are combined to calculate and determine whether the target station should adopt machine learning tidal level correction. The eight performance characteristic indices are input into the preset judgment rules and combined with the tidal physical mechanism constraints to make a comprehensive judgment on whether the target station should adopt machine learning tidal level correction. The applicability evaluation conclusion is output and a result report is generated and output. The result report includes at least the values ​​of the eight indices, the threshold comparison results, and the final judgment level.

[0007] Preferably, the input data may include training segment data and validation segment data. The calculation scope of this invention is the validation segment content, and does not involve the training segment; the training segment data is regarded as machine learning training data and is not included in the validation calculation scope.

[0008] Preferably, the Physical Fit Adequacy Index (PFA) is constructed based on the Nash–Sutcliffe efficiency coefficient, as shown in the following formula: ; In the above formula To verify the tidal observation data at time t, To verify the mean of the tidal observation data, For the physical harmonic analysis of tidal level data at time t, verification segment is performed. This means restricting the range of the function value within the parentheses to the interval [0,1].

[0009] Preferably, the tidal coverage index (TRC) is constructed by the relative deviation between the observed tidal range and the physical tidal range, as shown in the following formula: ; In the above formula To observe tidal range; The physical tide level data is the tidal range. This means restricting the range of the function value within the parentheses to the interval [0,1].

[0010] Preferably, the residual normality index RNI is constructed by combining the residual normality test and the residual distribution shape. A Shapiro-Wilk normality test is performed on the physical error sequence to obtain the significance probability; and skewness and excess kurtosis are calculated using the following formulas: ; In the above formula This represents the mean of the physical tide level error sequence; The error at time t represents the physical tide level. Skewness; Excess kurtosis; c1 and c2 are preset scale constants used to normalize skewness and excess kurtosis; P sw is the p-value for the Shapiro–Wilk normality test, n represents the number of samples in the physical tide level error sequence, and RNI is the residual normality index, which is used to characterize the degree of normality of the physical tide level residual sequence.

[0011] Preferably, the Physical Stability Index (PSI) is constructed using windowing error fluctuations. The verification segment is divided into K windows of length L, and the index set is used in the k-th window. The above calculation uses the following formula: ; In the above formula Let be the RMSE of the physical model within the k-th window; mean(r) is the mean RMSE of each window; std(r) is the standard deviation of the RMSE of each window; CV is called the coefficient of variation; Ik Represents the set of window indices.

[0012] Preferably, the relative gain index RGI is constructed by the relative decrease of the root mean square error, as shown in the following formula: ; In the above formula, RMSE P For physical model errors, RMSE M For the error corrected by machine learning, RGI represents the relative gain exponent, O t P represents the observed tide level at time t. t M represents the physical tide level at time t. t T represents the machine learning-corrected tide level at time t. v This represents the set of indices for the verification segment.

[0013] Preferably, the Consistency Improvement Index (CII) simultaneously characterizes the consistency improvement over the time window and the improvement at key tidal points, and calculates the RMSE of the physical model in the k-th window. P,k With machine learning model RMSE M,k The formula is as follows: ; In the above formula, C win This is expressed as the window consistency improvement rate; K is the number of windows; Represents the set of climaxes; This represents the set of verification periods; Let I be an indicator function; if the condition within the parentheses is true, it takes the value 1, otherwise it takes the value 0. Ipeak represents a binary indicator of whether the high tide point has improved; Ithrough represents a binary indicator of whether the low tide point has improved; w1, w2, and w3 are preset weighting coefficients, which, when added together, result in RMSE. M,k RMSE p,K These represent the root mean square error of the physical model in k-window computation and the root mean square error of the machine learning model, respectively.

[0014] Preferably, the significance confidence index (SCI) is constructed using paired nonparametric statistical tests, as shown in the following formula: ; In the above formula, This is expressed as the absolute error of the physical model at time t; To predict the absolute error of machine learning at time t; This represents the difference between the machine learning mean absolute error and the physical mean absolute error. <0 indicates that machine learning outperforms the physical model in the sense of mean absolute error; The expression is an indicator function; it takes the value 1 if the condition in parentheses is true, and 0 otherwise. Pw represents the p-value obtained from the Wilcoxon paired signed-rank test; SCI represents the significance confidence index.

[0015] Preferably, the Comprehensive Improved Reliability Index (CRI) is a weighted fusion of the gain normalization term, the consistency improvement term, the significance term, and the artifact penalty term, as shown in the following formula: ; In the above formula, G is the gain component of the relative gain RGI normalized to [0,1]; g0 is a preset gain scaling constant used to control the normalization intensity; α, β, γ, and δ are preset weighting coefficients that add up to 1; S art In response to the issue of artifact penalties, the CII stated that it is comprehensively improving its reliability system.

[0016] Preferably, the artifact penalty S art It includes at least three types of consistency criteria: amplitude compression, excessive smoothing, and peak shape degradation, as shown in the following formula: ; In the above formula, ρ σ The amplitude ratio is determined by the ratio of the standard deviation of the measured tide level in the validation section to the standard deviation of the machine learning tide level; ρ r The roughness ratio is determined by the ratio of the standard deviation of the measured difference sequence to the standard deviation of the machine learning difference sequence; ρ p The ratio of peak counts is determined by combining the measured tidal peak index set with the machine learning peak index; c1, c2, and c3 are preset deduction coefficients. , , The preset threshold constant is used; std(·) is the standard deviation operator in statistics; M represents the validation segment machine learning sequence; O represents the validation segment measured sequence; P M P represents the set of peak indices of the machine learning sequence in the validation segment; O ΔO represents the set of peak indices of the observed sequence in the validation segment, ΔO represents the first difference of the measured sequence, and ΔM represents the first difference of the machine learning sequence.

[0017] Preferably, the result report includes site identification information, time range, training length, window length, values ​​and ranking of eight indices, and the final applicability level and reasoning based on the judgment rules, and supports outputting the report in tabular form.

[0018] The present invention further provides a tidal level correction applicability evaluation system for tidal stations, the system comprising: The alignment and preprocessing module is used to perform unified time reference alignment and preprocessing on the measured tide level time series. A physical error sequence construction module is built to construct physical error sequences. The Machine Learning Error Sequence Building Module is used to build machine learning error sequences. The module for obtaining physical-side performance characteristic indices and machine learning-side performance characteristic indices is used to obtain physical-side performance characteristic indices and machine learning-side performance characteristic indices. The output module is used to determine the target site based on the physical performance characteristic index and the machine learning performance characteristic index, combined with the constraints of tidal physical mechanism, and output the applicability evaluation conclusion.

[0019] The technical effects and advantages of this invention are as follows: The method and system for evaluating the applicability of tidal level correction at tidal stations use the physical tidal level output by an arbitrary unsteady harmonic analysis physical model as a constraint framework and the corrected tidal level of a machine learning model as a comparison object. By constructing error sequences of measured tidal levels, physical tidal levels, and machine learning tidal levels, eight novel performance indices are calculated respectively. Thus, under the constraint of physical mechanisms, a quantitative judgment result on whether to recommend the use of machine learning correction is given. Based on the physical tidal level constructed by tidal harmonic analysis, the system simultaneously considers the sufficiency of the physical model's fit, tidal range coverage, residual statistical properties, and time stability. It also combines information such as the relative error gain, time consistency improvement degree, statistical significance, and comprehensive improvement reliability of the machine learning model to form a standardized, multi-index evaluation system. The system compares and analyzes the physical model and the machine learning model, and outputs an objective judgment result on whether to recommend the use of machine learning correction, thus providing a reliable basis for tidal level prediction and feature acquisition under tidal conditions. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 This is a schematic diagram of the GUI operation interface of the present invention; Figure 3 This is a schematic diagram of the report exported by the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] This invention provides, for example Figure 1 The method for evaluating the applicability of tidal level correction at a tidal station, as shown, includes the following steps: S1, Input tidal observation data, use harmonic analysis to calculate physical tide level, and use machine learning algorithm to obtain corrected tide level sequence, and perform time alignment and outlier processing on relevant data; S2, construct error sequences for the physical model and the machine learning model respectively, calculate performance characteristic indicators such as physical fit sufficiency index, tidal range coverage index, residual normality index, physical stability index, relative gain index, consistency improvement index, significance confidence index, and comprehensive improved reliability index, and obtain physical effectiveness and improved reliability evaluation results by combining statistical test methods; S3, based on physical mechanism constraints and multi-index threshold rules, comprehensively determines whether machine learning should be used at the site and generates an applicability evaluation report; In this implementation case, tidal characteristics are recalculated and captured, using the publicly available database of the University of Hawaii Sea Level Center as a verification example; Based on the observed tide level sequence of tidal stations, an evaluation process of physical tide level—machine learning-corrected tide level—multi-index comprehensive judgment is constructed, which is divided into S1, S2 and S3 as shown in the figure. The input data adopts the form of structured table, which includes at least the Time column, the Physical model output column, the Observed tide level column, and the Machine Learning-corrected tide level column. In order to ensure the repeatability of the indicators and the feasibility of engineering implementation, the training segment and the evaluation segment are strictly separated: by default, the first 744 hours are the training segment, and the subsequent data are used as an independent evaluation segment to calculate the physical model error and the machine learning correction error and output the conclusion. Phase S1 involves tidal data input and dual-path tidal level calculation. First, station tidal level data is imported into the system, and field integrity is verified. If necessary fields such as Time, Physical, Observed, or ML are missing, the calculation stops and an error is displayed. Subsequently, in engineering implementation, two input methods are allowed: first, directly inputting a result file that already contains Physical and ML data; second, having an external module perform tidal harmonic analysis to obtain the Physical sequence, and then having an external machine learning module output the ML sequence before importing it into the system. Regardless of the method used, the three sequences (physical, observed, and corrected) are aligned along the same time axis, and corresponding subsequences are extracted from the evaluation segment to form the physical residual sequence and the corrected residual sequence, providing a unified data foundation for subsequent index calculations. Phase S2 is the evaluation index calculation phase, comprising two links: physical mechanism usability index and machine learning gain credibility index, converging at the end into a comprehensive credibility improvement index. Firstly, regarding the fit of the physical model itself, a "physical fit sufficiency index" is calculated. This index normalizes and scores the physical model's ability to interpret observed tidal levels within the evaluation segment, providing a graded explanation based on thresholds to determine whether the physical framework of the station is sufficiently good and whether there is still room for improvement. Secondly, a "tidal range coverage index" is calculated. This index measures the physical model's ability to characterize extreme value amplitudes by comparing the consistency between observed and physical tidal ranges within the evaluation segment, avoiding the situation where "the fit is good near the mean." However, the amplitude of high and low tides is significantly distorted; third, calculate the "residual normality index" to perform a normality test on the physical residuals and combine skewness and kurtosis characteristics for comprehensive scoring, which is used to characterize whether there is still a significant systematic structure in the physical model residuals, and then judge whether the machine learning correction has learnable space; fourth, calculate the "physical stability index", divide the evaluation segment into fixed time windows (default 168 hours) to calculate the RMSE of each window, and use the degree of dispersion between windows to characterize whether the physical model error drifts significantly over time; when the stability is insufficient, it will indicate that the site may have changes in the physical error structure caused by the enhancement of non-stationary processes, changes in boundary conditions, or regional nonlinear effects; The four physical performance characteristic indices are, in order: Physical Fit Adequacy Index (PFA), Tidal Coverage Index (TRC), Residual Normality Index (RNI), and Physical Stability Index (PSI). Their calculation principles are as follows: The Physical Fit Adequacy Index (PFA) is constructed based on the Nash–Sutcliffe efficiency coefficient, and the formula is as follows: ; In the above formula To verify the tidal observation data at time t, To verify the mean of the tidal observation data, For the physical harmonic analysis of tidal level data at time t, verification segment is performed. This means restricting the range of the function value within the parentheses to the interval [0,1].

[0023] The Tidal Cover Index (TRC) is constructed by the relative deviation between observed tidal range and physical tidal range, and the formula is as follows: ; In the above formula To observe tidal range; The physical tide level data is the tidal range. This means restricting the range of the function value within the parentheses to the interval [0,1].

[0024] The residual normality index (RNI) is constructed by combining the residual normality test and the residual distribution shape. The significance probability is obtained by performing the Shapiro–Wilk normality test on the physical error sequence; and skewness and excess kurtosis are calculated using the following formulas: ; In the above formula This represents the mean of the physical tide level error sequence; The error at time t represents the physical tide level. Skewness; Excess kurtosis; c1 and c2 are preset scale constants used to normalize skewness and excess kurtosis; P sw is the p-value for the Shapiro–Wilk normality test, n represents the sample size of the physical tide level error series, and RNI is an abbreviation for Residual Normality Index, which is used to characterize the degree of normality of the physical tide level residual series and is a comprehensive index.

[0025] The Physical Stability Index (PSI) is constructed using windowing error fluctuations. The validation segment is divided into K windows of length L, and the index set is used in the k-th window. The above calculation uses the following formula: ; In the above formula is the RMSE of the physical model within the k-th window; mean(r) is the mean RMSE of each window; std(r) is the standard deviation of the RMSE of each window; CV is called the coefficient of variation.

[0026] In the machine learning gain evaluation process, the "relative gain index" is first calculated, using the percentage decrease in RMSE over the evaluation segment as the core metric. Simultaneously, the magnitude of the decrease in MAE is given to measure whether the overall error has been substantially reduced after correction. Recommendations such as "highly recommended / recommended / optional / not recommended" are provided according to preset levels. Next, the consistency improvement index is calculated, again using a 168-hour window to compare the RMSE performance of the physical and corrected results across each window, and the percentage of windows where the correction outperforms the physical is statistically analyzed. Simultaneously, peak and trough identification is performed on the high and low tides of the observation sequence (setting a minimum interval to match tidal cycle characteristics), and the absolute error at the peak and trough is compared. Whether or not there is improvement is determined to avoid pseudo-improvements such as "a slight decrease in overall RMSE but the destruction of key features at high or low points." Furthermore, a "significance credibility index" is calculated, using the absolute error sequence of the evaluation segment as the object for paired significance testing: when the mean absolute error of the corrected result is not better than the physical result, the index is directly judged as low credibility; when there is indeed improvement, it is mapped to a credibility score based on the significance level and an explanation is given, statistically constraining the misjudgment of accidental improvements. The four performance characteristic indices on the machine learning side are, in order: Relative Gain Index (RGI), Consistent Improvement Index (CII), Significance Credibility Index (SCI), and Comprehensive Improvement Reliability Index (CRI). Their principles are as follows: The relative gain exponent (RGI) is constructed based on the relative decrease in the root mean square error, as shown in the following formula: ; In the above formula, RMSE P For physical model errors, RMSE M The error is corrected by machine learning, RGI is the relative gain exponent, and O is the error. t Let P be the observed tide level at time t. t Let M be the physical tide level at time t. t This is the machine learning-corrected tide level at time t. Therefore, this formula calculates the relative gain of the machine learning, T. v This represents the set of indices for our verification segment.

[0027] The Consistency Improvement Index (CII) simultaneously characterizes consistency improvement over time windows and improvement at key tidal points, and calculates the RMSE of the physical model in the k-th window. P,k With machine learning model RMSE M,k The formula is as follows: ; In the above formula, C win This is expressed as the window consistency improvement rate; K is the number of windows; Represents the set of climaxes; This represents the set of verification periods; Let I be an indicator function; if the condition within the parentheses is true, it takes the value 1, otherwise it takes the value 0. Ipeak represents a binary indicator of whether the high tide point has improved; Ithrough represents a binary indicator of whether the low tide point has improved; w1, w2, and w3 are preset weighting coefficients, which, when added together, result in RMSE. M,k RMSE p,K These represent the root mean square error of the physical model in k-window computation and the root mean square error of the machine learning model, respectively. RMSE is an abbreviation for root mean square error.

[0028] The Significance Confidence Index (SCI) is constructed using paired nonparametric statistical tests, and the formula is as follows: ; In the above formula, This is expressed as the absolute error of the physical model at time t; To predict the absolute error of machine learning at time t; This represents the difference between the machine learning mean absolute error and the physical mean absolute error. <0 indicates that machine learning outperforms the physical model in the sense of mean absolute error; It is represented as an indicator function, taking the value 1 if the condition in parentheses is true, and 0 otherwise. Pw refers to the p-value obtained from the Wilcoxon paired signed-rank test; SCI is the significance confidence index, which is a more objective indicator for determining the confidence level of a significance test.

[0029] The Comprehensive Improved Reliability Index (CRI) is a weighted fusion of the gain normalization term, consistency improvement term, significance term, and artifact penalty term, as shown in the following formula: ; In the above formula, G is the gain component of the relative gain RGI normalized to [0,1]; g0 is a preset gain scaling constant used to control the normalization intensity; α, β, γ, and δ are preset weighting coefficients that add up to 1; S art The artifact penalty, or CII, stands for Comprehensive Improvement Reliability Institution. It is a comprehensive index obtained by weighting the gain normalization term, consistency improvement term, significance term, and artifact penalty term. Its significance lies in comprehensively judging whether the modification and improvement of the machine learning scheme is reliable.

[0030] Artifact Punishment S art It includes at least three types of consistency criteria: amplitude compression, excessive smoothing, and peak shape degradation. The calculation principle formula is as follows: ; In the above formula, ρ σ The amplitude ratio is determined by the ratio of the standard deviation of the measured tide level in the validation section to the standard deviation of the machine learning tide level; ρ rThe roughness ratio is determined by the ratio of the standard deviation of the measured difference sequence to the standard deviation of the machine learning difference sequence; ρ p The ratio of peak counts is determined by combining the measured tidal peak index set with the machine learning peak index; c1, c2, and c3 are preset deduction coefficients. , , The preset threshold constant is used; std(·) is the standard deviation operator in statistics; M represents the validation segment machine learning sequence; O represents the validation segment measured sequence; P M P represents the set of peak indices of the machine learning sequence in the validation segment; O ΔO represents the set of peak indices of the observed sequence in the validation segment; ΔO represents the first difference of the measured sequence; and ΔM represents the first difference of the machine learning sequence.

[0031] To improve engineering reliability and prevent algorithmic artifacts such as over-smoothing, a "Comprehensive Improvement Credibility Index" is introduced in the synthesis phase. This index weights and fuses relative gain, consistency improvement, and significance credibility, while additionally calculating artifact detection: comparing the standard deviation ratio, differential roughness, and peak count ratio between the corrected and observed sequences. If the correction result shows significant amplitude compression, over-smoothing, or missing peaks, the comprehensive credibility is reduced to ensure that the recommended conclusions not only pursue error reduction but also meet the fidelity requirements of tidal morphology characteristics. Finally, a complete set of evaluation index results is formed as input for the S3 stage judgment.

[0032] Phase S3 is the judgment and report output phase combining mechanism and performance characteristics. Key results such as physical fit adequacy, relative gain, significance credibility, and consistent improvement are mapped to discrete scores. A comprehensive scoring system is applied according to the following rules: reducing the necessity of introducing machine learning when the physical model is already excellent; increasing the recommendation level when the machine learning error decreases significantly and statistically; and decreasing the recommendation level when the improvement is inconsistent or insignificant. The final decision level (e.g., strongly recommended, recommended, optional, not recommended, strongly not recommended) and a list of reasons are output, ensuring the conclusion is interpretable and traceable. A results report is then generated: on the one hand, it displays the values, levels, and explanations of each index on the interface; on the other hand, it can be exported as an Excel report file. The report includes at least an overall conclusion page, a key statistics page, and a conclusion page, and the values, levels, significance test results, and recommendation reasons of each indicator are permanently output, facilitating batch site evaluation, horizontal comparison, and archiving management. Through the above process, this implementation method achieves an integrated judgment on whether the physical model is effective, whether the machine learning correction is truly effective, and whether the correction has stable and credible promotional value, meeting the engineering, interpretable, and reproducible evaluation requirements of the machine learning correction effect for tidal sites.

[0033] The present invention further provides a tidal level correction applicability evaluation system for tidal stations, the system comprising: The alignment and preprocessing module is used to perform unified time reference alignment and preprocessing on the measured tide level time series. A physical error sequence construction module is built to construct physical error sequences. The Machine Learning Error Sequence Building Module is used to build machine learning error sequences. The module for obtaining physical-side performance characteristic indices and machine learning-side performance characteristic indices is used to obtain physical-side performance characteristic indices and machine learning-side performance characteristic indices. The output module is used to determine the target site based on the physical performance characteristic index and the machine learning performance characteristic index, combined with the constraints of tidal physical mechanism, and output the applicability evaluation conclusion.

[0034] like Figure 2 This invention provides a GUI interface that allows users to manipulate input data and, through calculation, obtain various visual indices, an eight-index result table, and the final conclusion. The results can also be output as a table.

[0035] like Figure 3 This is the content of the table of results derived from this invention, which describes the results of each index, corresponding characteristics, and final conclusions.

[0036] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0037] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.

[0038] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0039] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0040] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0041] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

[0042] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for evaluating the applicability of tidal level correction at tidal stations, characterized in that: include: The measured tide level time series were aligned with a unified time reference and preprocessed. Construct physical error sequences and machine learning error sequences respectively; Obtain the physical-side performance characteristic index and the machine learning-side performance characteristic index; Based on the physical performance characteristic index and the machine learning performance characteristic index, the target site is determined by the constraints of tidal physical mechanism, and the applicability evaluation conclusion is output.

2. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 1, characterized in that: The physical performance characteristic indices include: physical fit adequacy index, tidal range coverage index, residual normality index, and physical stability index. The machine learning side performance characteristic indices include: relative gain index, consistency improvement index, significance confidence index, and comprehensive improvement reliability index.

3. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 2, characterized in that: The physical fit adequacy index includes: Based on the Nash–Sutcliffe efficiency coefficient, the formula is as follows: ; In the formula: This indicates the verification segment of tidal observation data at time t. This represents the mean of the tidal observation data for the verification section. This represents the physical harmonic analysis of tidal level data at time t. This means restricting the range of the function value within the parentheses to the interval [0,1].

4. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 3, characterized in that: The tidal range coverage index includes: The formula is constructed by observing the relative deviation between tidal range and physical tidal range, as follows: ; In the formula: This represents the tidal range as observed in the data. This represents the tidal range, which is the physical tidal data. This means restricting the range of the function value within the parentheses to the interval [0,1].

5. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 2, characterized in that: The residual normality index includes: The residual normality test and residual distribution shape are used together to construct a model. A Shapiro–Wilk normality test is performed on the physical error sequence to obtain the significance probability. Skewness and excess kurtosis are then calculated using the following formulas: ; In the formula: This represents the mean of the physical tide level error sequence; The error at time t represents the physical tide level. Indicates skewness; Indicates excess kurtosis; c1 and c2 represent preset scale constants used to normalize skewness and excess kurtosis; P sw denoted by p-value for the Shapiro–Wilk normality test, n represents the number of samples in the physical tide level error sequence, and RNI is the residual normality index, used to characterize the degree of normality of the physical tide level residual sequence.

6. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 3, characterized in that: The physical stability index includes: By constructing a windowing error fluctuation model, the validation segment is divided into K windows of length L, and the index set is used in the k-th window. The above calculation uses the following formula: ; In the formula: Let represent the RMSE of the physical model within the k-th window; mean(r) represents the mean RMSE of each window; std(r) represents the standard deviation of the RMSE of each window; CV represents the coefficient of variation; I k Represents the set of window indices.

7. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 3, characterized in that: The relative gain index includes: Constructed based on the relative decrease of the root mean square error, the formula is as follows: ; In the formula: RMSE P RMSE represents the error in the physical model. M RGI represents the error after machine learning correction, and O represents the relative gain exponent. t P represents the observed tide level at time t. t M represents the physical tide level at time t. t T represents the machine learning-corrected tide level at time t. v This represents the set of indices for the verification segment.

8. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 3, characterized in that: The consistency improvement index includes: Characterizing time window consistency improvements and key tidal point improvements, the RMSE of the physical model is calculated in the k-th window. P,k With machine learning model RMSE M,k The formula is as follows: ; In the formula, C win This is expressed as the window consistency improvement rate; K represents the number of windows. Represents the set of climaxes; This represents the set of verification periods; This is represented as an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise. peak A binary indicator representing whether the climax point has been improved; I through A binary indicator representing whether the low tide point has improved; w1, w2, and w3 represent preset weighting coefficients, satisfying the condition that the sum of the three is 1, indicating RMSE. M,k RMSE p,K These represent the root mean square error of the physical model in k-window computation and the root mean square error of the machine learning model, respectively.

9. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 3, characterized in that: The significance and credibility index includes: The paired nonparametric statistical test is used to construct the test, and the formula is as follows: ; In the formula: This is expressed as the absolute error of the physical model at time t; This represents the absolute error of the machine learning prediction at time t; This represents the difference between the machine learning mean absolute error and the physical mean absolute error. <0 indicates that machine learning outperforms the physical model in the sense of mean absolute error; The expression is an indicator function; it takes the value 1 if the condition in parentheses is true, and 0 otherwise. Pw represents the p-value obtained from the Wilcoxon paired signed-rank test; SCI represents the significance confidence index.

10. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 9, characterized in that: The comprehensive reliability enhancement index includes: The weighted fusion of the gain normalization term, consistency improvement term, significance term, and artifact penalty term is as follows: ; In the formula: G represents the gain component of the relative gain RGI normalized to [0,1]; g0 represents the preset gain scaling constant, used to control the normalization intensity; α, β, γ, δ represent preset weighting coefficients, which add up to 1; S art The term "CII" indicates a penalty for artifacts, and "CII" signifies a comprehensive improvement in reliability.

11. The method for evaluating the applicability of tidal level correction at a tidal station according to claim 10, characterized in that: The artifact penalty includes: The consistency criteria for the three mechanisms of amplitude compression, excessive smoothing, and peak shape degradation are as follows: ; In the formula, ρ σ The amplitude ratio is determined by the ratio of the standard deviation of the measured tide level in the validation section to the standard deviation of the machine learning tide level; ρ r The roughness ratio is determined by the ratio of the standard deviation of the measured difference sequence to the standard deviation of the machine learning difference sequence; ρ p The ratio of peak counts is determined by combining the measured tidal peak index set with the machine learning peak index; c1, c2, and c3 are preset deduction coefficients. , , The preset threshold constant is used; std(·) is the standard deviation operator in statistics; M represents the validation segment machine learning sequence; O represents the validation segment measured sequence; P M P represents the set of peak indices of the machine learning sequence in the validation segment; O ΔO represents the set of peak indices of the observed sequence in the validation segment, ΔO represents the first difference of the measured sequence, and ΔM represents the first difference of the machine learning sequence.

12. A tidal level correction applicability evaluation system for tidal stations, characterized in that: The system includes: The alignment and preprocessing module is used to perform unified time reference alignment and preprocessing on the measured tide level time series. A physical error sequence construction module is built to construct physical error sequences. The Machine Learning Error Sequence Building Module is used to build machine learning error sequences. The module for obtaining physical-side performance characteristic indices and machine learning-side performance characteristic indices is used to obtain physical-side performance characteristic indices and machine learning-side performance characteristic indices. The output module is used to determine the target site based on the physical performance characteristic index and the machine learning performance characteristic index, combined with the constraints of tidal physical mechanism, and output the applicability evaluation conclusion.