A method and device for constructing a probability prediction model for taking fresh water in an estuary

By constructing a machine learning model based on historical salinity and water-air coupling sequences, the problem of low accuracy in predicting the probability of freshwater extraction in estuary areas was solved, enabling precise prediction of the timing of freshwater extraction and scientific allocation of water resources.

CN120493558BActive Publication Date: 2025-11-18SUN YAT SEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510657874.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-11-18
Estimated Expiration
2045-05-21

Smart Images

  • Figure CN120493558B_ABST
    Figure CN120493558B_ABST
Patent Text Reader

Abstract

The application discloses a kind of estuary area to take fresh water probability prediction model construction method and device, to solve the existing estuary area to take fresh water probability prediction technology is mainly based on historical hydrological observation data and fixed salinity threshold to carry out empirical judgment, resulting in the accuracy of taking fresh water probability prediction is lower technical problem.Method includes that the historical salinity sequence obtained is binarized based on pre-set salinity threshold, generates taking fresh water situation sequence;According to multiple pre-set lag time, historical salinity sequence and the historical water-air coupling sequence obtained, determine target lag time salinity sequence and target lag time water-air coupling sequence;Using pre-set model configuration item according to target lag time salinity sequence, target lag time water-air coupling sequence and taking fresh water situation sequence, constructs multiple initial estuary area to take fresh water probability prediction model;Each initial estuary area to take fresh water probability prediction model is screened based on pre-set model evaluation strategy, determines target estuary area to take fresh water probability prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrological and water resources application technology, and in particular to a method and apparatus for constructing a freshwater extraction probability prediction model in estuary areas. Background Technology

[0002] Estuarine areas are sensitive zones where fresh and saltwater meet, and the probability of obtaining freshwater (i.e., the likelihood of acquiring usable freshwater resources) directly impacts water supply for coastal cities, agricultural irrigation, and ecological balance. Due to the dynamic interaction of factors such as tides, runoff, and saltwater intrusion, the spatial and temporal distribution of freshwater resources is highly complex. Accurate prediction of freshwater extraction probability has significant engineering value for water resource management, salinization control, and emergency dispatch.

[0003] To meet this need, the freshwater extraction technology has emerged. This technology refers to a water resource allocation technique used in estuaries, through the scientific management of water conservancy projects and the utilization of tidal and saline tide patterns, to extract freshwater from river channels during periods of low or no saline tide influence to satisfy the water needs of industry, agriculture, and residential use. Its importance is particularly evident during dry seasons or periods of high salinity, when water intakes in estuaries are susceptible to salinity and unable to draw water normally. By accurately predicting the probability of freshwater extraction and applying the freshwater extraction technology in conjunction with saline tide avoidance, the efficiency and security of water resource utilization in estuaries can be significantly improved.

[0004] Existing techniques for predicting the probability of freshwater extraction in estuaries mainly rely on historical hydrological observation data and empirical judgments based on fixed salinity thresholds (such as 0.5‰). They estimate the extraction period by statistically analyzing the correlation between tidal cycles, runoff, and salinity. However, this method requires manually setting fixed salinity thresholds and depends on expert experience to determine the standard of "acceptable freshwater." It cannot adapt to saltwater intrusion caused by climate change, resulting in low accuracy in predicting the probability of freshwater extraction. Summary of the Invention

[0005] This invention provides a method and apparatus for constructing a freshwater extraction probability prediction model in estuary areas, which addresses the technical problem that existing freshwater extraction probability prediction technologies in estuary areas mainly rely on historical hydrological observation data and fixed salinity thresholds for empirical judgment, resulting in low accuracy in predicting freshwater extraction probability.

[0006] The first aspect of this invention provides a method for constructing a freshwater extraction probability prediction model in estuary areas, comprising:

[0007] Historical salinity sequences and historical water-air coupling sequences are obtained, and the historical salinity sequences are binarized based on a preset salinity threshold to generate a freshwater extraction sequence.

[0008] Based on multiple preset lag times, the historical salinity sequence, and the historical water-gas coupling sequence, determine the target lag salinity sequence and the target lag water-gas coupling sequence.

[0009] Multiple initial freshwater extraction probability prediction models for estuary areas are constructed using preset model configuration items based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction sequence.

[0010] Based on the pre-set model evaluation strategy, the prediction models for freshwater extraction probability in each initial estuary area are screened to determine the prediction model for freshwater extraction probability in the target estuary area.

[0011] Optionally, determining the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence based on multiple preset lag times, the historical salinity sequence, and the historical water-gas coupling sequence includes:

[0012] The historical salinity sequence and the historical water-gas coupling sequence are subjected to time lag processing using the preset lag times respectively, and the time-lag salinity sequence and time-lag water-gas coupling sequence corresponding to each preset lag time are determined.

[0013] The historical salinity sequence is correlated with each of the lag salinity sequences, and the first correlation value corresponding to each lag salinity sequence is output.

[0014] The historical salinity sequence is correlated with each of the time-lapse water-gas coupling sequences, and the second correlation value corresponding to each of the time-lapse water-gas coupling sequences is output.

[0015] The lag salinity sequence corresponding to the largest first correlation value is selected as the optimal lag salinity sequence;

[0016] The time-delayed water-gas coupling sequence corresponding to the largest second correlation value is selected as the optimal time-delayed water-gas coupling sequence.

[0017] The optimal time-lapse salinity sequence and the optimal time-lapse water-gas coupling sequence are preprocessed respectively to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence.

[0018] Optionally, the preprocessing of the optimal time-lapse salinity sequence and the optimal time-lapse water-gas coupling sequence to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence includes:

[0019] The missing value removal operation is performed on the optimal lag salinity sequence and the optimal lag water-gas coupling sequence respectively to determine the initial lag salinity sequence and the initial lag water-gas coupling sequence;

[0020] The initial time-delay salinity sequence and the initial time-delay water-gas coupling sequence are normalized respectively to determine the intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence;

[0021] The intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence are standardized respectively to determine the target time-delay salinity sequence and the target time-delay water-gas coupling sequence.

[0022] Optionally, the preset model configuration includes a random seed and a callback function; the step of using the preset model configuration to construct multiple initial estuarine freshwater extraction probability prediction models based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction sequence includes:

[0023] Using the target time-lapse salinity sequence and the target time-lapse water-air coupling sequence as forecasting factors, and the freshwater extraction sequence as the prediction target, a freshwater extraction probability prediction model for the estuary area is constructed.

[0024] The freshwater extraction probability prediction model in the estuary area is repeatedly trained based on random seeds and callback functions to determine multiple initial freshwater extraction probability prediction models in the estuary area.

[0025] Optionally, the step of screening the initial freshwater extraction probability prediction models for each of the pre-set model evaluation strategies to determine the freshwater extraction probability prediction model for the target estuary area includes:

[0026] The pre-set model evaluation strategy is used to evaluate the performance of each of the initial estuary freshwater extraction probability prediction models, and to determine the accuracy, confusion matrix statistics, and model prediction consistency value of each of the initial estuary freshwater extraction probability prediction models.

[0027] The initial freshwater estuary probability prediction model corresponding to the highest accuracy, the largest confusion matrix statistic, and the largest model prediction consistency value was selected as the target freshwater estuary probability prediction model.

[0028] Optionally, it also includes:

[0029] When the salinity sequence to be measured and the water-gas coupling sequence to be measured are received, the salinity sequence to be measured and the water-gas coupling sequence to be measured are preprocessed respectively to generate the target salinity sequence and the target water-gas coupling sequence.

[0030] The target salinity sequence and the target water-air coupling sequence are input into the freshwater extraction probability prediction model of the target estuary area to generate freshwater extraction probability prediction results for the estuary area.

[0031] The second aspect of this invention provides an apparatus for constructing a freshwater extraction probability prediction model in estuary areas, comprising:

[0032] The acquisition module is used to acquire historical salinity sequences and historical water-air coupling sequences, and to perform binarization processing on the historical salinity sequences based on a preset salinity threshold to generate a freshwater extraction sequence.

[0033] The determination module is used to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence based on multiple preset lag times, the historical salinity sequence, and the historical water-gas coupling sequence;

[0034] The construction module is used to construct multiple initial freshwater extraction probability prediction models for estuary areas based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction situation sequence using preset model configuration items.

[0035] The screening module is used to screen the freshwater extraction probability prediction models for each of the initial estuary areas based on the preset model evaluation strategy, and to determine the freshwater extraction probability prediction model for the target estuary area.

[0036] A computer device provided in a third aspect of the present invention includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method for constructing a freshwater sampling probability prediction model in the estuary area as described in any of the preceding claims.

[0037] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed, it implements the steps of the method for constructing a freshwater extraction probability prediction model in estuary areas as described in any of the preceding claims.

[0038] The fifth aspect of the present invention provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein, when the program instructions are executed by a computer, the computer performs the steps of the method for constructing a freshwater sampling probability prediction model in the estuary area as described in any of the preceding claims.

[0039] As can be seen from the above technical solutions, the present invention has the following advantages:

[0040] The above-mentioned technical solution of the present invention provides a method for constructing a freshwater extraction probability prediction model in estuary areas. First, historical salinity sequences and historical water-air coupling sequences are obtained, and the historical salinity sequences are binarized based on preset salinity thresholds to generate a freshwater extraction sequence. Next, target lag salinity sequences and target lag water-air coupling sequences are determined based on multiple preset lag times, historical salinity sequences, and historical water-air coupling sequences. Using preset model configuration items, multiple initial freshwater extraction probability prediction models for estuary areas are constructed based on the target lag salinity sequences, target lag water-air coupling sequences, and freshwater extraction sequence. Finally, based on a preset model evaluation strategy, each initial freshwater extraction probability prediction model for estuary areas is screened to determine the target freshwater extraction probability prediction model for estuary areas. Based on the above scheme, and combined with preset lag time and preset model configuration items, the obtained historical salinity sequence and historical water-air coupling sequence are processed to obtain multiple initial freshwater extraction probability prediction models for estuary areas. The process of selecting the freshwater extraction probability prediction models for target estuary areas through preset model evaluation strategies is as follows: This invention directly uses the obtained freshwater extraction probability prediction model for target estuary areas to predict the freshwater extraction probability, and predicts the best time for freshwater extraction in advance. It does not require manually setting a fixed salinity threshold, which can effectively cope with saltwater intrusion and thus improve the accuracy of freshwater extraction probability prediction. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating the steps of a method for constructing a freshwater extraction probability prediction model in estuary areas, as provided in Embodiment 1 of the present invention.

[0043] Figure 2 This is a schematic diagram of the short-term lightening forecast confusion matrix provided in Embodiment 1 of the present invention;

[0044] Figure 3 This is a schematic diagram of the forecast results of the test set of the short-term forecast model with different lead times provided in Embodiment 1 of the present invention;

[0045] Figure 4 This is a schematic diagram of the confusion matrix for medium- to long-term lightening forecasting provided in Embodiment 1 of the present invention;

[0046] Figure 5 This is a schematic diagram of the forecast results of the test set of the 7-day long-term forecast model for light-dampening conditions provided in Embodiment 1 of the present invention.

[0047] Figure 6 This is a schematic diagram of the forecast results of the test set of the 15-day long-term forecast model for light-dampening conditions provided in Embodiment 1 of the present invention.

[0048] Figure 7 This is a schematic diagram of the forecast results of the test set of the 30-day long-term forecast model for light-dampening conditions provided in Embodiment 1 of the present invention.

[0049] Figure 8 This is a flowchart of the steps for making predictions using a freshwater extraction probability prediction model for a target estuary area, as provided in Embodiment 2 of the present invention.

[0050] Figure 9 This is a structural block diagram of a device for constructing a freshwater extraction probability prediction model in estuary areas, as provided in Embodiment 3 of the present invention. Detailed Implementation

[0051] This invention provides a method and apparatus for constructing a freshwater extraction probability prediction model in estuary areas, which addresses the technical problem that existing freshwater extraction probability prediction technologies in estuary areas mainly rely on historical hydrological observation data and fixed salinity thresholds for empirical judgment, resulting in low accuracy in predicting freshwater extraction probability.

[0052] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0053] Please see Figure 1 , Figure 1 The flowchart illustrates the steps of constructing a freshwater extraction probability prediction model in estuary areas, as provided in Embodiment 1 of the present invention.

[0054] This invention provides a method for constructing a freshwater extraction probability prediction model in estuary areas, comprising:

[0055] Step 101: Obtain historical salinity sequences and historical water-air coupling sequences, and perform binarization processing on the historical salinity sequences based on preset salinity thresholds to generate freshwater extraction sequences.

[0056] It should be noted that the historical hourly salinity series (historical salinity series) of the main water intakes or water plants in the target estuary area during the dry season is collected. The obtained historical salinity series is divided into two categories according to national drinking water safety standards: exceeding the standard and not exceeding the standard. Specifically, salinity below a preset salinity threshold (250 mg / L) is recorded as not exceeding the standard (freshwater intake is permissible), and salinity above 250 mg / L is recorded as exceeding the standard (freshwater intake is not permissible). This is transformed into a binary classification problem, forming a new series, denoted as the freshwater intake sequence. Specifically, the salinity data in the obtained historical salinity series is binarized: values ​​exceeding 250 mg / L are recorded as 1, and values ​​below 250 mg / L are recorded as 0, resulting in a freshwater intake sequence composed of 0s and 1s.

[0057] For example, the Modaomen waterway was selected as the target estuary area for the study. Hourly salinity data (historical salinity series) of the Pinggang station during the dry season from 2019 to 2023 were collected. The salinity series were divided into two categories according to national drinking water safety standards: exceeding the standard and not exceeding the standard. Salinity below 250 mg / L was recorded as acceptable for freshwater intake (not exceeding the standard), and salinity not lower than 250 mg / L was recorded as unacceptable for freshwater intake (exceeding the standard), thus forming a freshwater intake sequence. An example of the compiled freshwater intake sequence is shown in Table 1.

[0058] Table 1. Examples of light-dip sequence

[0059]

[0060] Furthermore, sequences of other factors related to salinity exceeding the target area during the same period were obtained. According to relevant studies, saline intrusion frequently occurs in estuaries due to the combined influence of multiple factors such as runoff, tidal currents, wind speed, and wind direction, leading to salinity exceeding the target area. Therefore, historical water-air coupled sequences during the same period as the historical salinity sequence were collected. The historical water-air coupled sequences include historical runoff sequences, historical tidal level sequences, and historical wind speed and direction sequences.

[0061] For example, hourly flow rates at Ma Kou station, hourly tide levels at San Zao station, and wind speed data in the 10m direction (u and v directions) of the Macau region from the European Centre for Medium-Range Weather Forecasts (ECMWF) ERA5-Land hourly data from 1950 to present dataset were collected for the same period from 2019 to 2023. The correlation between the salinity series and the collected flow rate, tide level, and wind speed series in the u and v directions was analyzed. The optimal lag sequence with the highest correlation to the salinity series was selected. The selected forecast factors, with a lead time of 1-6 hours, are shown in Table 2. The optimal lag sequence required for the target lead time model was compiled, missing values ​​were removed, and normalization was performed. Where S is the hourly average salinity at Pinggang Station (mg / L); F is the hourly average flow rate at Makou Station (m3 / s); T is the hourly tide level at Sanzao Station (m); Wu is the east-west projection of the wind speed in Macau (m / s), and Wv is the north-south projection of the wind speed in Macau (m / s); S(t-1) represents the salinity one hour ahead of the salinity forecast value, and other variables follow the same principle.

[0062] Table 2 Forecasting factors for the model of lightness probability at different lead times

[0063]

[0064] Step 102: Determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence based on multiple preset lag times, historical salinity sequences, and historical water-gas coupling sequences.

[0065] Specifically, step 102 may include the following sub-steps S21-S26:

[0066] Step S21: Use each preset lag time to perform lag processing on the historical salinity sequence and the historical water-gas coupling sequence to determine the lag salinity sequence and lag water-gas coupling sequence corresponding to each preset lag time.

[0067] Step S22: Calculate the correlation between the historical salinity series and each lag salinity series, and output the first correlation value corresponding to each lag salinity series;

[0068] Step S23: Calculate the correlation between the historical salinity sequence and each time-lapsed water-gas coupling sequence, and output the second correlation value corresponding to each time-lapsed water-gas coupling sequence;

[0069] Step S24: Select the lag salinity sequence corresponding to the largest first correlation value as the optimal lag salinity sequence;

[0070] Step S25: Select the time-delayed water-gas coupling sequence corresponding to the largest second correlation value as the optimal time-delayed water-gas coupling sequence;

[0071] The time-delayed water-air coupling sequence includes the time-delayed runoff sequence, the time-delayed tide level sequence, and the time-delayed wind speed and direction sequence.

[0072] The optimal time-delayed water-air coupling sequence includes the optimal time-delayed runoff sequence, the optimal time-delayed tidal level sequence, and the optimal time-delayed wind speed and direction sequence.

[0073] It should be noted that, firstly, multiple preset lag times are used to perform time lag processing on the historical salinity series, historical runoff series, historical tide level series, and historical wind speed and direction series, to obtain the time-lapsed salinity series, time-lapsed runoff series, time-lapsed tide level series, and time-lapsed wind speed and direction series corresponding to each preset lag time. For example, if the preset lag times are 3h and 5h, then the time-lapsed salinity series, time-lapsed runoff series, time-lapsed tide level series, and time-lapsed wind speed and direction series with a lag of 3 hours, and the time-lapsed salinity series, time-lapsed runoff series, time-lapsed tide level series, and time-lapsed wind speed and direction series with a lag of 5 hours can be obtained.

[0074] Furthermore, the correlation between different time-lapse sequences (time-lapse salinity sequence, time-lapse runoff sequence, time-lapse tide sequence, and time-lapse wind speed and direction sequence) and historical salinity sequences is analyzed. The optimal time lag for salinity, runoff, tide, and wind speed and direction sequences required by different forecast period models is determined by using a cross-correlation analysis method combined with Gaussian processes. For example, assuming there are currently time-lapse runoff sequences with a time lag of 3 hours and time-lapse runoff sequences with a time lag of 5 hours, the correlation value between the time-lapse runoff sequence with a time lag of 3 hours and the historical salinity sequence is calculated, as well as the correlation value between the time-lapse runoff sequence with a time lag of 5 hours and the historical salinity sequence. If the correlation value between the time-lapse runoff sequence with a time lag of 3 hours and the historical salinity sequence is the largest, then the time-lapse runoff sequence with a time lag of 3 hours is selected as the optimal time-lapse runoff sequence. Similarly, the optimal time-lapse salinity sequence, optimal time-lapse tide sequence, and optimal time-lapse wind speed and direction sequence can be obtained.

[0075] Furthermore, by introducing a Gaussian weighting factor to enhance the ability to capture nonlinear relationships, the weighting term automatically decays outliers far from the mean, focusing on the main distribution area of ​​the data, thus more accurately reflecting the true correlation pattern. This can significantly improve the robustness and interpretability of traditional cross-correlation analysis under nonlinear and noisy data. The specific process of correlation calculation is as follows:

[0076] ;

[0077] in, For historical salinity series and preset lag time The correlation values ​​between the corresponding time-lapse sequences include the first correlation value and the second correlation value, where N is the total length of the time series. These are the characteristic values ​​of the historical salinity series; It is a time-lapse sequence, which includes time-lapse salinity sequence, time-lapse runoff sequence, time-lapse tidal level sequence, and time-lapse wind speed and direction sequence; The standard deviation of the historical salinity series is used for standardization and Gaussian weight calculation; The standard deviation of the time-delay sequence is used for standardization and Gaussian weight calculation. The mean of the historical salinity series; The mean of the time-delay sequence; This is the lag period, i.e., the preset lag time; As a Gaussian weighting term, a Gaussian weight is assigned to the cross-covariance at each time point. The weights are determined by... and The degree to which it deviates from its mean determines this.

[0078] Step S26: Preprocess the optimal time-delay salinity sequence and the optimal time-delay water-gas coupling sequence respectively to determine the target time-delay salinity sequence and the target time-delay water-gas coupling sequence.

[0079] The target time-lapse water-air coupling sequence includes the target time-lapse runoff sequence, the target time-lapse tidal level sequence, and the target time-lapse wind speed and direction sequence.

[0080] Furthermore, step S26 may include the following sub-steps S261-S263:

[0081] Step S261: Perform missing value removal operations on the optimal lag salinity sequence and the optimal lag water-gas coupling sequence respectively to determine the initial lag salinity sequence and the initial lag water-gas coupling sequence;

[0082] Step S262: Normalize the initial time-delay salinity sequence and the initial time-delay water-gas coupling sequence respectively to determine the intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence;

[0083] Step S263: Standardize the intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence respectively to determine the target time-delay salinity sequence and the target time-delay water-gas coupling sequence.

[0084] It should be noted that different sequences with lag time compared to the salinity sequence are organized into a single data frame. Specifically, the optimal lag salinity sequence and the optimal lag water-air coupling sequence are preprocessed separately. This process includes removing missing values ​​and normalizing them. The mean and standard deviation of each feature are calculated, and then the feature values ​​are converted into standardized values ​​with a mean of 0 and a standard deviation of 1 to eliminate the impact of differences in the units and numerical ranges between different features on the model performance.

[0085] Step 103: Using the preset model configuration items, construct multiple initial freshwater extraction probability prediction models for the estuary area based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction situation sequence.

[0086] The preset model configuration options include a random seed and a callback function.

[0087] It should be noted that by determining the feature variable X and the target variable Y, and using the collected salinity, runoff, tide level, wind direction and speed sequences (i.e., the target time-lapse salinity sequence, target time-lapse runoff sequence, target time-lapse tide level sequence, and target time-lapse wind speed and direction sequence) as the forecasting factors (X), and taking the salinity situation as the prediction target (Y), and dividing the training set and test set in a 6:4 ratio, the method of integrating multiple time series features can better capture the dynamic pattern of salinity changes and its correlation with other variables, providing richer information for the model and helping to improve the accuracy of prediction.

[0088] Furthermore, the freshwater harvesting probability prediction model established in this invention is a machine learning model, with a structure resembling a Gated Recurrent Unit (GRU). Specifically, the aforementioned feature variables and target variables are input into the model to train and construct a freshwater harvesting probability prediction model (the initial freshwater harvesting probability prediction model for the estuary area). The model employs a two-layer GRU structure: the first layer has 256 units and returns a sequence, while the second layer has 128 units and does not return a sequence. This structural design can progressively extract feature information from the time series, while using a Dropout layer to prevent overfitting. Finally, a fully connected layer with a sigmoid activation function outputs the prediction result, making it suitable for binary classification problems. During model training, an EarlyStopping callback function is used to monitor the validation set loss (val_loss). When the validation set loss no longer decreases within five consecutive epochs, training is stopped early, and the model's optimal weights are restored (restore_best_weights=True). This mechanism avoids overtraining, saves computational resources, and ensures good model performance on the test set. Before training each model, a random seed (seed(s)) is set to ensure that the training process is repeatable, which facilitates debugging and verification of model performance and subsequent selection of the optimal model.

[0089] Specifically, step 103 may include the following sub-steps S31-S32:

[0090] Step S31: Using the target time-lapse salinity sequence and the target time-lapse water-air coupling sequence as forecasting factors, and the freshwater extraction situation sequence as the prediction target, construct a freshwater extraction probability prediction model for the estuary area.

[0091] Step S32: Based on random seeds and callback functions, repeatedly train the freshwater sourcing probability prediction model in the estuary area to determine multiple initial freshwater sourcing probability prediction models in the estuary area.

[0092] For example, the collected salinity, runoff, tide level, and wind direction and speed sequences (i.e., target time-lapse salinity sequence, target time-lapse runoff sequence, target time-lapse tide level sequence, and target time-lapse wind speed and direction sequence) are all used as forecast factors (X), and freshwater availability is used as the prediction target (Y). The training and test sets are divided in a 6:4 ratio. A gated recurrent unit (GRU) is used to input the above feature variables and target variables into the model to train and construct a freshwater availability prediction model (an initial freshwater availability probability prediction model for estuaries). The model adopts a two-layer GRU structure: the first GRU has 256 units and returns the sequence, while the second GRU has 128 units and does not return a sequence. This structural design can progressively extract feature information from the time series, while using a Dropout layer to prevent overfitting. Finally, a fully connected layer with a sigmoid activation function outputs the prediction result, making it suitable for binary classification problems. During model training, the EarlyStopping callback function is used to monitor the validation set loss (val_loss). When the validation set loss no longer decreases within 5 consecutive epochs, training is stopped early, and the model's optimal weights are restored (restore_best_weights=True). This mechanism avoids overtraining, saves computational resources, and ensures good model performance on the test set. Before training each model, 10 random seeds (seed(s)) are set, and 10 relatively stable models are trained for each epoch and temporarily stored.

[0093] Step 104: Based on the pre-set model evaluation strategy, screen the freshwater extraction probability prediction models for each initial estuary area and determine the freshwater extraction probability prediction model for the target estuary area.

[0094] Specifically, step 104 may include the following sub-steps S41-S42:

[0095] Step S41: Using a pre-set model evaluation strategy, evaluate the performance of the freshwater extraction probability prediction model for each initial estuary area, and determine the accuracy, confusion matrix statistics, and model prediction consistency value of each initial estuary area freshwater extraction probability prediction model.

[0096] Step S42: Select the initial estuary freshwater probability prediction model corresponding to the highest accuracy, the largest confusion matrix statistic, and the largest model prediction consistency value as the target estuary freshwater probability prediction model.

[0097] It should be noted that, to more comprehensively evaluate the model's performance in predicting the probability of lightness, accuracy, MCC, and κ are combined to comprehensively measure model performance. Accuracy directly reflects the model's prediction accuracy on the overall data. However, in evaluating binary classification models, while accuracy is intuitive, it is insufficient in scenarios with imbalanced classes or requiring more rigorous evaluation. Therefore, this invention adopts a pre-defined model evaluation strategy, which includes accuracy calculation and the use of Matthews Correlation Coefficient (MCC) and Cohen's Kappa (κ) as higher-order confusion matrix derived indicators, providing a more comprehensive and robust method for evaluating model performance.

[0098] Furthermore, accuracy: Accuracy is one of the most intuitive performance metrics, defined as the proportion of samples correctly predicted by the model out of the total number of samples. Its calculation formula is as follows:

[0099] ;

[0100] Wherein, TP (True Positives) represents the number of samples where salinity is not exceeded and the model correctly predicted that salinity is not exceeded; TN (True Negatives) represents the number of samples where salinity exceeds the standard and the model correctly predicted that salinity exceeds the standard; FP (False Positives) represents the number of samples where salinity exceeds the standard and the model incorrectly predicted that salinity is not exceeded; and FN (False Negatives) represents the number of samples where salinity is not exceeded and the model incorrectly predicted that salinity exceeds the standard.

[0101] Furthermore, MCC (Matthews Correlation Coefficient): MCC is a statistic based on the confusion matrix, specifically the confusion matrix statistic. It comprehensively considers four classes of results: true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), primarily focusing on evaluating the overall predictive ability of a binary classification model. Its core idea is to measure the consistency between the model's predictions and the actual results through the concept of covariance. The formula for calculating MCC (confusion matrix statistic) is:

[0102] ;

[0103] The indicators in the formula are the same as those in the accuracy calculation formula. The numerator reflects the difference between correct and incorrect classifications, and the denominator is the standardization factor to eliminate the influence of class imbalance. The MCC value ranges between [-1, 1]. If the MCC value is 1, it indicates that the model makes perfect predictions (all samples are classified correctly). If the MCC value is 0, it indicates that the model is equivalent to random guessing. If the MCC value is -1, it indicates that the model makes completely inverse predictions (all predictions are opposite to the true labels).

[0104] Furthermore, κ (Cohen's Kappa): Cohen's Kappa, or model prediction consistency score, measures the consistency between model predictions and true labels, and is particularly useful for assessing consistency among annotators or eliminating the influence of random guessing in model predictions. Its core is to compare the observed consistency rate (p0, i.e., the model's accuracy) with the random consistency rate (p...). e Assuming the model predictions are independent of the true labels, the expected consistency rate (calculated by random classification) quantifies the model's ability to outperform random guessing. The formula for calculating the model prediction consistency value is as follows:

[0105] ;

[0106] ;

[0107] Among them, κ can standardize the difference between the observed consistency rate and the random consistency rate, and is extremely robust to imbalanced data. The value of κ ranges between [-1, 1]. If the value of MCC is 1, it indicates perfect consistency (the model prediction matches the true label perfectly). If the value of MCC is 0, it indicates that the consistency is the same as the random guess. If the value of MCC is -1, it indicates complete inconsistency (the model prediction is completely opposite to the true label).

[0108] Furthermore, from the models with the set random number of seeds, i.e. the initial estuary freshwater leaching probability prediction models, the initial estuary freshwater leaching probability prediction model with the highest test set accuracy and the highest MCC value (confusion matrix statistic) and κ value (model prediction consistency value) is selected as the target estuary freshwater leaching probability prediction model, and saved as pb format (Protocol Buffer) or h5 format (Hierarchical Data Format version 5).

[0109] For example, 10 models with different forecast periods were validated respectively, and the accuracy and AUC statistics of the selected optimal model are shown in Table 3.

[0110] Table 3. Evaluation of the accuracy of short-term light condensation probability forecasts

[0111]

[0112] Further, please refer to Figure 2 Taking models with forecast periods of 1 hour, 6 hours, 12 hours, and 24 hours as examples, the forecast confusion matrix of the model is as follows: Figure 2 As shown, the confusion matrix can display the correctly predicted and incorrectly predicted samples with salinity exceeding or not exceeding the standard in the training and test sets. It can be clearly observed that only a very small number of samples are predicted inaccurately.

[0113] Further, please refer to Figure 3 , Figure 3 The accuracy of salinity predictions for different forecast periods (1 hour, 6 hours, 12 hours, and 24 hours) is shown. Red dots represent incorrect predictions, and green dots represent correct predictions. The figure shows that correct predictions (green) account for the vast majority, indicating high accuracy across all forecast periods. Especially with longer forecast periods (e.g., 24 hours), the predictions remain highly accurate, demonstrating the model's robustness and stability. Furthermore, the distribution of incorrect predictions is relatively dispersed, without a concentrated error trend, indicating that the model adapts well to data changes across different time periods and has strong generalization ability. The table shows that the model exhibits good performance across all forecast periods. Particularly in the 1-hour forecast, the accuracy is 0.9701, and the MCC and κ values ​​reach 0.9395, demonstrating extremely high classification accuracy in the short term. It can effectively identify whether salinity exceeds the standard within the forecast period, thus determining whether water can be drawn. Even with the prediction timeframe extended to 24 hours, the model's test set accuracy was 0.8758, indicating that the model can still effectively identify whether salinity exceeds the standard when handling predictions over a longer time period. Furthermore, the accuracy on the training set was consistently slightly higher than that on the test set, demonstrating that the model learns patterns from the training data well and has good fitting capabilities. Overall, the model exhibits stable and reliable performance across prediction tasks across different timeframes.

[0114] Furthermore, we will further construct medium- and long-term binary classification forecast models, such as... Figure 4 The figure shows the confusion matrix for whether the salinity exceeds the standard for forecast periods of 7 days (168h), 15 days (360h), and 30 days (720h). The number of samples in four different scenarios (predicted exceedance but actually exceedance, predicted exceedance but actually not exceedance, predicted not exceedance but actually exceedance, predicted not exceedance but actually not exceedance) is counted. The number of samples counted on the main diagonal represents the cases where the prediction is accurate. As can be seen from the figure, in different forecast periods, the number of accurate predictions in both the training and test sets is much greater than the number of incorrect predictions.

[0115] Furthermore, Figures 5-7The performance of the model in predicting fading over medium- to long-term forecast periods (7 days, 15 days, and 30 days) is demonstrated. Analysis shows that even within a 30-day forecast period, the model maintains high prediction accuracy, indicating excellent long-term predictive ability and stability. Across all forecast periods, the number of incorrect predictions (represented by red dots) is relatively small and evenly distributed, without concentrated errors. This demonstrates that the model effectively captures data characteristics and reduces errors when processing long-term series data. Furthermore, the consistent performance across different forecast periods indicates its adaptability to forecasting needs at various time scales, making it highly valuable for practical applications.

[0116] Furthermore, the accuracy evaluation of models with different lead times is shown in Table 4, and all models demonstrated good predictive ability. This invention uses hourly-scale data for long-term forecasting, achieving relatively accurate and refined forecasts for the long lead time. In particular, with a 15-day lead time, the accuracy of the test set reached 0.719, which is a relatively high value, indicating that the model has a good ability to distinguish whether the sample salinity exceeds the standard in medium-term forecasting.

[0117] Table 4. Accuracy Evaluation of Medium- and Long-Term Freshwater Extraction Forecast Model

[0118]

[0119] As can be seen from the above, whether it is short-term (within 24 hours) or medium-to-long-term (7 days, 15 days, 30 days), the prediction model can be constructed using the binary classification and machine learning algorithm fusion method for predicting the probability of freshwater harvesting in estuaries proposed in this invention to achieve accurate prediction. When the latest data is collected, the best selected model can be called to predict whether freshwater can be harvested in the target estuary area in the future.

[0120] For comparison of technical effectiveness, existing technologies can be referenced. Freshwater extraction to avoid salinity refers to a water resource allocation technology in estuaries that, through the scientific management of water conservancy projects and the utilization of tidal and salinity tide patterns, extracts freshwater from river channels during periods of low or no salinity tide influence to meet the water needs of industry, agriculture, and residential use. During dry seasons or periods of high salinity tide activity, water intakes in estuaries are easily affected by salinity tides, leading to disruptions in water intake. Currently, there is considerable research on salinity tide forecasting, but less on predicting the probability of freshwater extraction. Therefore, developing a method to accurately predict the probability of freshwater extraction is of paramount importance for ensuring water security in estuaries.

[0121] To address the aforementioned problems, this invention proposes a method for constructing a freshwater extraction probability prediction model for estuary areas. Specifically, historical salinity sequences of the target area are collected and classified into two categories: exceeding the standard and not exceeding the standard, forming a freshwater extraction probability sequence. Further, sequences of other factors related to salinity exceeding the standard in the target area during the same period are collected and preprocessed accordingly. All collected sequences except for the freshwater extraction sequence are input into a machine learning model to construct a freshwater extraction probability prediction model, with the freshwater extraction probability being the prediction target. The model accuracy is verified, and the best model is saved. Newly collected data is input, and the saved model is invoked to predict the future freshwater extraction probability of the target area. By predicting the optimal time for freshwater extraction in advance, scientifically scheduling water conservancy projects, and optimizing water resource allocation, this method can effectively cope with saltwater intrusion and ensure stable and sustainable water supply in estuary areas.

[0122] In this embodiment of the invention, a method for constructing a freshwater extraction probability prediction model for estuary areas is provided. First, historical salinity sequences and historical water-air coupling sequences are acquired, and the historical salinity sequences are binarized based on a preset salinity threshold to generate a freshwater extraction sequence. Next, target lag salinity sequences and target lag water-air coupling sequences are determined based on multiple preset lag times, historical salinity sequences, and historical water-air coupling sequences. Using preset model configuration items, multiple initial freshwater extraction probability prediction models for estuary areas are constructed based on the target lag salinity sequences, target lag water-air coupling sequences, and freshwater extraction sequence. Finally, based on a preset model evaluation strategy, each initial freshwater extraction probability prediction model for estuary areas is screened to determine the target freshwater extraction probability prediction model for estuary areas. Based on the above scheme, and combined with preset lag time and preset model configuration items, the obtained historical salinity sequence and historical water-air coupling sequence are processed to obtain multiple initial freshwater extraction probability prediction models for estuary areas. The process of selecting the freshwater extraction probability prediction models for target estuary areas through preset model evaluation strategies is as follows: This invention directly uses the obtained freshwater extraction probability prediction model for target estuary areas to predict the freshwater extraction probability, and predicts the best time for freshwater extraction in advance. It does not require manually setting a fixed salinity threshold, which can effectively cope with saltwater intrusion and thus improve the accuracy of freshwater extraction probability prediction.

[0123] For better explanation, refer to Figure 8 The flowchart illustrates the steps of predicting the freshwater extraction probability in a target estuary area using a prediction model provided in Embodiment 2 of the present invention. This process may include the following steps:

[0124] Step 801: When the salinity sequence to be measured and the water-gas coupling sequence to be measured are received, preprocess the salinity sequence to be measured and the water-gas coupling sequence to be measured respectively to generate the target salinity sequence and the target water-gas coupling sequence.

[0125] Step 802: Input the target salinity sequence and the target water-air coupling sequence into the freshwater extraction probability prediction model for the target estuary area to generate the freshwater extraction probability prediction results for the estuary area.

[0126] The measured water-air coupling sequence includes the measured runoff sequence, the measured tidal level sequence, and the measured wind speed and direction sequence.

[0127] It should be noted that the latest salinity, runoff, tide, and wind speed and direction sequences to be measured are collected and organized into the corresponding formats according to the processing principles mentioned above, namely, removing missing values, normalizing, and standardizing. The optimal model (the freshwater extraction probability prediction model for the target estuary area) saved for the required forecast period is then called to predict the freshwater extraction probability, thereby providing a reference for the allocation of regional water resources.

[0128] In this embodiment of the invention, the freshwater extraction probability prediction model of the target estuary area is invoked to predict the freshwater extraction probability, so as to know the best time for freshwater extraction in advance, scientifically schedule water conservancy projects, optimize water resource allocation, effectively cope with saltwater intrusion, and ensure the stable and sustainable water supply of the estuary area.

[0129] Please see Figure 9 , Figure 9 This is a structural block diagram of a device for constructing a freshwater extraction probability prediction model in estuary areas, as provided in Embodiment 3 of the present invention.

[0130] This invention provides a device for constructing a freshwater extraction probability prediction model in estuary areas, comprising:

[0131] The acquisition module 901 is used to acquire historical salinity sequences and historical water-air coupling sequences, and to perform binarization processing on the historical salinity sequences based on a preset salinity threshold to generate a freshwater extraction sequence.

[0132] The determination module 902 is used to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence based on multiple preset lag times, historical salinity sequences and historical water-gas coupling sequences;

[0133] Module 903 is used to construct multiple initial freshwater extraction probability prediction models for estuary areas based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction situation sequence using preset model configuration items.

[0134] The screening module 904 is used to screen the freshwater extraction probability prediction models for each initial estuary area based on the pre-set model evaluation strategy, and to determine the freshwater extraction probability prediction model for the target estuary area.

[0135] Furthermore, module 901 includes:

[0136] The first submodule is used to perform time-lapse processing on the historical salinity sequence and the historical water-gas coupling sequence using each preset time lag, and to determine the time-lapsed salinity sequence and time-lapsed water-gas coupling sequence corresponding to each preset time lag.

[0137] The second submodule is used to calculate the correlation between the historical salinity series and each lag salinity series, and output the first correlation value corresponding to each lag salinity series.

[0138] The third submodule is used to calculate the correlation between the historical salinity sequence and each time-lapsed water-gas coupling sequence, and output the second correlation value corresponding to each time-lapsed water-gas coupling sequence.

[0139] The fourth submodule is used to select the lag salinity sequence corresponding to the largest first correlation value as the optimal lag salinity sequence;

[0140] The fifth submodule is used to select the time-delayed water-gas coupling sequence corresponding to the largest second correlation value as the optimal time-delayed water-gas coupling sequence;

[0141] The sixth submodule is used to preprocess the optimal time-lapse salinity sequence and the optimal time-lapse water-gas coupling sequence respectively, and to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence.

[0142] Furthermore, the sixth submodule is specifically used for:

[0143] Missing values ​​were removed from the optimal lag salinity sequence and the optimal lag water-gas coupling sequence, respectively, to determine the initial lag salinity sequence and the initial lag water-gas coupling sequence.

[0144] The initial time-delay salinity sequence and the initial time-delay water-gas coupling sequence were normalized respectively to determine the intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence.

[0145] The intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence are standardized respectively to determine the target time-delay salinity sequence and the target time-delay water-gas coupling sequence.

[0146] Furthermore, the pre-configured model settings include a random seed and a callback function; module 903 is specifically used for:

[0147] Using the target time-lapse salinity sequence and the target time-lapse water-air coupling sequence as forecasting factors, and the freshwater extraction situation sequence as the prediction target, a prediction model for the probability of freshwater extraction in the estuary area is constructed.

[0148] The freshwater sourcing probability prediction model in the estuary area was repeatedly trained based on random seeds and callback functions to determine multiple initial freshwater sourcing probability prediction models in the estuary area.

[0149] Furthermore, the filtering module 904 is specifically used for:

[0150] The pre-set model evaluation strategy was used to evaluate the performance of the freshwater leaching probability prediction model for each initial estuary area, and to determine the accuracy, confusion matrix statistics and model prediction consistency value of each initial estuary area freshwater leaching probability prediction model.

[0151] The initial freshwater estuary probability prediction model corresponding to the highest accuracy, the largest confusion matrix statistic, and the largest model prediction consistency value was selected as the target freshwater estuary probability prediction model.

[0152] In one optional device embodiment, it further includes:

[0153] When the salinity sequence to be measured and the water-gas coupling sequence to be measured are received, preprocessing is performed on the salinity sequence to be measured and the water-gas coupling sequence to be measured respectively to generate the target salinity sequence and the target water-gas coupling sequence.

[0154] The target salinity sequence and the target water-air coupling sequence are input into the freshwater extraction probability prediction model for the target estuary area to generate freshwater extraction probability prediction results for the estuary area.

[0155] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, modules, and sub-modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0156] This invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the steps of the method for constructing a freshwater sampling probability prediction model in the estuary area as described in any of the above embodiments.

[0157] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the method for constructing a freshwater sampling probability prediction model in the estuary area as described in any of the above embodiments.

[0158] This invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the method for constructing a freshwater sampling probability prediction model in the estuary area as described in any of the above embodiments.

[0159] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0161] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a predictive model for the probability of freshwater harvesting in estuary areas, characterized in that, include: Historical salinity sequences and historical water-air coupling sequences are obtained, and the historical salinity sequences are binarized based on a preset salinity threshold to generate a freshwater extraction sequence. Based on multiple preset lag times, the historical salinity sequence, and the historical water-gas coupling sequence, determine the target lag salinity sequence and the target lag water-gas coupling sequence. Multiple initial freshwater extraction probability prediction models for estuary areas are constructed using preset model configuration items based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction sequence. Based on the pre-set model evaluation strategy, the freshwater extraction probability prediction models for each of the initial estuary areas are screened to determine the freshwater extraction probability prediction model for the target estuary area. The step of determining the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence based on multiple preset lag times, the historical salinity sequence, and the historical water-gas coupling sequence includes: The historical salinity sequence and the historical water-gas coupling sequence are subjected to time lag processing using the preset lag times respectively, and the time-lag salinity sequence and time-lag water-gas coupling sequence corresponding to each preset lag time are determined. The historical salinity sequence is correlated with each of the lag salinity sequences, and the first correlation value corresponding to each lag salinity sequence is output. The historical salinity sequence is correlated with each of the time-lapse water-gas coupling sequences, and the second correlation value corresponding to each of the time-lapse water-gas coupling sequences is output. The lag salinity sequence corresponding to the largest first correlation value is selected as the optimal lag salinity sequence; The time-delayed water-gas coupling sequence corresponding to the largest second correlation value is selected as the optimal time-delayed water-gas coupling sequence. The optimal time-lapse salinity sequence and the optimal time-lapse water-gas coupling sequence are preprocessed respectively to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence.

2. The method for constructing a freshwater extraction probability prediction model in estuary areas according to claim 1, characterized in that, The step of preprocessing the optimal time-lapse salinity sequence and the optimal time-lapse water-gas coupling sequence to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence includes: The missing value removal operation is performed on the optimal lag salinity sequence and the optimal lag water-gas coupling sequence respectively to determine the initial lag salinity sequence and the initial lag water-gas coupling sequence; The initial time-delay salinity sequence and the initial time-delay water-gas coupling sequence are normalized respectively to determine the intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence; The intermediate time-delay salinity sequence and the intermediate time-delay water-gas coupling sequence are standardized respectively to determine the target time-delay salinity sequence and the target time-delay water-gas coupling sequence.

3. The method for constructing a freshwater extraction probability prediction model in estuary areas according to claim 1, characterized in that, The preset model configuration includes a random seed and a callback function; the preset model configuration is used to construct multiple initial estuarine freshwater extraction probability prediction models based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction sequence, including: Using the target time-lapse salinity sequence and the target time-lapse water-air coupling sequence as forecasting factors, and the freshwater extraction sequence as the prediction target, a freshwater extraction probability prediction model for the estuary area is constructed. The freshwater extraction probability prediction model in the estuary area is repeatedly trained based on random seeds and callback functions to determine multiple initial freshwater extraction probability prediction models in the estuary area.

4. The method for constructing a freshwater extraction probability prediction model in estuary areas according to claim 1, characterized in that, The pre-set model evaluation strategy is used to screen the initial freshwater extraction probability prediction models for each estuary area to determine the target estuary area freshwater extraction probability prediction model, including: The pre-set model evaluation strategy is used to evaluate the performance of each of the initial estuary freshwater extraction probability prediction models, and to determine the accuracy, confusion matrix statistics, and model prediction consistency value of each of the initial estuary freshwater extraction probability prediction models. The initial freshwater estuary probability prediction model corresponding to the highest accuracy, the largest confusion matrix statistic, and the largest model prediction consistency value was selected as the target freshwater estuary probability prediction model.

5. The method for constructing a freshwater extraction probability prediction model in estuary areas according to claim 1, characterized in that, Also includes: When the salinity sequence to be measured and the water-gas coupling sequence to be measured are received, the salinity sequence to be measured and the water-gas coupling sequence to be measured are preprocessed respectively to generate the target salinity sequence and the target water-gas coupling sequence. The target salinity sequence and the target water-air coupling sequence are input into the freshwater extraction probability prediction model of the target estuary area to generate freshwater extraction probability prediction results for the estuary area.

6. A device for constructing a freshwater extraction probability prediction model in estuary areas, applied to the method for constructing a freshwater extraction probability prediction model in estuary areas as described in claim 1, characterized in that, include: The acquisition module is used to acquire historical salinity sequences and historical water-air coupling sequences, and to perform binarization processing on the historical salinity sequences based on a preset salinity threshold to generate a freshwater extraction sequence. The determination module is used to determine the target time-lapse salinity sequence and the target time-lapse water-gas coupling sequence based on multiple preset lag times, the historical salinity sequence, and the historical water-gas coupling sequence; The construction module is used to construct multiple initial freshwater extraction probability prediction models for estuary areas based on the target time-lapse salinity sequence, the target time-lapse water-air coupling sequence, and the freshwater extraction situation sequence using preset model configuration items. The screening module is used to screen the freshwater extraction probability prediction models for each of the initial estuary areas based on the preset model evaluation strategy, and to determine the freshwater extraction probability prediction model for the target estuary area.

7. A computer device, characterized in that, The system includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the method for constructing a freshwater extraction probability prediction model in the estuary area as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the method for constructing a freshwater extraction probability prediction model in the estuary area as described in any one of claims 1-5.

9. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by a computer, the computer performs the method for constructing a freshwater extraction probability prediction model in the estuary area as described in any one of claims 1-5.