Dissolved oxygen concentration prediction method and device based on machine learning

A dissolved oxygen concentration prediction model was constructed by using machine learning methods. By using time series data to capture the nonlinear relationship between dissolved oxygen and other water quality parameters in the aquatic environment, the problem of high computational cost and limited prediction performance in traditional methods was solved, and more efficient and accurate dissolved oxygen concentration prediction was achieved.

CN121983167APending Publication Date: 2026-05-05GUANGZHOU INST OF GEOGRAPHY GUANGDONG ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU INST OF GEOGRAPHY GUANGDONG ACAD OF SCI
Filing Date
2025-12-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional methods for predicting dissolved oxygen concentration rely on complex physical, chemical, and biological process equations, which are computationally expensive and fail to capture the complex nonlinear relationships between dissolved oxygen and other water quality parameters in the aquatic environment, resulting in limited predictive performance.

Method used

Machine learning methods were employed to construct and train a model using daily lag time series of dissolved oxygen concentration and water quality parameters. This model captures and expresses the complex nonlinear relationship between dissolved oxygen and other water quality parameters in the aquatic environment, and constructs a dissolved oxygen concentration prediction model.

Benefits of technology

It improves the accuracy and efficiency of dissolved oxygen concentration prediction, better captures nonlinear relationships in complex aquatic environments, and provides more accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983167A_ABST
    Figure CN121983167A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of dissolved oxygen concentration prediction, in particular to a dissolved oxygen concentration prediction method and device based on machine learning, computer equipment and a storage medium, and the method comprises the steps: obtaining a dissolved oxygen concentration daily scale lag characteristic time sequence and a water quality parameter daily scale lag characteristic time sequence of a sample region; performing combination and training set division according to the dissolved oxygen concentration daily scale lag characteristic time sequence and the water quality parameter daily scale lag characteristic time sequence, and constructing a plurality of training sets; performing model construction and model training according to the plurality of training sets by adopting a machine learning method to obtain a dissolved oxygen concentration prediction model; and obtaining a water quality parameter daily scale lagging characteristic time sequence of the target area, inputting the water quality parameter daily scale lagging characteristic time sequence of the target area into the dissolved oxygen concentration prediction model to carry out dissolved oxygen concentration prediction, and obtaining a dissolved oxygen concentration prediction result of the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dissolved oxygen concentration prediction technology, and in particular to a machine learning-based method, apparatus, computer device, and storage medium for predicting dissolved oxygen concentration. Background Technology

[0002] Dissolved oxygen is a key indicator for assessing surface water environmental quality, especially crucial for the health of estuarine and nearshore marine ecosystems. Low-dissolved oxygen water flowing into the ocean negatively impacts marine habitats and nearshore environments, while the dramatic changes in coastal river water quality and the complex aquatic environment make dissolved oxygen concentration prediction a challenging task.

[0003] Traditional water quality prediction models rely on complex physical, chemical, and biological process equations, which, while capable of effectively simulating pollutant diffusion and spatial distribution, are complex to construct, computationally expensive, and extremely demanding in their input data requirements. Traditional statistical methods using multiple linear regression, while capable of quantifying the long-term impact of individual environmental factors on dissolved oxygen, fail to capture and express the complex nonlinear relationships between dissolved oxygen and other water quality parameters, resulting in limited predictive performance. Summary of the Invention

[0004] Based on this, the purpose of this invention is to provide a machine learning-based method, device, computer equipment, and storage medium for predicting dissolved oxygen concentration. This method utilizes daily lag time series of dissolved oxygen concentration and daily lag time series of water quality parameters for model construction and training, capturing and expressing the complex nonlinear relationship between dissolved oxygen and other water quality parameters in the aquatic environment, thereby improving the accuracy and efficiency of dissolved oxygen concentration prediction.

[0005] In a first aspect, embodiments of this application provide a dissolved oxygen concentration prediction method based on machine learning, comprising the following steps: The daily lag time series of dissolved oxygen concentration and the daily lag time series of water quality parameters in the sample area are obtained. The daily lag time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days. The daily lag time series of water quality parameters includes water quality parameter data for several consecutive days. Based on the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters, several training sets are constructed by combining and dividing the training sets; machine learning methods are used to construct and train the model based on the several training sets to obtain a dissolved oxygen concentration prediction model. Obtain the daily lag characteristic time series of water quality parameters in the target area, input the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration, and obtain the dissolved oxygen concentration prediction result of the target area.

[0006] Secondly, embodiments of this application provide a dissolved oxygen concentration prediction device based on machine learning, comprising: The data acquisition module is used to acquire the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters in the sample area. The daily lag characteristic time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days; the daily lag characteristic time series of water quality parameters includes water quality parameter data for several consecutive days. The model training module is used to combine and divide the training set according to the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters to construct several training sets; and to use machine learning methods to construct and train the model based on the several training sets to obtain a dissolved oxygen concentration prediction model. The dissolved oxygen concentration prediction module is used to obtain the daily lag characteristic time series of water quality parameters in the target area, input the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration and obtain the dissolved oxygen concentration prediction result of the target area.

[0007] Thirdly, embodiments of this application provide a computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the machine learning-based dissolved oxygen concentration prediction method as described in the first aspect.

[0008] Fourthly, embodiments of this application provide a storage medium storing a computer program that, when executed by a processor, implements the steps of the machine learning-based dissolved oxygen concentration prediction method described in the first aspect.

[0009] In this application embodiment, a machine learning-based method, apparatus, computer device, and storage medium for predicting dissolved oxygen concentration are provided. The method utilizes daily-scale lag feature time series of dissolved oxygen concentration and daily-scale lag feature time series of water quality parameters for model construction and training, thereby capturing and expressing the complex nonlinear relationship between dissolved oxygen and other water quality parameters in the aquatic environment, and improving the accuracy and efficiency of dissolved oxygen concentration prediction.

[0010] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0011] Figure 1 A flowchart illustrating a machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application; Figure 2 This is a flowchart illustrating step S2 in a machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application. Figure 3 This is a flowchart illustrating step S23 of a machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application. Figure 4 This is a flowchart illustrating step S232 of a machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application. Figure 5 This is a flowchart illustrating step S24 of a machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application. Figure 6 This is a flowchart illustrating step S244 of a machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application. Figure 7 A schematic diagram of the structure of a machine learning-based dissolved oxygen concentration prediction device provided in one embodiment of this application; Figure 8 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation

[0012] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0013] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0014] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0015] Please see Figure 1 , Figure 1 The flowchart illustrates a machine learning-based dissolved oxygen concentration prediction method according to an embodiment of this application. The method includes the following steps: S1: Obtain the daily lag time series of dissolved oxygen concentration and water quality parameters in the sample area.

[0016] The execution entity of the machine learning-based dissolved oxygen concentration prediction method is a prediction device (hereinafter referred to as the prediction device). In an optional embodiment, the prediction device may be a computer device, a server, or a server cluster composed of multiple computer devices.

[0017] In this embodiment, the prediction device obtains the daily-scale lag characteristic time series of dissolved oxygen concentration and the daily-scale lag characteristic time series of water quality parameters in the sample area. The daily-scale lag characteristic time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days; the daily-scale lag characteristic time series of water quality parameters includes water quality parameter data for several consecutive days.

[0018] In an optional embodiment, the predictive device extracts dissolved oxygen concentration data and water quality parameter data from a preset database. The dissolved oxygen concentration data and water quality parameter data are collected by water quality sensors deployed in the estuary water body and transmitted to a central data server via an Internet of Things (IoT) communication module. The central data server performs format verification on the collected data and stores it in the database.

[0019] Furthermore, the central data server cleans the format-validated dissolved oxygen concentration and water quality parameter data. First, the central data server identifies and marks missing values ​​caused by power outages or equipment malfunctions. For a few missing values, linear interpolation is used to impute them to maintain data continuity. Next, the central data server averages and aggregates the imputed dissolved oxygen concentration and water quality parameter data into daily-scale data to smooth short-term noise and highlight daily trends, generating daily-scale dissolved oxygen concentration time-series datasets and daily-scale water quality parameter time-series datasets. Based on these datasets, the central data server constructs lag features, generating daily-scale lag feature time series datasets for dissolved oxygen concentration and water quality parameters.

[0020] S2: Combine and divide the daily lag characteristic time series of dissolved oxygen concentration and daily lag characteristic time series of water quality parameters to construct several training sets; use machine learning methods to construct and train the model based on the several training sets to obtain a dissolved oxygen concentration prediction model.

[0021] In this embodiment, the prediction device combines the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters and divides the training set to construct several training sets. Using machine learning methods, the model is constructed and trained based on the several training sets to obtain a dissolved oxygen concentration prediction model. This enables the machine learning model to learn the time dependence and lag effect of dissolved oxygen, thereby improving the accuracy of dissolved oxygen concentration prediction.

[0022] The dissolved oxygen concentration prediction model includes a first dissolved oxygen concentration prediction model and a second dissolved oxygen concentration prediction model; please refer to [link / reference]. Figure 2 , Figure 2 The flowchart of S2 in the machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application includes steps S21 to S24, as follows: S21: Calculate the standard deviation of dissolved oxygen concentration data for several days in each training set to obtain the standard deviation of dissolved oxygen concentration for each training set; classify and divide each training set according to the standard deviation of dissolved oxygen concentration and the preset dissolved oxygen concentration threshold to obtain several first training sets and several second training sets.

[0023] In this embodiment, the prediction device calculates the standard deviation of dissolved oxygen concentration data for several days in each training set to obtain the standard deviation of dissolved oxygen concentration for each training set.

[0024] The prediction device categorizes the training sets based on the standard deviation of dissolved oxygen concentration and a preset dissolved oxygen concentration threshold, resulting in several first training sets and several second training sets. Specifically, if the standard deviation of dissolved oxygen concentration is less than or equal to the dissolved oxygen concentration threshold, the dissolved oxygen concentration of the corresponding training set is considered to have stable fluctuations and is used as the first training set for training a first machine learning model. The first machine learning model employs a Bagging-Boosting model, which includes several XGBoost learners.

[0025] If the standard deviation of dissolved oxygen concentration is greater than the dissolved oxygen concentration threshold, the dissolved oxygen concentration of the corresponding training set is considered to fluctuate drastically, and is used as the first training set for training the second machine learning model, wherein the second machine learning model is a stacked model.

[0026] S22: Input several sets of the first training set into a preset first machine learning model for model training to obtain a first dissolved oxygen concentration prediction model.

[0027] The first machine learning model uses a Bagging-Boosting model, which includes several XGBoost learners.

[0028] In this embodiment, the prediction device inputs several sets of the first training set into a preset first machine learning model for model training to obtain a first dissolved oxygen concentration prediction model.

[0029] Please see Figure 3 , Figure 3 The flowchart of S23 in the machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application includes steps S221 to S223, as follows: S221: Sampling with replacement is performed on several of the first training sets to generate several first samples; the first samples are used as the input datasets for the current iteration, and each input dataset is input into each XGBoost learner in the current iteration to predict dissolved oxygen concentration, thereby obtaining the dissolved oxygen concentration prediction data of each input dataset in the current iteration.

[0030] In this embodiment, the prediction device performs sampling with replacement on several first training sets to generate several first samples, wherein the first samples include several water quality parameter data.

[0031] The prediction device uses the first sample as the input dataset for the current iteration, and inputs each input dataset into each XGBoost learner in the current iteration to predict dissolved oxygen concentration, thereby obtaining the dissolved oxygen concentration prediction data for each input dataset in the current iteration.

[0032] S222: Based on the input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the current iteration, and the dissolved oxygen concentration prediction data of the corresponding input datasets of the previous iteration, train each XGBoost learner of the current iteration to obtain each XGBoost learner of the next iteration.

[0033] In this embodiment, the prediction device trains each XGBoost learner for the current iteration based on each input dataset of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input dataset of the current iteration, and the dissolved oxygen concentration prediction data of the corresponding input dataset of the previous iteration, to obtain each XGBoost learner for the next iteration.

[0034] Please see Figure 4 , Figure 4 The flowchart of S222 in the machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application includes step S2221, as follows: S2221: Based on the input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the previous iteration, and the preset regularization loss function, obtain the loss value of each XGBoost learner in the current iteration.

[0035] The regularization loss function is:

[0036] In the formula, For the first t The loss value of the XGBoost learner in the next iteration. The dissolved oxygen concentration prediction data is the corresponding input dataset for the current iteration. The dissolved oxygen concentration prediction data is from the corresponding input dataset of the previous iteration. For decision tree functions, Given the input dataset for the current iteration, This is a regularization term.

[0037] In this embodiment, the prediction device obtains the loss value of each XGBoost learner in the current iteration based on the input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the previous iteration, and a preset regularization loss function.

[0038] S223: Use the first sample as the input dataset for the next iteration, input each input dataset into each XGBoost learner in the current iteration, and repeat the dissolved oxygen concentration prediction and model training to obtain the first dissolved oxygen concentration prediction model.

[0039] In this embodiment, the prediction device uses the first sample as the input dataset for the next iteration, inputs each input dataset into each XGBoost learner in the current iteration, and repeatedly performs dissolved oxygen concentration prediction and model training to obtain the first dissolved oxygen concentration prediction model.

[0040] S23: Construct a second machine learning model based on several sets of the second training set; input several sets of the second training set into the first machine learning model for model training to obtain a second dissolved oxygen concentration prediction model.

[0041] The second machine learning model includes a basic learner layer and a meta-learning layer. The basic learner layer includes several sub-dissolved oxygen concentration prediction models. The meta-learning layer includes a random forest model. The sub-dissolved oxygen concentration prediction models are machine learning models constructed using the water quality parameter data as independent variables and the dissolved oxygen concentration data as dependent variables. They include decision tree models, highly random tree models, gradient boosting models, and extreme gradient boosting models.

[0042] In this embodiment, the prediction device constructs a second machine learning model based on several second training sets; the several second training sets are input into the first machine learning model for model training to obtain a second dissolved oxygen concentration prediction model.

[0043] Please see Figure 5 , Figure 5 The flowchart of S23 in the machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application includes steps S231 to S234, as follows: S231: Perform sampling with replacement on several sets of the second training set to generate several sets of second samples; construct a model based on the several sets of second samples to obtain several sub-dissolved oxygen concentration prediction models.

[0044] In this embodiment, the prediction device performs sampling with replacement on several second training sets to generate several second samples, wherein the second samples include several water quality parameter data. The prediction device constructs a model based on several second samples to obtain several sub-dissolved oxygen concentration prediction models. The sub-dissolved oxygen concentration prediction models are machine learning models constructed using the water quality parameter data as independent variables and the dissolved oxygen concentration data as dependent variables, including decision tree models, extremely random tree models, gradient boosting models, and extreme gradient boosting models.

[0045] S232: Input several sets of the second training set into several sub-dissolved oxygen concentration prediction models for model training, and obtain several sub-dissolved oxygen concentration prediction models after training and the dissolved oxygen concentration prediction data corresponding to each second sample output by the several sub-dissolved oxygen concentration prediction models after training.

[0046] In this embodiment, the prediction device inputs several second training sets into several sub-dissolved oxygen concentration prediction models for model training, thereby obtaining several trained sub-dissolved oxygen concentration prediction models and the dissolved oxygen concentration prediction data corresponding to each second sample output by the several trained sub-dissolved oxygen concentration prediction models.

[0047] Specifically, both the decision tree model and the extremely random tree model use minimizing the mean squared error function as the objective function, wherein the mean squared error function is:

[0048] In the formula, To minimize the mean square error, The number of elements in the second training set. This is the dissolved oxygen concentration prediction data corresponding to the second training set. The dissolved oxygen concentration label data is the data obtained from the daily lag feature time series of dissolved oxygen concentration corresponding to the second training set.

[0049] The gradient boosting model uses the mean squared error function as the objective function, wherein the mean squared error function is:

[0050]

[0051] In the formula, This is the mean square error. This represents the dissolved oxygen concentration prediction data output by the gradient boosting model in the m-th iteration. For learning rate, This is the data processing function for the gradient boosting model in the m-th iteration.

[0052] The extreme gradient boosting model uses a regularized loss function as its objective function, wherein the regularization term in the regularized loss function used in the extreme gradient boosting model is:

[0053]

[0054] In the formula, For regularization terms, These are complexity control parameters. The L2 regularization coefficient is... The leaf weight vector, =( , ,..., ), The L1 regularization coefficient is... , These are the first and second derivatives of the loss function with respect to the model's predicted values, respectively.

[0055] S233: Perform prediction vector transformation on the dissolved oxygen concentration prediction data corresponding to each second sample output by each sub-dissolved oxygen concentration prediction model after training to obtain the dissolved oxygen concentration prediction vector corresponding to each second sample; combine the dissolved oxygen concentration prediction vector corresponding to the same second sample to obtain the training feature matrix of each second sample.

[0056] In this embodiment, the prediction device performs prediction vector transformation on the dissolved oxygen concentration prediction data corresponding to each second sample output by each sub-dissolved oxygen concentration prediction model after training, to obtain the dissolved oxygen concentration prediction vector corresponding to each second sample, as described below:

[0057]

[0058]

[0059]

[0060] In the formula, , , , These are the dissolved oxygen concentration prediction vectors for each of the second samples. , , , These are the corresponding prediction vector transformation functions.

[0061] The prediction device combines the dissolved oxygen concentration prediction vectors corresponding to the same second sample to obtain the training feature matrix of each second sample.

[0062] S234: Obtain dissolved oxygen concentration label data corresponding to each of the second samples; construct and train a random forest model based on the training feature matrix of each of the second samples and the dissolved oxygen concentration label data to obtain the trained random forest model.

[0063] In this embodiment, the prediction device obtains dissolved oxygen concentration label data corresponding to each of the second samples, wherein the dissolved oxygen concentration label data is obtained from the daily lag feature time series of the dissolved oxygen concentration corresponding to the second sample.

[0064] The prediction device constructs and trains a random forest model based on the training feature matrix of each second sample and the dissolved oxygen concentration label data, and obtains the trained random forest model.

[0065] Please see Figure 6 , Figure 6 The flowchart of S234 in the machine learning-based dissolved oxygen concentration prediction method provided in one embodiment of this application includes steps S2341 to S2343, as follows: S2341: Perform sampling with replacement on the training feature matrix of each second sample to generate several third samples.

[0066] In this embodiment, the prediction device performs sampling with replacement on the training feature matrix of each second sample to generate several third samples, wherein the third samples include several candidate features, and the candidate features are dissolved oxygen concentration prediction vectors.

[0067] S2342: Randomly select several candidate features for each of the third samples to construct the candidate feature set of the current node of the decision tree corresponding to each of the third samples; according to the gain maximization principle, select the target feature as the node for splitting based on the feature values ​​of several candidate features in the candidate feature set of the current node; stop splitting when the preset recursive splitting condition is met; construct each decision tree; combine the decision trees to construct the random forest model.

[0068] In this embodiment, the prediction device randomly selects several candidate features for each of the third samples to construct a candidate feature set for the current node of the decision tree corresponding to each of the third samples. Based on the gain maximization principle, the target feature is selected as a node for splitting according to the feature values ​​of several candidate features in the candidate feature set of the current node. When the preset recursive splitting condition is met, the splitting stops, and each decision tree is constructed. The decision trees are then combined to construct the random forest model.

[0069] S2343: Input the training feature matrix of each of the second samples and the dissolved oxygen concentration label data into the random forest model for training to obtain the trained random forest model.

[0070] In this embodiment, the prediction device inputs several of the third samples into the random forest model for training, obtaining a trained random forest model. The training process aims to find the optimal model parameters that minimize the error between the random forest's prediction results and the true target variable.

[0071] In an optional embodiment, the prediction device calculates the SHAP value of each water quality parameter for a single predicted value (local) or the entire model (global) based on a trained first dissolved oxygen concentration prediction model, a second dissolved oxygen concentration prediction model, and a SHAP interpreter. Water quality parameter analysis is performed based on Shapley values ​​from cooperative game theory to obtain water quality parameter analysis results, wherein the water quality parameter analysis results include both positive and negative dissolved oxygen driving results.

[0072] S3: Obtain the daily lag characteristic time series of water quality parameters in the target area, input the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration, and obtain the dissolved oxygen concentration prediction result of the target area.

[0073] In this embodiment, the prediction device obtains the daily lag characteristic time series of water quality parameters in the target area, inputs the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration, and obtains the dissolved oxygen concentration prediction result of the target area.

[0074] Specifically, the prediction device obtains a river identifier for the target area, which can be a monitoring station ID or a river name, indicating whether the river in the target area has stable or volatile flow. Based on the river identifier, the prediction device confirms that the daily lag characteristic time series of water quality parameters input to the target area is an object in the dissolved oxygen concentration prediction model, that is, determines that the input object is either the first dissolved oxygen concentration prediction model or the second dissolved oxygen concentration prediction model. The prediction device then performs dissolved oxygen concentration prediction based on the daily lag characteristic time series of water quality parameters in the target area and the confirmed input object, obtaining the dissolved oxygen concentration prediction result for the target area.

[0075] By utilizing daily lag time series of dissolved oxygen concentration and water quality parameters, model construction and training are carried out to capture and express the complex nonlinear relationship between dissolved oxygen and other water quality parameters in the aquatic environment, thereby improving the accuracy and efficiency of dissolved oxygen concentration prediction.

[0076] Please refer to Figure 7 , Figure 7 This is a schematic diagram of a machine learning-based dissolved oxygen concentration prediction device provided in one embodiment of this application. The device can be implemented entirely or partially through software, hardware, or a combination of both. The device 7 includes: The data acquisition module 71 is used to acquire the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters in the sample area. The daily lag characteristic time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days; the daily lag characteristic time series of water quality parameters includes water quality parameter data for several consecutive days. The model training module 72 is used to combine and divide the training set according to the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters to construct several training sets; and to use machine learning methods to construct and train the model according to the several training sets to obtain a dissolved oxygen concentration prediction model. The dissolved oxygen concentration prediction module 73 is used to obtain the daily lag characteristic time series of water quality parameters in the target area, input the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration and obtain the dissolved oxygen concentration prediction result of the target area.

[0077] In this embodiment, the prediction device obtains the daily-scale lag characteristic time series of dissolved oxygen concentration and the daily-scale lag characteristic time series of water quality parameters in the sample area through the data acquisition module. The daily-scale lag characteristic time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days; the daily-scale lag characteristic time series of water quality parameters includes water quality parameter data for several consecutive days. Through the model training module, the device combines and divides the daily-scale lag characteristic time series of dissolved oxygen concentration and the daily-scale lag characteristic time series of water quality parameters to construct several training sets. Using machine learning methods, the device constructs and trains a model based on the several training sets to obtain a dissolved oxygen concentration prediction model. Through the dissolved oxygen concentration prediction module, the device obtains the daily-scale lag characteristic time series of water quality parameters in the target area, and inputs this time series into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration in the target area, thereby obtaining the predicted dissolved oxygen concentration result for the target area. By utilizing daily lag time series of dissolved oxygen concentration and water quality parameters, model construction and training are carried out to capture and express the complex nonlinear relationship between dissolved oxygen and other water quality parameters in the aquatic environment, thereby improving the accuracy and efficiency of dissolved oxygen concentration prediction.

[0078] Please refer to Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. The computer device 8 includes: a processor 81, a memory 82, and a computer program 83 stored in the memory 82 and executable on the processor 81; the computer device can store multiple instructions, which are adapted to be loaded and executed by the processor 81. Figures 1 to 6 The method steps of the illustrated embodiment can be found in the following documentation for detailed execution. Figures 1 to 6 The specific details of the illustrated embodiments will not be elaborated here.

[0079] The processor 81 may include one or more processing cores. The processor 81 connects to various parts of the server using various interfaces and lines. It executes various functions and processes data of the machine learning-based dissolved oxygen concentration prediction device 7 by running or executing instructions, programs, code sets, or instruction sets stored in memory 82, and by accessing data in memory 82. Optionally, the processor 81 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 81 may integrate one or more of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 81.

[0080] The memory 82 may include random access memory (RAM) or read-only memory. Optionally, the memory 82 may include a non-transitory computer-readable storage medium. The memory 82 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 82 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch instructions), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. Optionally, the memory 82 may also be at least one storage device located remotely from the aforementioned processor 81.

[0081] This application also provides a storage medium that can store multiple instructions. These instructions are applicable to being loaded and executed by a processor using the method steps described in Embodiments 1 to 4 above. For details of the execution process, please refer to the specific descriptions of Embodiments 1 to 4, which will not be repeated here.

[0082] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. In the embodiments, each functional unit and module can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. In addition, the specific names of each functional unit and module are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0083] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0084] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the algorithm. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0085] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0086] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0087] Furthermore, in the various embodiments of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0088] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms.

[0089] This invention is not limited to the above-described embodiments. If any modifications or variations to this invention do not depart from the spirit and scope of this invention, and if such modifications and variations fall within the scope of the claims and equivalent technologies of this invention, then this invention also intends to include such modifications and variations.

Claims

1. A machine learning-based method for predicting dissolved oxygen concentration, characterized in that, Includes the following steps: The daily lag time series of dissolved oxygen concentration and the daily lag time series of water quality parameters in the sample area are obtained. The daily lag time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days. The daily lag time series of water quality parameters includes water quality parameter data for several consecutive days. Based on the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters, several training sets are constructed by combining and dividing the training sets; machine learning methods are used to construct and train the model based on the several training sets to obtain a dissolved oxygen concentration prediction model. Obtain the daily lag characteristic time series of water quality parameters in the target area, input the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration, and obtain the dissolved oxygen concentration prediction result of the target area.

2. The dissolved oxygen concentration prediction method based on machine learning according to claim 1, characterized in that: The dissolved oxygen concentration prediction model includes a first dissolved oxygen concentration prediction model and a second dissolved oxygen concentration prediction model. The method employs machine learning to build and train a model based on several training sets to obtain a dissolved oxygen concentration prediction model, including the following steps: The standard deviation of dissolved oxygen concentration data for several days in each training set is calculated to obtain the standard deviation of dissolved oxygen concentration for each training set; based on the standard deviation of dissolved oxygen concentration in each training set and the preset dissolved oxygen concentration threshold, the training sets are classified and divided to obtain several first training sets and several second training sets. Several sets of the first training set are input into a preset first machine learning model for model training to obtain a first dissolved oxygen concentration prediction model. A second machine learning model is constructed based on several sets of the second training set; the second set of the second training set is input into the first machine learning model for model training to obtain a second dissolved oxygen concentration prediction model.

3. The dissolved oxygen concentration prediction method based on machine learning according to claim 2, characterized in that: The first machine learning model includes several XGBoost learners; The step of inputting a plurality of the first training sets into a preset first machine learning model for model training to obtain a first dissolved oxygen concentration prediction model includes the following steps: Sampling with replacement is performed on several of the first training sets to generate several first samples; the first samples are used as the input datasets for the current iteration, and each input dataset is input into each XGBoost learner for the current iteration to predict dissolved oxygen concentration, thereby obtaining the dissolved oxygen concentration prediction data for each input dataset in the current iteration, wherein the first sample includes several water quality parameter data. Based on the input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the current iteration, and the dissolved oxygen concentration prediction data of the corresponding input datasets of the previous iteration, train each XGBoost learner of the current iteration to obtain each XGBoost learner of the next iteration. The first sample is used as the input dataset for the next iteration. Each input dataset is input into each XGBoost learner in the current iteration. Dissolved oxygen concentration prediction and model training are repeated to obtain the first dissolved oxygen concentration prediction model.

4. The dissolved oxygen concentration prediction method based on machine learning according to claim 3, characterized in that, The step of training each XGBoost learner in the current iteration based on the input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the current iteration, and the dissolved oxygen concentration prediction data of the corresponding input datasets of the previous iteration, to obtain each XGBoost learner in the next iteration, includes the following steps: Based on the input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the current iteration, the dissolved oxygen concentration prediction data of the corresponding input datasets of the previous iteration, and a preset regularization loss function, the loss value of each XGBoost learner in the current iteration is obtained; based on the loss value, each XGBoost learner in the current iteration is trained to obtain each XGBoost learner in the next iteration, wherein the regularization loss function is: In the formula, For the first t The loss value of the XGBoost learner in the next iteration. The dissolved oxygen concentration prediction data is the corresponding input dataset for the current iteration. The dissolved oxygen concentration prediction data is from the corresponding input dataset of the previous iteration. For decision tree functions, Given the input dataset for the current iteration, This is a regularization term.

5. The dissolved oxygen concentration prediction method based on machine learning according to claim 4, characterized in that: The second machine learning model includes a basic learner layer and a meta-learning layer. The basic learner layer includes several sub-dissolved oxygen concentration prediction models; the meta-learning layer includes a random forest model. The step of constructing a second machine learning model based on several second training sets, using the water quality parameter data as independent variables and the dissolved oxygen concentration data as dependent variables, and then inputting the several second training sets into the first machine learning model for model training to obtain a second dissolved oxygen concentration prediction model includes the following steps: Sampling with replacement is performed on several second training sets to generate several second samples; a model is constructed based on several second samples to obtain several sub-dissolved oxygen concentration prediction models, wherein the second samples include several water quality parameter data; Several second training sets are respectively input into several sub-dissolved oxygen concentration prediction models for model training, to obtain several sub-dissolved oxygen concentration prediction models after training and the dissolved oxygen concentration prediction data corresponding to each second sample output by several sub-dissolved oxygen concentration prediction models after training; The dissolved oxygen concentration prediction data corresponding to each second sample output by each sub-dissolved oxygen concentration prediction model after training are transformed into prediction vectors to obtain the dissolved oxygen concentration prediction vectors corresponding to each second sample; the dissolved oxygen concentration prediction vectors corresponding to the same second sample are combined to obtain the training feature matrix of each second sample. Obtain dissolved oxygen concentration label data corresponding to each of the second samples; construct and train a random forest model based on the training feature matrix of each of the second samples and the dissolved oxygen concentration label data to obtain the trained random forest model.

6. The dissolved oxygen concentration prediction method based on machine learning according to claim 5, characterized in that: The sub-dissolved oxygen concentration prediction model is a machine learning model constructed using the water quality parameter data as independent variables and the dissolved oxygen concentration data as dependent variables, including decision tree model, extremely random tree model, gradient boosting model and extreme gradient boosting model; Both the decision tree model and the extremely random tree model use minimizing the mean squared error function as the objective function, wherein the mean squared error function is: In the formula, To minimize the mean square error, The number of elements in the second training set. This is the dissolved oxygen concentration prediction data corresponding to the second training set. The dissolved oxygen concentration label data is the data obtained from the daily lag feature time series of dissolved oxygen concentration corresponding to the second training set. The gradient boosting model uses the mean squared error function as the objective function, and the extreme gradient boosting model uses the regularized loss function as the objective function. The mean squared error function is: In the formula, This is the mean square error. This represents the dissolved oxygen concentration prediction data output by the gradient boosting model in the m-th iteration. For learning rate, This is the data processing function for the gradient boosting model in the m-th iteration.

7. The dissolved oxygen concentration prediction method based on machine learning according to claim 6, characterized in that, The step of constructing and training a random forest model based on the training feature matrix of each of the second samples and the dissolved oxygen concentration label data includes the following steps: The training feature matrix of each second sample is sampled with replacement to generate several third samples, wherein the third samples include several candidate features, and the candidate features are dissolved oxygen concentration prediction vectors. For each of the third samples, several candidate features are randomly selected to construct the candidate feature set of the current node of the decision tree corresponding to each of the third samples; according to the gain maximization principle, the target feature is selected as the node for splitting based on the feature values ​​of several candidate features in the candidate feature set of the current node; when the preset recursive splitting condition is met, the splitting stops, and each decision tree is constructed. The decision trees are combined to construct the random forest model. The training feature matrix of each second sample and the dissolved oxygen concentration label data are input into the random forest model for training to obtain the trained random forest model.

8. A dissolved oxygen concentration prediction device based on machine learning, characterized in that, include: The data acquisition module is used to acquire the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters in the sample area. The daily lag characteristic time series of dissolved oxygen concentration includes dissolved oxygen concentration data for several consecutive days; the daily lag characteristic time series of water quality parameters includes water quality parameter data for several consecutive days. The model training module is used to combine and divide the training set according to the daily lag characteristic time series of dissolved oxygen concentration and the daily lag characteristic time series of water quality parameters to construct several training sets; and to use machine learning methods to construct and train the model based on the several training sets to obtain a dissolved oxygen concentration prediction model. The dissolved oxygen concentration prediction module is used to obtain the daily lag characteristic time series of water quality parameters in the target area, input the daily lag characteristic time series of water quality parameters in the target area into the dissolved oxygen concentration prediction model to predict the dissolved oxygen concentration and obtain the dissolved oxygen concentration prediction result of the target area.

9. A computer device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the machine learning-based dissolved oxygen concentration prediction method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a computer program that, when executed by a processor, implements the steps of the machine learning-based dissolved oxygen concentration prediction method as described in any one of claims 1 to 7.