Watershed pollution prediction method based on multi-source data fusion

By constructing a multi-source data sequence and random forest regression model, the problem of inaccurate prediction caused by single data in traditional river basin pollution monitoring is solved, and accurate prediction and stable prevention and control of river basin pollution are achieved.

CN120296576APending Publication Date: 2025-07-11HYDRAULIC SCI RES INST OF SICHUAN PROVINCE
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510354488.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional river basin pollution monitoring relies on single data, resulting in inaccurate predictions and lack of targeted prevention and control measures, and is unable to adapt to dynamic changes in the basin environment.

Method used

Water quality, runoff, land use and meteorological data sequences are constructed, features are extracted and spliced into fusion feature vectors, random forest regression models are used for training and verification, and models that meet the conditions are selected for prediction.

Benefits of technology

It provides comprehensive and rich information, improves the accuracy and stability of river basin pollution prediction, reduces prediction errors, and ensures the reliability of the model in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296576A_ABST
    Figure CN120296576A_ABST
Patent Text Reader

Abstract

The invention discloses a watershed pollution prediction method based on multi-source data fusion, and belongs to the technical field of watershed pollution prediction. According to the method, water quality, runoff, land utilization and meteorological data sequences are firstly constructed, and then features of the sequences are extracted and spliced to form a fusion feature vector. And forming a training set by using a plurality of fusion feature vectors, inputting the training set into a random forest regression model for training, performing performance evaluation on the trained model by using a verification set, and finally using the model meeting evaluation conditions to predict the drainage basin pollution degree. According to the invention, through multi-source data fusion and scientific model training and evaluation processes, the accuracy of drainage basin pollution prediction is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of watershed pollution prediction, and particularly relates to a method for predicting watershed pollution based on multi-source data fusion. Background Art

[0002] With the rapid development of industrialization and urbanization, the problem of watershed pollution has become increasingly serious. Traditional methods for monitoring and preventing watershed pollution often have problems such as single data, inaccurate prediction, and lack of pertinence in prevention and control measures. On the one hand, relying solely on single water quality monitoring data is difficult to comprehensively reflect the actual situation of watershed pollution, because watershed pollution is affected by multiple factors, such as upstream pollution source emissions, runoff changes, climate conditions, and surrounding land use patterns. On the other hand, most existing pollution prediction models use fixed weights to synthesize various factors and cannot adapt to the dynamic changes of the watershed environment, resulting in a large deviation between the prediction results and the actual situation. This not only affects the accurate judgment of the watershed pollution situation but also makes it difficult for pollution prevention and control measures to achieve the expected effect and cannot effectively protect the watershed ecological environment. Summary of the Invention

[0003] Aiming at the above deficiencies in the prior art, a method for predicting watershed pollution based on multi-source data fusion provided by the present invention solves the problems of single monitoring data and inaccurate prediction in the prior art for watershed pollution.

[0004] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: A method for predicting watershed pollution based on multi-source data fusion, comprising the following steps:

[0005] S1. Construct a water quality data sequence, a runoff data sequence, a land use data sequence, and a meteorological data sequence;

[0006] S2. Extract features from each sequence, splice the features of each sequence to obtain a fused feature vector;

[0007] S3. Use multiple fused feature vectors as a training set, input the training set into a random forest regression model to obtain a trained random forest regression model;

[0008] S4. Use multiple fused feature vectors as a validation set, and use the validation set to evaluate the performance of the trained random forest regression model;

[0009] S5. Use the random forest regression model that meets the evaluation conditions to predict the degree of watershed pollution.

[0010] Further, the S1 includes the following sub-steps:

[0011] S11. Collect historical and real-time water quality data of the upstream watershed and the watershed to be monitored, and construct a water quality data sequence;

[0012] S12. Collect the historical and real-time runoff data of the basin, and construct a runoff data sequence;

[0013] S13. Obtain the land use type data around the basin, and construct a land use data sequence;

[0014] S14. Collect the meteorological data of the basin, and construct a meteorological data sequence.

[0015] Furthermore, the water quality data in S11 includes: Chemical Oxygen Demand (COD), Ammonia Nitrogen (NH3-N), and Total Phosphorus (TP);

[0016] The land use type data in S13 includes: cultivated land, construction land, and forest land;

[0017] The meteorological data in S14 includes: water, air temperature, and wind speed.

[0018] Furthermore, S2 includes the following sub-steps:

[0019] S21. Use the principal component analysis method to extract the principal components from the water quality data sequence, and obtain the water quality principal component features;

[0020] S22. Extract the mean, standard deviation, and change trend slope of the runoff data sequence, and construct the runoff features;

[0021] S23. Use one-hot encoding to convert the land use data sequence into a numerical vector, and perform principal component analysis on the numerical vector to obtain the land principal component features;

[0022] S24. Perform Fourier transform on the meteorological data sequence to extract its periodic frequency, and obtain the meteorological features;

[0023] S25. Concatenate the water quality principal component features, runoff features, land principal component features, and meteorological features to obtain a fused feature vector.

[0024] Furthermore, S3 is specifically: Use the fused feature vector of the training set as the input and the pollution degree of the polluted basin as the output, and input them into the random forest regression model. During the training process, adjust the parameters of the random forest.

[0025] Furthermore, the parameters of the random forest include: the number of decision trees, the maximum depth.

[0026] Furthermore, S4 is specifically: Calculate the error index between the predicted basin pollution degree and the real basin pollution degree of the random forest regression model on the validation set. When the error index is less than the error threshold, the verification of the random forest regression model is completed.

[0027] Furthermore, the calculation formula of the error index is: Or Among them, MSE is the error index, y i is the predicted pollution degree of the basin, is the actual pollution degree of the basin, N is the number of samples in the validation set, and i is a positive integer.

[0028] In summary, the beneficial effects of the present invention are as follows:

[0029] 1. By constructing water quality, runoff, land use, and meteorological data sequences, the present invention breaks through the limitations of traditional single-data monitoring. Different types of data reflect the basin environmental conditions from multiple dimensions. For example, water quality data directly presents the content of pollutants, runoff affects the diffusion of pollutants, land use patterns determine the distribution of pollution sources, and meteorological conditions act on the migration and transformation of pollutants. The fusion of multi-source data provides comprehensive and rich information, laying a solid foundation for accurate prediction.

[0030] 2. The present invention extracts features from each data sequence and splices them into a fused feature vector, fully mining the key information and potential relationships behind different data. The random forest regression model can effectively handle complex non-linear relationships, learn and train the fused feature vector. Compared with relying on single data or simple models, it can more accurately capture the complex mapping between basin pollution and multiple factors, greatly improving the prediction accuracy, reducing the prediction error, and making the prediction results more consistent with the actual pollution situation.

[0031] 3. The present invention uses the validation set to evaluate the performance of the trained random forest regression model, testing the performance of the model on data not involved in training from different perspectives. By screening through evaluation indicators, a model that meets the conditions is selected to ensure that the model has good generalization ability, avoid overfitting, enhance the reliability and stability of the model in practical applications, and provide continuous and accurate support for basin pollution prediction. Description of the Drawings

[0032] Figure 1 is a flowchart of a basin pollution prediction method based on multi-source data fusion. Detailed Embodiments

[0033] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0034] As Figure 1 shown, a basin pollution prediction method based on multi-source data fusion includes the following steps:

[0035] S1. Construct water quality data sequences, runoff data sequences, land use data sequences, and meteorological data sequences;

[0036] S2. Extract features from each sequence, splice the features of each sequence to obtain a fused feature vector;

[0037] S3. Use multiple fused feature vectors as the training set, input the training set into the random forest regression model to obtain the trained random forest regression model;

[0038] S4. Use multiple fused feature vectors as the validation set, and use the validation set to evaluate the performance of the trained random forest regression model;

[0039] S5. Use the random forest regression model that meets the evaluation conditions to predict the pollution degree of the basin.

[0040] In this embodiment, S1 includes the following sub-steps:

[0041] S11. Collect historical and real-time water quality data of the upstream basin and the basin to be monitored, and construct a water quality data sequence;

[0042] S12. Collect historical and real-time runoff data of the basin, and construct a runoff data sequence;

[0043] S13. Obtain land use type data around the basin, and construct a land use data sequence;

[0044] S14. Collect meteorological data of the basin, and construct a meteorological data sequence.

[0045] In this embodiment, the water quality data in S11 includes: chemical oxygen demand COD, ammonia nitrogen NH3-N, and total phosphorus TP;

[0046] The land use type data in S13 includes: cultivated land, construction land, and forest land;

[0047] The meteorological data in S14 includes: water, air temperature, and wind speed.

[0048] In this embodiment, S2 includes the following sub-steps:

[0049] S21. Use the principal component analysis method to extract the principal components of the water quality data sequence to obtain the water quality principal component features;

[0050] S22. Extract the mean, standard deviation, and change trend slope of the runoff data sequence to construct runoff features;

[0051] S23. Use one-hot encoding to convert the land use data sequence into a numerical vector, and perform principal component analysis on the numerical vector to obtain the land principal component features;

[0052] S24. Perform a Fourier transform on the meteorological data sequence to extract its periodic frequency and obtain meteorological features;

[0053] S25. Concatenate the water quality principal component features, runoff features, land principal component features, and meteorological features to obtain a fused feature vector.

[0054] In this embodiment, the specific operation of S3 is as follows: Use the fused feature vector of the training set as the input and the pollution degree of the polluted basin as the output, and input them into the random forest regression model. During the training process, adjust the parameters of the random forest.

[0055] In this embodiment, the parameters of the random forest include: the number of decision trees and the maximum depth.

[0056] In this embodiment, the specific operation of S4 is as follows: Calculate the error index between the predicted basin pollution degree and the true basin pollution degree of the random forest regression model on the validation set. When the error index is less than the error threshold, the verification of the random forest regression model is completed.

[0057] In this embodiment, the calculation formula of the error index is: Or where MSE is the error index, y i is the predicted basin pollution degree, is the true basin pollution degree, N is the number of samples in the validation set, and i is a positive integer.

[0058] By constructing water quality, runoff, land use, and meteorological data sequences, the present invention breaks through the limitations of traditional single-data monitoring. Different types of data reflect the basin environmental conditions from multiple dimensions. For example, water quality data directly presents the content of pollutants, runoff affects the diffusion of pollutants, land use patterns determine the distribution of pollution sources, and meteorological conditions act on the migration and transformation of pollutants. The multi-source data fusion provides comprehensive and rich information, laying a solid foundation for accurate prediction.

[0059] The present invention extracts features from each data sequence and concatenates them into a fused feature vector, fully exploring the key information and potential connections behind different data. The random forest regression model can effectively handle complex non-linear relationships, learn and train the fused feature vector. Compared with relying on single data or simple models, it can more accurately capture the complex mapping between basin pollution and multiple factors, greatly improving the prediction accuracy, reducing the prediction error, and making the prediction results more consistent with the actual pollution situation.

[0060] The present invention uses a validation set to evaluate the performance of the trained random forest regression model and examines the performance of the model on data not involved in training from different perspectives. By screening the models through evaluation metrics, models that meet the conditions are selected to ensure that the models have good generalization ability, avoid overfitting, enhance the reliability and stability of the models in practical applications, and provide continuous and accurate support for watershed pollution prediction.

[0061] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting basin pollution based on multi-source data fusion, characterized in that, It includes the following steps: S1. Construct water quality data sequences, runoff data sequences, land use data sequences and meteorological data sequences; S2. Extract features from each sequence, splice the features of each sequence to obtain a fused feature vector; S3. Use multiple fused feature vectors as a training set, input the training set into a random forest regression model to obtain a trained random forest regression model; S4. Use multiple fused feature vectors as a validation set, and use the validation set to evaluate the performance of the trained random forest regression model; S5. Use the random forest regression model that meets the evaluation conditions to predict the pollution degree of the watershed.

2. The method for predicting basin pollution based on multi-source data fusion according to claim 1, wherein The S1 includes the following sub-steps: S11. Collect historical and real-time water quality data of the upstream watershed and the watershed to be monitored, and construct a water quality data sequence; S12. Collect historical and real-time runoff data of the watershed, and construct a runoff data sequence; S13. Obtain land use type data around the watershed, and construct a land use data sequence; S14. Collect meteorological data of the watershed, and construct a meteorological data sequence.

3. The method for predicting basin pollution based on multi-source data fusion according to claim 2, wherein, The water quality data in S11 includes: chemical oxygen demand (COD), ammonia nitrogen (NH3-N) and total phosphorus (TP); The land use type data in S13 includes: cultivated land, construction land and forest land; The meteorological data in S14 includes: water, air temperature and wind speed.

4. The multi-source data fusion-based watershed pollution prediction method according to claim 1, wherein, The S2 includes the following sub-steps: S21. Use the principal component analysis method to extract the principal components of the water quality data sequence to obtain water quality principal component features; S22. Extract the mean, standard deviation and change trend slope of the runoff data sequence to construct runoff features; S23. Use one-hot encoding to convert the land use data sequence into a numerical vector, and perform principal component analysis on the numerical vector to obtain land principal component features; S24. Perform Fourier transform on the meteorological data sequence to extract its periodic frequency to obtain meteorological features; S25. Splice the water quality principal component features, runoff features, land principal component features and meteorological features to obtain a fused feature vector.

5. The method for predicting basin pollution based on multi-source data fusion according to claim 1, wherein The S3 is specifically: use the fused feature vector of the training set as the input, and the pollution degree of the polluted watershed as the output, and input it into the random forest regression model. During the training process, adjust the parameters of the random forest.

6. The multi-source data fusion-based watershed pollution prediction method according to claim 5, wherein The parameters of the random forest include: the number of decision trees and the maximum depth.

7. The method for predicting basin pollution based on multi-source data fusion according to claim 1, wherein The S4 is specifically: calculate the error index between the predicted watershed pollution degree and the real watershed pollution degree of the random forest regression model on the validation set. When the error index is less than the error threshold, the verification of the random forest regression model is completed.

8. The multi-source data fusion based watershed pollution prediction method according to claim 7, characterized in that, The calculation formula for the error index is as follows: or where MSE is the error index, y i is the predicted pollution level of the basin, is the actual pollution level of the basin, N is the number of samples in the validation set, and i is a positive integer.

Citation Information

Cited By

  • Soil heavy metal pollution cause accurate identification method and system

    CN120763552A

  • A method and system for accurately identifying the cause of heavy metal pollution in soil

    CN120763552B

  • Pollution source groundwater environment quality dynamic evaluation method and device based on multi-source data fusion

    CN120893904A

  • Basin scale agricultural non-point source pollution simulation method

    CN120995944A

  • A method for simulating agricultural non-point source pollution at a watershed scale

    CN120995944B