Intelligent monitoring system for pollution of river course around rice planting area based on remote sensing information
Patent Information
- Application Number
- CN202610956544.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
但是,水稻种植区周边河道的遥感数据容易受到背景环境的动态变化影响,例如季节变化、灌溉和排水等农业活动、降雨和气温等气候条件变化,也会导致河道遥感数据发生变化
接着,获取各监测点各时刻的平均无污染可能性,将一个监测点一个时刻的各种遥感监测数据组成该监测点该时刻的数据点,利用平均无污染可能性对各监测点各时刻的数据点进行筛选得到数据点集合,对数据点集合中的数据点进行聚类获取各数据点聚簇;对一个数据点聚簇中的时刻聚类获取该数据点聚簇的各时刻聚簇,根据每个数据点聚簇的各时刻聚簇判断当前时刻对应的各数据点聚簇,这里通过遥感监测数据的周期性表现得到当前时刻对应的各数据点聚簇,进而最后利用当前时刻对应的各数据点聚簇中各种遥感监测数据对各监测点当前时刻的各种遥感监测数据进行修正,以降低背景环境对遥感监测数据的影响,提高河道污染监测结果的准确性。
Smart Images

Figure CN122799271A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of river pollution monitoring technology, specifically to a smart monitoring system for river pollution around rice-growing areas based on remote sensing information. Background Technology
[0002] Rice-growing areas typically employ flooded cultivation, resulting in close water connections between paddy fields and surrounding waterways. Pollution in these waterways can lead to contaminants entering the paddy fields through irrigation systems, impacting crop growth. Furthermore, nutrients and pesticides used in agricultural activities (such as fertilization and pesticide application) can enter waterways through surface runoff or drainage ditches, causing pollution. Therefore, monitoring and timely remediation of pollution in waterways surrounding rice-growing areas are crucial for ensuring rice production quality and maintaining aquatic ecosystem security.
[0003] Remote sensing technology is playing an increasingly important role in river pollution monitoring to achieve large-scale automated monitoring. Existing remote sensing-based river pollution monitoring methods mainly analyze river turbidity and chlorophyll concentration by monitoring spectral reflectance, and determine pollution levels by setting anomaly thresholds. However, remote sensing data of rivers surrounding rice-growing areas are easily affected by dynamic changes in the background environment, such as seasonal changes, agricultural activities like irrigation and drainage, and climatic conditions like rainfall and temperature, all of which can cause changes in the remote sensing data. Existing monitoring methods often struggle to distinguish between remote sensing data under background environmental conditions and those affected by river pollution, leading to insufficient accuracy in pollution monitoring results. Summary of the Invention
[0004] To address the aforementioned technical problems, the present invention aims to provide a smart monitoring system for river pollution around rice-growing areas based on remote sensing information. The specific technical solution adopted is as follows: One embodiment of the present invention provides a smart monitoring system for river pollution around rice-growing areas based on remote sensing information. The system includes: The data preprocessing module is used to acquire various remote sensing monitoring data at each monitoring point at each time, and to record any type of remote sensing monitoring data as the target monitoring data; and to acquire the change in the target monitoring data at a monitoring point at a given time. The importance acquisition module is used to cluster the target monitoring data of each monitoring point in history to obtain each cluster; the change in the target monitoring data in a cluster is used as the x-axis, and the change in other remote sensing monitoring data is used as the y-axis to obtain the coordinate system corresponding to the other remote sensing monitoring data of the cluster; the importance of the target monitoring data is obtained by using the coordinate system corresponding to the other remote sensing monitoring data of each cluster. The data filtering module is used to obtain the probability of no pollution of the target monitoring data at each monitoring point at each time based on the coordinate system corresponding to various other remote sensing monitoring data of each cluster; to take the average probability of no pollution of various remote sensing monitoring data at a monitoring point at a time as the average probability of no pollution of that monitoring point at that time; to form the data points of that monitoring point at that time by combining various remote sensing monitoring data at a monitoring point at a time; and to filter the data points of each monitoring point at each time using the average probability of no pollution to obtain a set of data points. The data correction module is used to cluster data points in the data point set according to the importance of each type of remote sensing monitoring data to obtain data point clusters; to cluster time points within a data point cluster to obtain time point clusters for that data point cluster; to determine the data point clusters corresponding to the current time based on the time point clusters of each data point cluster; to correct the various remote sensing monitoring data of each monitoring point at the current time using the various remote sensing monitoring data in the data point clusters corresponding to the current time; and to monitor river pollution using the corrected various remote sensing monitoring data of each monitoring point at the current time.
[0005] Preferably, obtaining the change in target monitoring data at a monitoring point at a given time includes: The difference between the target monitoring data of a monitoring point at a certain time and the target monitoring data of the same monitoring point at the previous time is taken as the change in the target monitoring data of that monitoring point at that time.
[0006] Preferably, the coordinate system corresponding to the various remote sensing monitoring data of the cluster is obtained by using the change in target monitoring data within a cluster as the x-axis and the changes in other types of remote sensing monitoring data as the y-axis, including: In a cluster, the target monitoring data in the cluster are arranged in ascending order as the x-axis of a coordinate system. Another type of remote sensing monitoring data is denoted as the remote sensing monitoring data to be constructed. Among all the remote sensing monitoring data to be constructed, the remote sensing monitoring data to be constructed that have the same monitoring point and time as each target monitoring data in the cluster are found. The monitoring point and time are then matched with each target monitoring data on the x-axis as the y-axis to obtain the coordinate system corresponding to the remote sensing monitoring data to be constructed in the cluster. Similarly, the coordinate systems corresponding to other remote sensing monitoring data in the cluster are obtained.
[0007] Preferably, the importance of the target monitoring data is obtained by using the coordinate system corresponding to various other remote sensing monitoring data of each cluster, including: Within a cluster, the Freedman-Diaconis rule is used to obtain the interval length of the x-coordinate in the coordinate system, and this interval length is used to divide the x-coordinate into various intervals. The standard deviation of the y-coordinates corresponding to each x-coordinate within an interval of the coordinate system corresponding to a certain type of remote sensing monitoring data is calculated and denoted as the y-coordinate standard deviation. A negative correlation mapping is performed using an exponential function with a base of the natural constant on the mean of the y-coordinate standard deviations corresponding to each interval of the coordinate system corresponding to this type of remote sensing monitoring data to obtain the correlation between the cluster's target monitoring data and this type of remote sensing monitoring data. Finally, the correlation between the cluster's target monitoring data and other types of remote sensing monitoring data is calculated. The mean of the correlation is used as the overall correlation of the cluster target monitoring data; the mean of the target monitoring data of each monitoring point at the current time is obtained and recorded as the average target monitoring data at the current time; the distance between the average target monitoring data at the current time and the cluster center of a cluster is added to the hyperparameter and the reciprocal is calculated to obtain the reciprocal distance of the cluster; the ratio of the reciprocal distance of a cluster to the sum of the reciprocals of the distances of all clusters is calculated to obtain the influence weight of the cluster target monitoring data; the importance of the target monitoring data is obtained by weighting and summing the overall correlation of the cluster target monitoring data using the influence weights of the cluster target monitoring data.
[0008] Preferably, the probability of no contamination of the target monitoring data at each monitoring point at each time is obtained based on the coordinate system corresponding to various other remote sensing monitoring data of each cluster, including: In a coordinate system, a point is formed by combining an abscissa and its corresponding ordinate. A point is formed by combining target monitoring data from a monitoring point at a given time with another type of remote sensing data; this point is denoted as the analysis point. Within the interval of the coordinate system containing the analysis point, the average distance between the analysis point and other points within the interval is calculated as the average distance of the analysis point at that time. A negative correlation mapping is performed using an exponential function with a base of the natural constant to the ratio of the average distance of the analysis point to the average distance of all points within the interval of the coordinate system containing the analysis point, yielding the relative distribution density of the analysis point at that time. The mean of the relative distribution density of the points formed by the target monitoring data at that time and other remote sensing data is taken as the probability of the target monitoring data at that time being uncontaminated.
[0009] Preferably, the data points at each monitoring point at each time point are filtered using the average probability of no pollution to obtain a set of data points, including: The average probability of no pollution at each monitoring point at each time point is processed using a box plot to obtain the lower limit of the normal range of the average probability of no pollution, which is used as the screening threshold. Data points at each monitoring point at each time point whose average probability of no pollution is less than the screening threshold are removed, and the remaining data points are retained to form a data point set.
[0010] Preferably, the data points in the data point set are clustered according to the importance of each type of remote sensing monitoring data to obtain data point clusters, including: The importance of various remote sensing monitoring data is used as the weight of each type of remote sensing monitoring data. Combined with the Euclidean distance calculation model, the weighted Euclidean distance between two data points is calculated. Based on the weighted Euclidean distance between every two data points in the data point set, the data points in the data point set are clustered to obtain each data point cluster.
[0011] Preferably, determining the cluster of data points at the current time based on the clusters of each data point cluster at each time step includes: For a data point cluster, the time interval between the earliest and latest times of a time cluster within that data point cluster is taken as the time interval corresponding to that time cluster. The time intervals corresponding to each time cluster of the data point cluster are arranged in chronological order, and the average time interval between any two time intervals is taken as the average time interval of the data point cluster. The average duration of the time intervals corresponding to each time cluster of the data point cluster is calculated as the average duration of the data point cluster. Starting from the latest time in the data point cluster, a window with a duration equal to the average duration is constructed. The sliding step is the sum of the average time interval and the average duration. The window slides from the starting time. If the current time is within the window, then the data point cluster is the data point cluster corresponding to the current time.
[0012] Preferably, the remote sensing monitoring data of each monitoring point at the current time is corrected using various remote sensing monitoring data in the clusters of data points corresponding to the current time, including: The data point clusters corresponding to the current time are denoted as the data point cluster to be analyzed. The mean of a type of remote sensing monitoring data within a data point cluster to be analyzed is obtained as the standard value for that type of remote sensing monitoring data in that data point cluster. The time periods corresponding to each time cluster of the data point cluster to be analyzed are arranged in chronological order, and the standard deviation of the time interval between every two time periods is obtained, denoted as the time interval standard deviation of the data point cluster to be analyzed. The standard deviation of the time duration corresponding to each time cluster of the data point cluster to be analyzed is calculated, denoted as the time duration standard deviation of the data point cluster to be analyzed. The normalized value of the time interval standard deviation of the data point cluster to be analyzed, the reciprocal of the hyperparameter, and the time duration standard deviation of the data point cluster to be analyzed are then used as the standard deviation of the time duration. The existence probability of a cluster of data points to be analyzed is obtained by adding the difference to the normalized value of the reciprocal of the hyperparameter; the existence probability of a cluster of data points to be analyzed is obtained by comparing it with the sum of the existence probabilities of all clusters of data points to be analyzed; the standard values of the target monitoring data of each cluster of data points to be analyzed are weighted and summed using the weights of each cluster of data points to be analyzed to obtain the adjustment parameter of the target monitoring data at the current time; the difference between the target monitoring data of a monitoring point at the current time and the adjustment parameter of the target monitoring data at the current time is obtained to obtain the corrected target monitoring data of the monitoring point at the current time, and similarly, the corrected remote sensing monitoring data of each monitoring point at the current time are obtained.
[0013] The embodiments of the present invention have at least the following beneficial effects: This application obtains various remote sensing monitoring data of each monitoring point at each time, and then obtains the change in the target monitoring data of a monitoring point at a time. Then, it clusters the target monitoring data of each monitoring point in history to obtain each cluster, and then obtains the importance of the target monitoring data based on the cluster. Similarly, it obtains the importance of various remote sensing monitoring data. That is, based on the correlation performance of each type of remote sensing monitoring data with other types of remote sensing monitoring data, it adjusts the influence weight of each type of remote sensing monitoring data on the identification of historical data of similar background environmental conditions, so as to minimize the interference of pollution and environmental conditions on the judgment of the real background environment and improve the accuracy of the identification of historical data of similar background environmental conditions. Next, the average probability of no pollution at each monitoring point at each time is obtained. Various remote sensing data from a monitoring point at a given time are combined to form a data point set for that monitoring point at that time. The average probability of no pollution is used to filter the data points at each monitoring point at each time to obtain a data point set. The data points in the data point set are then clustered to obtain data point clusters. For each data point cluster, time clusters are obtained. Based on the time clusters of each data point cluster, the data point clusters corresponding to the current time are determined. Here, the periodicity of the remote sensing data is used to obtain the data point clusters corresponding to the current time. Finally, the various remote sensing data in the data point clusters corresponding to the current time are used to correct the various remote sensing data of each monitoring point at the current time, in order to reduce the impact of the background environment on the remote sensing data and improve the accuracy of river pollution monitoring results. Attached Figure Description
[0014] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a system block diagram of a smart monitoring system for river pollution around rice-growing areas based on remote sensing information, provided as an embodiment of the present invention. Detailed Implementation
[0016] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a smart monitoring system for river pollution around rice-growing areas based on remote sensing information proposed in accordance with the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0018] The following description, in conjunction with the accompanying drawings, details the specific scheme of the intelligent monitoring system for river pollution around rice-growing areas provided by this invention. Example
[0019] The main application scenario of this invention is as follows: Rivers surrounding rice-growing areas are easily affected by factors such as rice paddy irrigation and drainage, seasonal changes, and weather variations, leading to fluctuations in remote sensing monitoring data. These data changes are caused by changes in background environmental conditions, rather than river pollution. Traditional monitoring methods only determine the presence of pollution by setting thresholds for the monitoring data, which cannot effectively distinguish between background environmental changes and actual pollution, resulting in insufficient accuracy of monitoring results. Therefore, this application corrects the monitoring data to obtain more accurate monitoring data, thereby enabling the monitoring of river pollution.
[0020] Please see Figure 1 The diagram illustrates a system block diagram of a smart monitoring system for river pollution around rice-growing areas based on remote sensing information, provided by an embodiment of the present invention. The system includes the following modules: The data preprocessing module is used to acquire various remote sensing monitoring data at each monitoring point at each time, and to record any type of remote sensing monitoring data as the target monitoring data; and to acquire the change in the target monitoring data at a monitoring point at a given time.
[0021] Because single-type remote sensing data has strong limitations, such as spectral data being unable to identify changes in river water levels, radar data being unable to determine specific pollution types, and thermal infrared data mainly reflecting thermal pollution, the remote sensing information collected in this application includes spectral data, radar data, and thermal infrared data.
[0022] Therefore, monitoring points were set up at 10m intervals in stable water bodies, which were defined manually and not covered by vegetation. Then, drones equipped with multispectral and thermal infrared cameras were used to collect multispectral and thermal infrared data for each monitoring point, respectively. Radar data for each monitoring point was collected using a radar satellite. The multispectral data consisted of surface reflectance data for each spectral band after atmospheric correction using an atmospheric radiative transfer model (Sen2Cor for Sentinel-2). The thermal infrared data consisted of temperature data converted from the collected raw radiance data (current technology, achievable through remote sensing processing software such as ENVI). The radar data consisted of the backscattering coefficient after processing the raw signal. The monitoring equipment collected remote sensing data weekly. To facilitate analysis and comparison of different data types, the collected data were first normalized. This allows us to obtain the surface reflectance data, temperature data, and backscattering coefficient at each monitoring point in a stable water body area at each time. For ease of explanation, the surface reflectance data, temperature data, and backscattering coefficient are collectively referred to as different types of remote sensing monitoring data, and any type of remote sensing monitoring data is recorded as the target monitoring data.
[0023] In order to identify the patterns of changes in remote sensing data caused by seasonal changes in the background environment and environmental changes such as different climates, various remote sensing monitoring data from the past three years were used as historical data for analysis.
[0024] Meanwhile, in order to analyze the temporal changes of various remote sensing monitoring data, it is necessary to preprocess the various remote sensing monitoring data to obtain the amount of data change; specifically, the difference between the target monitoring data of a monitoring point at a certain time and the target monitoring data of the previous time at the same monitoring point is taken as the amount of change of the target monitoring data of the monitoring point at that time.
[0025] The specific calculation model is as follows: , in, This represents the change in the type a remote sensing data (target monitoring data) at time i of the c-th monitoring point. This represents the a-th type of remote sensing data at the c-th monitoring point at the i-th time. This represents the a-th type of remote sensing data at the time preceding the i-th time at the c-th monitoring point. From this, we can obtain the change in each type of remote sensing data at each time for each monitoring point.
[0026] The importance acquisition module is used to cluster the target monitoring data of each monitoring point in history to obtain each cluster; the change in the target monitoring data in a cluster is used as the x-axis and the change in other remote sensing monitoring data is used as the y-axis to obtain the coordinate system corresponding to the other remote sensing monitoring data of the cluster; the importance of the target monitoring data is obtained by using the coordinate system corresponding to the other remote sensing monitoring data of each cluster.
[0027] Because different remote sensing monitoring data are affected by different factors, such as radar data mainly reflecting river water level and thermal infrared data possibly being affected by air temperature, the background environmental conditions reflected by different remote sensing monitoring data have different focuses and degrees of reflection. For example, water level changes not only affect radar data but also cause changes in spectral data. However, it is difficult to determine whether the cause is weather, water level change, or pollution based solely on multispectral data. Therefore, in order to obtain historical data under similar background environmental conditions, it is first necessary to analyze the importance of each type of remote sensing monitoring data in the background environmental condition similarity analysis.
[0028] If every change in the target monitoring data leads to significant changes in other types of remote sensing monitoring data, it indicates that the monitoring data is of higher importance. A single data change may be accidental, so it is necessary to analyze the behavior of multiple data changes together.
[0029] Because the actual meaning of the same change may be different when the data values of the target monitoring data are different, such as the inundation structure and flow velocity of the river channel when the water level is different, which will lead to different changes in other types of remote sensing monitoring data. Therefore, the k-means clustering method is first used to cluster the target monitoring data of all monitoring points at all times in history to obtain each cluster. The clustering condition is the difference between the values of each target monitoring data, where the k value is determined by the elbow method.
[0030] Furthermore, it is necessary to analyze the correlation between changes in target monitoring data and changes in other types of remote sensing monitoring data within each cluster; thereby, a coordinate system of "change in target monitoring data - change in other types of remote sensing monitoring data" is constructed. Using the change in target monitoring data within a cluster as the x-axis and the changes in various other types of remote sensing monitoring data as the y-axis, the coordinate systems corresponding to the other types of remote sensing monitoring data within that cluster are obtained.
[0031] Specifically, within a cluster, the target monitoring data within the cluster are arranged in ascending order, forming the x-axis of a coordinate system. Another type of remote sensing monitoring data is designated as the data to be constructed. Among all the data to be constructed, those with the same monitoring points and times as the target monitoring data within the cluster are identified. These monitoring points and times are then mapped to the target monitoring data on the x-axis, forming the y-axis, thus obtaining the coordinate system corresponding to the data to be constructed for that cluster. Similarly, the coordinate systems corresponding to all other types of remote sensing monitoring data within the cluster are obtained. Therefore, for a cluster, using the target monitoring data as the x-axis, the coordinate system corresponding to each other type of remote sensing monitoring data can be obtained.
[0032] Since historical remote sensing monitoring data may be affected by various factors, the change in the same target monitoring data may correspond to multiple values for the change in other types of remote sensing monitoring data (corresponding to different historical moments). If the vertical coordinates of the data points (changes in other types of remote sensing monitoring data) under the change in the same target monitoring data have a high degree of concentration, it indicates that the two types of remote sensing monitoring data have a strong correlation.
[0033] Therefore, the importance of the target monitoring data is obtained by utilizing the coordinate systems corresponding to various other remote sensing monitoring data within each cluster. Specifically, within a cluster, the interval length of the abscissa of the coordinate system is obtained using the Freedman-Diaconis rule, and the abscissa is divided into intervals using this interval length. The standard deviation of each ordinate corresponding to each abscissa within an interval of the coordinate system corresponding to a certain type of remote sensing monitoring data is calculated and denoted as the ordinate standard deviation. The mean of the ordinate standard deviations corresponding to each interval within the coordinate system corresponding to this type of remote sensing monitoring data is negatively correlated using an exponential function with a base of the natural constant to obtain the correlation between the cluster's target monitoring data and this type of remote sensing monitoring data. Finally, the importance of the cluster's target monitoring data is determined by comparing it with other various types of remote sensing monitoring data. The mean of the degree of change in correlation is used as the overall correlation of the cluster target monitoring data; the mean of the target monitoring data of each monitoring point at the current time is obtained and recorded as the average target monitoring data at the current time; the distance between the average target monitoring data at the current time and the cluster center of a cluster is added to the hyperparameter and the reciprocal is calculated to obtain the reciprocal distance of the cluster; the ratio of the reciprocal distance of a cluster to the sum of the reciprocals of the distances of all clusters is calculated to obtain the influence weight of the cluster target monitoring data; the importance of the target monitoring data is obtained by weighting and summing the overall correlation of the cluster target monitoring data using the influence weights of the cluster target monitoring data.
[0034] The calculation model for the correlation between changes in a cluster of target monitoring data and a type of remote sensing monitoring data is as follows: , in, This indicates the degree of correlation between the changes in the a-type remote sensing monitoring data (target monitoring data) and the b-type remote sensing monitoring data of the w-th cluster; The standard deviation of the ordinates corresponding to the x-coordinates in the u-th interval of the coordinate system composed of the a-type and b-type remote sensing data of the w-th cluster (that is, the coordinate system corresponding to the b-type remote sensing data of the w-th cluster) is also the standard deviation of the ordinates corresponding to the u-th interval. The smaller the standard deviation, the more concentrated the distribution of the ordinates, that is, the stronger the correlation between the two types of remote sensing data. This represents the number of intervals in the coordinate system corresponding to the b-th type of remote sensing data in the w-th cluster; exp is an exponential function with the natural constant e as the base, used for negative correlation mapping to reflect the logical relationship that the smaller the standard deviation, the greater the correlation. Then, within a cluster, the mean value of the correlation between the cluster's target monitoring data and other types of remote sensing monitoring data is calculated as the overall correlation of the cluster's target monitoring data. (The overall correlation of the monitoring data of the w-th cluster target).
[0035] The specific calculation model for the influence weight of clustered target monitoring data is as follows: , in, This represents the influence weight of the monitoring data of the w-th cluster target (the a-th type of remote sensing monitoring data); This represents the distance between the average target monitoring data at the current moment (the average of the a-th type of remote sensing monitoring data from all monitoring points at the current moment) and the cluster center of the w-th cluster. This is the reciprocal of the distance to the w-th cluster, indicating that the closer the distance, the more likely the current monitoring time is to be similar to the environmental state corresponding to the data in that cluster. This is a hyperparameter used to ensure that the denominator is not zero and the fraction is meaningful; This represents the total number of clusters, or the total number of clusters corresponding to the target monitoring data. Remote sensing data with a higher correlation to other types of remote sensing data has a stronger ability to represent background environmental conditions. However, the correlation between target monitoring data and other data at different numerical levels is not necessarily consistent. For example, changes in radar data at low water levels are mainly caused by rising water levels, which has a relatively large impact on spectral data. Conversely, changes in radar data at high water levels may be caused by changes in river surface roughness (such as suspended matter), which has a smaller impact on spectral data. Since the current state of river pollution is uncertain, it is impossible to directly determine which cluster the current average target monitoring data belongs to. Therefore, the influence weight of each cluster's target monitoring data is obtained based on the distance between the current average target monitoring data and the cluster center of each cluster.
[0036] The specific model for calculating the importance of target monitoring data is as follows: , in, This indicates the importance of the a-th type of remote sensing monitoring data (target monitoring data). This represents the influence weight of the monitoring data of the w-th cluster target (the a-th type of remote sensing monitoring data); This indicates the overall correlation of the monitoring data for the w-th cluster target; This represents the total number of clusters corresponding to the a-th type of remote sensing monitoring data. Here, a weighted sum is performed to obtain the importance of the a-th type of remote sensing monitoring data (target monitoring data), which is used as the importance of the target monitoring data in the background environmental condition similarity analysis.
[0037] Similarly, for other types of remote sensing monitoring data, the importance of each type of remote sensing monitoring data can be determined.
[0038] The data filtering module is used to obtain the probability of no pollution of the target monitoring data at each monitoring point at each time based on the coordinate system corresponding to various other remote sensing monitoring data of each cluster; to take the average probability of no pollution of various remote sensing monitoring data at a monitoring point at a time as the average probability of no pollution of that monitoring point at that time; to form the data points of that monitoring point at that time by combining various remote sensing monitoring data at a monitoring point at a time; and to filter the data points of each monitoring point at each time using the average probability of no pollution to obtain a set of data points.
[0039] Irrigation and drainage in rice-growing areas, fluctuations in river levels, and rainfall can all alter river turbidity, suspended solids concentration, and reflectivity. Therefore, remote sensing data from rivers under different background environmental conditions exhibit significant differences. Although the background environment is complex, its influence on remote sensing data follows certain patterns, such as seasonal variations. Since agricultural activities like irrigation and drainage, as well as weather changes like rainfall, are not strictly periodic, to obtain more accurate patterns in remote sensing data under similar background environmental conditions, it is necessary to identify historical data segments with similar background environmental conditions based on the characteristics of historical data and the relative importance of each type of remote sensing data.
[0040] Data from similar background environments under unpolluted conditions exhibit obvious similarities. Historical remote sensing monitoring data may include data from polluted river conditions. Therefore, before analyzing data patterns under normal background conditions, it is necessary to eliminate interference from data from polluted river conditions.
[0041] Based on the distance between the point corresponding to the change in each type of remote sensing data at each time point and other points within the corresponding interval of the coordinate system, we analyze the probability that each type of remote sensing data at each time point reflects the absence of pollution in the river at that time. The smaller the distance, the greater the probability of no pollution.
[0042] The probability of no contamination in the target monitoring data at each monitoring point at each time is obtained by using the coordinate system corresponding to various other remote sensing monitoring data of each cluster.
[0043] Specifically, in a coordinate system, a point is formed by combining an abscissa and its corresponding ordinate. A point is formed by combining target monitoring data from a monitoring point at a given time with another type of remote sensing monitoring data; this point is denoted as the analysis point. Within the interval of the coordinate system containing the analysis point, the average distance between the analysis point and other points within the interval is calculated as the average distance corresponding to the analysis point at that monitoring point at that time. A negative correlation mapping is performed using an exponential function with a base of the natural constant to the ratio of the average distance corresponding to the analysis point to the average of the average distances corresponding to all points within the interval of the coordinate system containing the analysis point, yielding the relative distribution density corresponding to the analysis point at that monitoring point at that time. The mean of the relative distribution density corresponding to all points formed by the target monitoring data from the monitoring point at that time and other remote sensing monitoring data is taken as the probability of the target monitoring data at that monitoring point being uncontaminated at that time.
[0044] The specific calculation model for the average distance between a monitoring point and the point to be analyzed at a given time is as follows: , in, Let represent the average distance between the c-th monitoring point and the point to be analyzed at time i. Let 'a' represent the target monitoring data (the a-th type of remote sensing data), and 'b' represent the b-th type of remote sensing data. Here, in the coordinate system corresponding to the b-th type of remote sensing data, the target monitoring data at time i of the c-th monitoring point and the b-th type of remote sensing data together constitute the point to be analyzed. Within the interval of the coordinate system containing the point to be analyzed, the distance between the point to be analyzed and the j-th point among the other points within the interval is calculated. , This indicates the number of points within the coordinate system interval containing the point to be analyzed. Additionally, if there is only one point within an interval, the average distance is simply recorded as 0.1.
[0045] The specific calculation model for relative distribution density is as follows: , in, This represents the relative distribution density of the points to be analyzed at the i-th time point of the c-th monitoring point; represents the average of the average distances between points within the interval of the coordinate system containing the point to be analyzed; exp represents an exponential function with the natural constant as its base. This indicates that the smaller the average distance, the more likely there is to be no pollution at that moment.
[0046] The specific calculation model for the probability of no pollution is as follows: , in, Let N represent the probability of the target monitoring data (type a remote sensing monitoring data) at time i of monitoring point c being uncontaminated. N represents the number of types of remote sensing monitoring data. For the same target monitoring data at the same time of the same monitoring point, it can form various points together with other types of monitoring data at the same time of the same monitoring point. In other words, the target monitoring data at time i of monitoring point c can form multiple points together with other types of remote sensing monitoring data. Let represent the relative distribution density of the point to be analyzed at the c-th monitoring point at the i-th time. By calculating the mean of the relative distribution density of each point, we can obtain the probability of the target monitoring data at the c-th monitoring point at the i-th time being free from pollution.
[0047] For a given monitoring point at a given time, the probability of no contamination for various remote sensing monitoring data at that monitoring point at that time can be obtained. Furthermore, the average probability of no contamination for all remote sensing monitoring data at that monitoring point at that time is calculated as the average probability of no contamination for that monitoring point at that time. In subsequent analysis, it is necessary to eliminate moments that may indicate river pollution. Therefore, various remote sensing monitoring data from a single monitoring point at a single moment are combined to form a data point for that monitoring point at that moment, which is a data point containing multidimensional data.
[0048] Furthermore, the data points at each monitoring point at each time point are filtered using the average probability of no pollution to obtain a data point set. Specifically, a box plot is used to process the average probability of no pollution at each monitoring point at each time point to obtain the lower limit of the normal range of the average probability of no pollution, which serves as the filtering threshold. Data points at each monitoring point at each time point whose average probability of no pollution is less than the filtering threshold are removed, and the remaining data points form the data point set. Data at monitoring points whose average probability of no pollution is less than the filtering threshold may correspond to river pollution data, and therefore are removed. This yields the data point set consisting of the remaining data points.
[0049] The data correction module is used to cluster data points in the data point set according to the importance of each type of remote sensing monitoring data to obtain data point clusters; to cluster time points within a data point cluster to obtain time point clusters for that data point cluster; to determine the data point clusters corresponding to the current time based on the time point clusters of each data point cluster; to correct the various remote sensing monitoring data of each monitoring point at the current time using the various remote sensing monitoring data in the data point clusters corresponding to the current time; and to monitor river pollution using the corrected various remote sensing monitoring data of each monitoring point at the current time.
[0050] The above process filters data points by monitoring point and time, resulting in a data point set. Further, it is necessary to cluster the data points in the data point set to obtain individual data point clusters.
[0051] Specifically, the importance of various remote sensing monitoring data is used as the weight of each type of remote sensing monitoring data, and the weighted Euclidean distance between two data points is calculated using a Euclidean distance calculation model. Based on the weighted Euclidean distance between every two data points in the data point set, the data points in the data point set are clustered to obtain each data point cluster.
[0052] The clustering method used is k-means clustering, where the value of k is determined by the elbow method. The specific calculation model for the weighted Euclidean distance is as follows: , in, This represents the weighted Euclidean distance between data point A and data point B (where one data point refers to a single moment at a monitoring point), and N represents the total number of remote sensing monitoring data types. This indicates the importance of the a-th type of remote sensing monitoring data, that is, its importance in the background environmental condition similarity analysis. This represents the value of the a-th type of remote sensing data corresponding to data point A. This represents the value of the a-th type of remote sensing data corresponding to data point B. Data within the same data point cluster may correspond to similar background environmental conditions.
[0053] Based on the data distribution characteristics within each data point cluster, the probability of being in each background environmental condition at the current moment is predicted. Firstly, changes in background environmental conditions typically exhibit strong periodicity, such as seasonal changes leading to rises or falls in river levels, variations in river temperature, or periodic irrigation activities. Therefore, by analyzing the periodic distribution of data within each data point cluster, the probability of being in the corresponding environmental condition at the current moment is calculated. Since the duration of the impact of different types of background environmental conditions on the data varies—for example, irrigation and drainage are often short-term, likely affecting only one monitoring moment at a time within a one-week monitoring interval—while the impact of seasonal changes is often continuous, typically affecting data for a period of time.
[0054] To avoid the persistent background environment being fragmented by sudden events (such as pollution), making it difficult to determine its periodic patterns, hierarchical clustering is used to cluster the time periods of each data point within the same data point cluster, obtaining the time period clusters for that data point cluster. The total number of clusters during clustering is determined by the maximum distance method; that is, the clustering results obtained before a sudden increase in distance are the final clustering results. Similarly, the time period clusters corresponding to each data point cluster can be obtained.
[0055] Furthermore, the clusters of data points corresponding to the current time can be determined based on the clusters of time points within each data point cluster. Specifically, for a data point cluster, the time interval between the earliest and latest times of a time point cluster within that cluster is taken as the time interval corresponding to that time point cluster; the time intervals corresponding to each time point cluster of that data point cluster are arranged in chronological order, and the average time interval between any two time intervals is taken as the average time interval of that data point cluster; the average duration of the time intervals corresponding to each time point cluster of that data point cluster is calculated; starting from the latest time point in that data point cluster, a window with a duration equal to the average duration is constructed, and the sum of the average time interval and the average duration is used as the sliding step. The window slides from the starting time, and if the current time point is within the window, then that data point cluster is the cluster of data points corresponding to the current time point.
[0056] It should be noted that the average time interval and average duration of each data point cluster are obtained. Then, a sliding window method is used to determine whether the current moment satisfies the periodic change of the background environmental conditions corresponding to the data point cluster. When the window slides to the current moment, if the current moment is within the window, it means that the current moment satisfies the periodic change of the background environmental conditions corresponding to the data point cluster. This is used to determine which data point cluster's background environmental conditions the current moment is likely to fall into. Thus, the data point clusters corresponding to the current moment can be obtained.
[0057] Next, the remote sensing monitoring data of each monitoring point at the current time are corrected using the various remote sensing monitoring data in the clusters of data points corresponding to the current time.
[0058] Specifically, the clusters of data points corresponding to the current time are denoted as the clusters of data points to be analyzed; the mean of a type of remote sensing monitoring data in a cluster to be analyzed is obtained as the standard value of that type of remote sensing monitoring data for that cluster; the time periods corresponding to each time cluster of the cluster to be analyzed are arranged in chronological order, and the standard deviation of the time interval between every two time periods is obtained, denoted as the time interval standard deviation of the cluster to be analyzed; the standard deviation of the time duration corresponding to each time cluster of the cluster to be analyzed is calculated, denoted as the time duration standard deviation of the cluster to be analyzed; the normalized value of the time interval standard deviation of the cluster to be analyzed, the reciprocal of the hyperparameter, and the time duration of the cluster to be analyzed are then calculated. The existence probability of a cluster of data points is obtained by adding the normalized value of the standard deviation and the reciprocal of the hyperparameter. The existence probability of this cluster is then compared with the sum of the existence probabilities of all clusters of data points to obtain the weight of the cluster. The standard values of the target monitoring data of each cluster are weighted and summed using the weights of each cluster to obtain the adjustment parameter of the target monitoring data at the current time. The difference between the target monitoring data at the current time and the adjustment parameter of the target monitoring data at the current time is calculated to obtain the corrected target monitoring data at the current time for that monitoring point. Similarly, the corrected remote sensing monitoring data at the current time for each monitoring point is obtained.
[0059] Standard deviation is denoted as , represents the standard value of the a-th type of remote sensing monitoring data in the s-th data point cluster to be analyzed; based on the data point cluster corresponding to the current time obtained above, it may be predicted to be under various background environmental conditions, and the more similar the time intervals and the closer the durations of the time periods corresponding to the data point clusters corresponding to the current time obtained above, the higher the credibility of the data point cluster corresponding to the current time obtained above.
[0060] Therefore, the probability of the existence of the clusters corresponding to the data points to be analyzed was calculated, and the specific calculation model is as follows: , in, The probability of the existence of the cluster corresponding to the s-th data point to be analyzed represents the probability that the current point is in the cluster of the s-th data point to be analyzed; norm represents the normalization function; This represents the standard deviation of the time interval between clusters of the data points to be analyzed. The smaller the standard deviation of the time interval, the more accurate the above judgment is; that is, the current moment is more likely to be under the background conditions corresponding to the cluster of the s-th data point to be analyzed. These are hyperparameters used to ensure that the fractions are meaningful. The standard deviation of the time length of the cluster of the data points to be analyzed is denoted as .
[0061] The specific calculation model for the adjustment parameters of the target monitoring data at the current moment is as follows: , in, The adjustment parameter represents the target monitoring data (type a remote sensing monitoring data) at the current moment. Let be the weight corresponding to the cluster of the s-th data point to be analyzed. Let G be the standard value of the a-th type of remote sensing monitoring data for the s-th data point cluster to be analyzed, and let G represent the number of data point clusters to be analyzed, or the number of possible background environment types at the current moment. Then, subtracting the adjustment parameter of the target monitoring data at the current moment from the target monitoring data at a monitoring point yields the corrected target monitoring data for that monitoring point at the current moment. Similarly, corrected values can be obtained for all types of remote sensing monitoring data at each monitoring point at the current moment.
[0062] The river pollution is monitored using corrected remote sensing data from various monitoring points at the current time. The presence of river pollution is determined by comparing the corrected remote sensing data with a set anomaly threshold, which is set by relevant professionals based on experience. An anomaly alarm is triggered for remote sensing data exceeding the threshold, and the type of abnormal data is reported.
[0063] In summary, this application, based on the strong regularity of remote sensing data under the influence of the background environment in its temporal variations, constructs a standard model for remote sensing monitoring data under the influence of the background environment by analyzing the regularity of historical remote sensing monitoring data. This model is used to predict the performance of remote sensing monitoring data under the current background environment, thereby correcting the actual remote sensing monitoring data. Judging the river pollution status based on the corrected data can effectively improve the accuracy of pollution monitoring results.
[0064] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0065] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0066] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A smart monitoring system for river pollution around rice-growing areas based on remote sensing information, characterized in that, The system includes: The data preprocessing module is used to acquire various remote sensing monitoring data at each monitoring point at each time, and to record any type of remote sensing monitoring data as the target monitoring data; and to acquire the change in the target monitoring data at a monitoring point at a given time. The importance acquisition module is used to cluster the target monitoring data of each monitoring point in history to obtain each cluster; the change in the target monitoring data in a cluster is used as the x-axis, and the change in other remote sensing monitoring data is used as the y-axis to obtain the coordinate system corresponding to the other remote sensing monitoring data of the cluster; the importance of the target monitoring data is obtained by using the coordinate system corresponding to the other remote sensing monitoring data of each cluster. The data filtering module is used to obtain the probability of no pollution of the target monitoring data at each monitoring point at each time based on the coordinate system corresponding to various other remote sensing monitoring data of each cluster; to take the average probability of no pollution of various remote sensing monitoring data at a monitoring point at a time as the average probability of no pollution of that monitoring point at that time; to form the data points of that monitoring point at that time by combining various remote sensing monitoring data at a monitoring point at a time; and to filter the data points of each monitoring point at each time using the average probability of no pollution to obtain a set of data points. The data correction module is used to cluster data points in the data point set according to the importance of each type of remote sensing monitoring data to obtain data point clusters; to cluster time points within a data point cluster to obtain time point clusters for that data point cluster; to determine the data point clusters corresponding to the current time based on the time point clusters of each data point cluster; to correct the various remote sensing monitoring data of each monitoring point at the current time using the various remote sensing monitoring data in the data point clusters corresponding to the current time; and to monitor river pollution using the corrected various remote sensing monitoring data of each monitoring point at the current time.
2. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information as described in claim 1, characterized in that, The acquisition of the change in target monitoring data at a monitoring point at a given time includes: The difference between the target monitoring data of a monitoring point at a certain time and the target monitoring data of the same monitoring point at the previous time is taken as the change in the target monitoring data of that monitoring point at that time.
3. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information as described in claim 1, characterized in that, The method of using the change in target monitoring data within a cluster as the x-axis and the changes in other types of remote sensing monitoring data as the y-axis to obtain the coordinate system corresponding to the other types of remote sensing monitoring data in that cluster includes: In a cluster, the target monitoring data in the cluster are arranged in ascending order as the x-axis of a coordinate system. Another type of remote sensing monitoring data is denoted as the remote sensing monitoring data to be constructed. Among all the remote sensing monitoring data to be constructed, the remote sensing monitoring data to be constructed that have the same monitoring point and time as each target monitoring data in the cluster are found. The monitoring point and time are then matched with each target monitoring data on the x-axis as the y-axis to obtain the coordinate system corresponding to the remote sensing monitoring data to be constructed in the cluster. Similarly, the coordinate systems corresponding to other remote sensing monitoring data in the cluster are obtained.
4. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information as described in claim 1, characterized in that, The method of determining the importance of target monitoring data by using the coordinate system corresponding to various other remote sensing monitoring data of each cluster includes: Within a cluster, the Freedman-Diaconis rule is used to obtain the interval length of the x-coordinate in the coordinate system, and this interval length is used to divide the x-coordinate into various intervals. The standard deviation of the y-coordinates corresponding to each x-coordinate within an interval of the coordinate system corresponding to a certain type of remote sensing monitoring data is calculated and denoted as the y-coordinate standard deviation. A negative correlation mapping is performed using an exponential function with a base of the natural constant on the mean of the y-coordinate standard deviations corresponding to each interval of the coordinate system corresponding to this type of remote sensing monitoring data to obtain the correlation between the cluster's target monitoring data and this type of remote sensing monitoring data. Finally, the correlation between the cluster's target monitoring data and other types of remote sensing monitoring data is calculated. The mean of the correlation is used as the overall correlation of the cluster target monitoring data; the mean of the target monitoring data of each monitoring point at the current time is obtained and recorded as the average target monitoring data at the current time; the distance between the average target monitoring data at the current time and the cluster center of a cluster is added to the hyperparameter and the reciprocal is calculated to obtain the reciprocal distance of the cluster; the ratio of the reciprocal distance of a cluster to the sum of the reciprocals of the distances of all clusters is calculated to obtain the influence weight of the cluster target monitoring data; the importance of the target monitoring data is obtained by weighting and summing the overall correlation of the cluster target monitoring data using the influence weights of the cluster target monitoring data.
5. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information as described in claim 1, characterized in that, The step of obtaining the probability of no contamination of target monitoring data at each monitoring point at each time based on the coordinate system corresponding to various other remote sensing monitoring data of each cluster includes: In a coordinate system, a point is formed by combining an abscissa and its corresponding ordinate. A point is formed by combining target monitoring data from a monitoring point at a given time with another type of remote sensing data; this point is denoted as the analysis point. Within the interval of the coordinate system containing the analysis point, the average distance between the analysis point and other points within the interval is calculated as the average distance of the analysis point at that time. A negative correlation mapping is performed using an exponential function with a base of the natural constant to the ratio of the average distance of the analysis point to the average distance of all points within the interval of the coordinate system containing the analysis point, yielding the relative distribution density of the analysis point at that time. The mean of the relative distribution density of the points formed by the target monitoring data at that time and other remote sensing data is taken as the probability of the target monitoring data at that time being uncontaminated.
6. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information according to claim 1, characterized in that, The data point set obtained by filtering data points at each monitoring point at each time point using the average probability of no pollution includes: The average probability of no pollution at each monitoring point at each time point is processed using a box plot to obtain the lower limit of the normal range of the average probability of no pollution, which is used as the screening threshold. Data points at each monitoring point at each time point whose average probability of no pollution is less than the screening threshold are removed, and the remaining data points are retained to form a data point set.
7. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information as described in claim 1, characterized in that, The process of clustering data points in the data point set based on the importance of each type of remote sensing monitoring data to obtain data point clusters includes: The importance of various remote sensing monitoring data is used as the weight of each type of remote sensing monitoring data. Combined with the Euclidean distance calculation model, the weighted Euclidean distance between two data points is calculated. Based on the weighted Euclidean distance between every two data points in the data point set, the data points in the data point set are clustered to obtain each data point cluster.
8. The intelligent monitoring system for river pollution around rice-growing areas based on remote sensing information as described in claim 1, characterized in that, The step of determining the cluster of data points corresponding to the current time based on the clusters of each data point cluster at each time step includes: For a data point cluster, the time interval between the earliest and latest times of a time cluster within that data point cluster is taken as the time interval corresponding to that time cluster. The time intervals corresponding to each time cluster of the data point cluster are arranged in chronological order, and the average time interval between any two time intervals is taken as the average time interval of the data point cluster. The average duration of the time intervals corresponding to each time cluster of the data point cluster is calculated as the average duration of the data point cluster. Starting from the latest time in the data point cluster, a window with a duration equal to the average duration is constructed. The sliding step is the sum of the average time interval and the average duration. The window slides from the starting time. If the current time is within the window, then the data point cluster is the data point cluster corresponding to the current time.
9. A smart monitoring system for river pollution around rice-growing areas based on remote sensing information, as described in claim 1, is characterized in that... The step of correcting the various remote sensing monitoring data of each monitoring point at the current time using various remote sensing monitoring data in the clusters of data points corresponding to the current time includes: The data point clusters corresponding to the current time are denoted as the data point cluster to be analyzed. The mean of a type of remote sensing monitoring data within a data point cluster to be analyzed is obtained as the standard value for that type of remote sensing monitoring data in that data point cluster. The time periods corresponding to each time cluster of the data point cluster to be analyzed are arranged in chronological order, and the standard deviation of the time interval between every two time periods is obtained, denoted as the time interval standard deviation of the data point cluster to be analyzed. The standard deviation of the time duration corresponding to each time cluster of the data point cluster to be analyzed is calculated, denoted as the time duration standard deviation of the data point cluster to be analyzed. The normalized value of the time interval standard deviation of the data point cluster to be analyzed, the reciprocal of the hyperparameter, and the time duration standard deviation of the data point cluster to be analyzed are then used as the standard deviation of the time duration. The existence probability of a cluster of data points to be analyzed is obtained by adding the difference to the normalized value of the reciprocal of the hyperparameter; the existence probability of a cluster of data points to be analyzed is obtained by comparing it with the sum of the existence probabilities of all clusters of data points to be analyzed; the standard values of the target monitoring data of each cluster of data points to be analyzed are weighted and summed using the weights of each cluster of data points to be analyzed to obtain the adjustment parameter of the target monitoring data at the current time; the difference between the target monitoring data of a monitoring point at the current time and the adjustment parameter of the target monitoring data at the current time is obtained to obtain the corrected target monitoring data of the monitoring point at the current time, and similarly, the corrected remote sensing monitoring data of each monitoring point at the current time are obtained.