Online environment monitoring and supervision method and system based on multi-source data fusion

By adopting multi-source data fusion methods in environmental online monitoring technology, including fuzzy clustering and weighted fusion algorithms, the limitations of data integration and anomaly detection in the prior art are solved, and data management efficiency and reliability of monitoring results are improved.

CN120123922AInactive Publication Date: 2025-06-10SHENZHEN WEIKE ECOLOGICAL ENVIRONMENT ENGINEERING CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510089629.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing online environmental monitoring technology has limitations in the flexibility and accuracy of data processing, especially in terms of data integration and anomaly detection, and lacks effective mechanisms to integrate and analyze data from different sources.

Method used

Using a multi-source data fusion method, we collect and standardize environmental data, use a fuzzy clustering algorithm to perform data grouping and similarity evaluation, identify and isolate abnormal data points, and finally use a weighted fusion algorithm to integrate the purified data.

Benefits of technology

It improves the efficiency of data management and data reliability, enhances the judgment of data typicality, reduces the interference of wrong data on analysis results, and ensures the comprehensiveness and representativeness of monitoring results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123922A_ABST
    Figure CN120123922A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of online monitoring and supervision, in particular to an environment online monitoring and supervision method and system based on multi-source data fusion, and the method comprises the following steps: collecting environment data from a sensor, including gas emission, water quality indexes and temperature readings, and generating an initial monitoring data set through time synchronization and data formatting; and performing unified standardization processing based on the initial monitoring data set to generate a standardized monitoring data set. According to the method, the accuracy of data classification is optimized through the application of the fuzzy clustering algorithm, and the similarity between the environmental data can be more effectively identified, so that the efficiency of data management is improved. The similarity evaluation further refines the comparison among the data points, enhances the judgment of the data typicality, allows the isolation of abnormal data, reduces the interference of error data on the analysis result, and improves the reliability of the monitoring data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online monitoring and supervision, and particularly to an environmental online monitoring and supervision method and system based on multi-source data fusion. Background Art

[0002] The environmental online monitoring and supervision method uses online monitoring technology to track the environmental conditions in real time, such as air and water quality, in order to detect and respond to environmental pollution incidents in a timely manner. This method relies on sensors to collect environmental data, which are analyzed in real time and used to evaluate environmental quality and monitor pollution trends. However, the existing environmental online monitoring technology has limitations in the flexibility and accuracy of data processing, especially in data integration and anomaly detection. The lack of an effective mechanism to integrate and analyze data from different sources may lead to insufficient or misinterpreted data, especially under rapidly changing environmental conditions. Therefore, improvements are needed. Summary of the Invention

[0003] The object of the present invention is to solve the drawbacks existing in the prior art, and to propose an environmental online monitoring and supervision method and system based on multi-source data fusion.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions. An environmental online monitoring and supervision method based on multi-source data fusion includes the following steps:

[0005] Collect environmental data from sensors, including gas emissions, water quality indicators, and temperature readings, and generate an initial monitoring data set through time synchronization and data formatting; based on the initial monitoring data set, perform unified standardization processing to generate a standardized monitoring data set;

[0006] Based on the standardized monitoring data set, use the fuzzy clustering algorithm to group the data, classify similar environmental data into the same category, and generate a clustering result; based on the clustering result, calculate the similarity between the data points in each category and the cluster center, and analyze the typicality of each data point through similarity to generate a similarity evaluation result;

[0007] According to the similarity evaluation result, set a threshold to screen out abnormal data points and generate an abnormal data identification result; based on the abnormal data identification result, isolate the abnormal data from the standardized monitoring data set to generate a purified monitoring data set;

[0008] Based on the purified monitoring data set, use the weighted fusion algorithm to summarize the purified data, assign weights to the data sources, and obtain the environmental monitoring result according to the weights of each data source.

[0009] Preferably, the step of obtaining the initial monitoring data set is:

[0010] Deploy environmental sensors to collect gas emissions, water quality indicators, and temperature readings, and obtain raw environmental data through regular sampling;

[0011] Based on the raw environmental data, perform data alignment, match the data collection time by calibrating the time stamps, and eliminate incorrect or illogical data points to generate an initial monitoring data set.

[0012] Preferably, the steps for obtaining the standardized monitoring data set are as follows:

[0013] Based on the initial monitoring data set, scale the values to a unified range and unify the data format of each sensor to generate data in a unified format;

[0014] According to the data in the unified format, integrate the data points to form a standardized monitoring data set.

[0015] Preferably, the steps for obtaining the clustering result are as follows:

[0016] Based on the standardized monitoring data set, use fuzzy clustering to calculate the membership degree of each data point to the clustering center. The calculation formula is:

[0017]

[0018] where M ij is the membership degree of data point i to clustering center j, x i is the data point, v j and v k are the clustering centers, C is the number of clusters, and m is the fuzzy coefficient;

[0019] According to the membership degree of each data point, determine which category it belongs to, and group the data points with high membership degrees into the same category to obtain the clustering result.

[0020] Preferably, the steps for obtaining the similarity evaluation result are as follows:

[0021] Based on the clustering result, obtain the data of each category, including the clustering center and all data points belonging to each category, to obtain the data and center information of each category;

[0022] Based on the data and center information of each category, calculate the similarity between each data point and the center of the category to which it belongs. The calculation formula is:

[0023]

[0024] where S ij represents the similarity between data point i and clustering center j, x i is the data point, c j is the clustering center, σ is the within-class variance, λ is the adjustment coefficient, and n jis the number of data points in the jth cluster, and DN is the maximum number of data points;

[0025] According to the similarity of each data point, its typicality in the category to which it belongs is analyzed to obtain the similarity evaluation result.

[0026] Preferably, the steps for obtaining the abnormal data identification result are:

[0027] Based on the similarity evaluation result, the threshold of abnormal data points is calculated, and the calculation formula is:

[0028]

[0029] Among them, S i is the similarity of data point i, is the average of all similarities, EN is the total number of data points, min(S i ) is the minimum similarity, T is the threshold;

[0030] The threshold is used to filter data points whose similarity is lower than the threshold and identify them as abnormal data, thereby obtaining an abnormal data identification result.

[0031] Preferably, the steps of obtaining the purification monitoring data set are:

[0032] Based on the abnormal data identification result, identify and list all data points marked as abnormal to obtain an abnormal data point list;

[0033] From the standardized monitoring data set, the abnormal data points in the abnormal data point list are removed one by one to obtain a purified monitoring data set.

[0034] Preferably, the steps for obtaining the environmental monitoring results are:

[0035] Based on the purified monitoring data set, the weight of each data source is calculated using the following formula:

[0036]

[0037] Among them, W i represents the weight of data source i, d i represents the deviation value between data source i and the purified data set, τ is the deviation adjustment parameter, and N is the total number of data sources;

[0038] The data sources are weighted and fused according to the weights, and the observation results of each data source are combined to obtain the environmental monitoring results.

[0039] The present invention provides a supervision system, comprising:

[0040] The data collection module collects gas emissions, water quality indicators, and temperature readings from environmental monitoring sensors, synchronizes the time and formats it, and generates an initial monitoring dataset.

[0041] The data standardization module uniformly adjusts the data format and range based on the initial monitoring dataset, performs data normalization, and obtains a standardized monitoring dataset.

[0042] The data clustering module uses a grouping strategy for the standardized monitoring dataset, classifies environmental data into multiple categories according to similarity, and calculates the similarity of data points in each category based on the distance from the class center, obtaining the clustering and similarity evaluation results.

[0043] The anomaly monitoring module sets a threshold to screen data points based on the clustering and similarity evaluation results, identifies abnormal data points that deviate from the normal, separates the data points from the standardized monitoring dataset, and generates an abnormal data identification result.

[0044] The data fusion module uses the purified data in the abnormal data identification result to summarize the data, assigns weights to each data source, and obtains the environmental monitoring result.

[0045] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0046] The present invention optimizes the accuracy of data classification through the application of the fuzzy clustering algorithm, can more effectively identify the similarity between environmental data, thereby improving the efficiency of data management. The similarity evaluation further refines the comparison between data points, enhances the judgment of data typicality, allows the isolation of abnormal data, reduces the interference of incorrect data on the analysis results, and improves the reliability of monitoring data. And using the weighted fusion algorithm to integrate the purified data realizes the effective fusion of different data sources, ensures the comprehensiveness and representativeness of the monitoring results, and thus optimizes the data basis for decision support. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] Please refer to Figure 1 , the present invention provides a technical solution, an online environmental monitoring and supervision method based on multi-source data fusion, including the following steps:

[0050] Collect environmental data from sensors, including gas emissions, water quality indicators, and temperature readings. Through time synchronization and data formatting, generate an initial monitoring dataset; Based on the initial monitoring dataset, perform unified standardization processing to generate a standardized monitoring dataset;

[0051] Based on the standardized monitoring dataset, use the fuzzy clustering algorithm to group the data, classify similar environmental data into the same category, and generate a clustering result; Based on the clustering result, calculate the similarity between the data points in each category and the cluster center, analyze the typicality of each data point through similarity, and generate a similarity evaluation result;

[0052] According to the similarity evaluation result, set a threshold to screen out abnormal data points and generate an abnormal data identification result; Based on the abnormal data identification result, isolate the abnormal data from the standardized monitoring dataset to generate a purified monitoring dataset;

[0053] Based on the purified monitoring dataset, use the weighted fusion algorithm to summarize the purified data, assign weights to data sources, and obtain the environmental monitoring result according to the weights of each data source.

[0054] The steps for obtaining the initial monitoring dataset are as follows:

[0055] Deploy environmental sensors to collect gas emissions, water quality indicators, and temperature readings, and obtain the original environmental data through timed sampling;

[0056] Based on the original environmental data, perform data alignment, match the data collection time by calibrating the time stamp, and eliminate incorrect or illogical data points to generate an initial monitoring dataset.

[0057] Specifically, various types of sensors capable of measuring gas emissions, water quality indicators, and temperature are installed in the designated area. The distribution density of the sensors is determined in combination with local industry data and previous experimental experience, and their installation altitude and horizontal coordinates are recorded. The sampling frequency is set by dividing the daily monitoring period into several fixed intervals. For example, the gas emissions can be observed in the interval between 0 and 500 milligrams per cubic meter and it is recorded whether it exceeds this range. If a value below 0 or greater than 500 appears, it is marked as an abnormal sample in the record and rechecked in subsequent processing. For water quality indicators such as pH, it is compared and recorded in segments with reference to the complete range of 0 to 14 to check whether each sampling falls within this range. For temperature, the sampling is compared item by item with reference to the assumed interval of 0°C to 90°C. If the temperature exceeds 90°C or is below 0°C, it is regarded as questionable and marked separately. The upper and lower limits of these intervals can be adjusted according to the conventional extreme values of the local environment or the data information obtained from previous on-site measurements. Assuming that historical monitoring data shows that the extreme maximum temperature in the local area is about 80°C and the minimum temperature is about -5°C, then the temperature thresholds can be set to -10°C and 90°C during actual deployment to include a certain range of fluctuations. For water quality indicators, the intermediate demarcation points can also be defined according to the statistical results of water quality parameters collected previously. When some water quality data falls into the suspicious interval, additional records are made. All measurement results are stored in the monitoring records of the software system in chronological order, and the original environmental data for subsequent analysis is formed through the above method.

[0058] Based on the obtained original environmental data, first, the timestamp of each record is read and compared with the pre-scheduled sampling time. If it is found that the deviation between the data record and the expected time exceeds the set comparison range, for example, more than two seconds, it is judged as a time mismatch and classified into the suspected invalid entries. Subsequently, the gas emissions, water quality indicators, and temperature values of adjacent sampling results are compared one by one. If the gas emissions are below 0 milligrams per cubic meter or exceed 500 milligrams per cubic meter, they are included in the suspicious list. If the water quality indicators are not within the previously defined reference interval, such as between pH 0 and 14, they are also added to this list. For temperature, if the recorded value is below -10°C or above 90°C, it is temporarily stored in the abnormal list. These ranges are calculated based on the previous empirical data and the actual environment of the site. After finishing the collation, all the records marked as suspicious are checked again. If they really cannot be classified into the normal range with any reasonable explanation, they are excluded from the available entries. Finally, the screened data is summarized in chronological order to generate the initial monitoring data set.

[0059] The steps to obtain the standardized monitoring data set are as follows:

[0060] Based on the initial monitoring data set, the values are scaled to a unified range, and at the same time, the data format of each sensor is unified to generate data in a unified format;

[0061] Integrate data points according to data in a unified format to form a standardized monitoring data set.

[0062] Specifically, according to the initial monitoring data set obtained previously, first determine the data range that needs to be scaled and read the values item by item. If it is found that some values have no corresponding upper and lower limits in historical observations or reference manuals, calculations need to be performed in combination with the previous maximum and minimum values. For example, for gas emissions, its minimum value v min and maximum value v max respectively correspond to the actual observed extremes in the current environmental monitoring stage. Define the scaled result v′ i Satisfy where v i represents the original value of a certain record, v min and v max are respectively the minimum and maximum observed values of this type of data statistically obtained in the entire initial monitoring data set. When standardizing water quality indicators, the same method can be used and their distribution within the range of 0 to 14 can be determined. For temperature, it can be scaled in the same way by referring to the existing environmental data statistical range, such as -10°C to 90°C. The data units of all sensors, such as dissolved oxygen concentration or temperature measurement, need to be checked item by item against the previously sorted sensor description document to determine the unified unit and then perform conversion. For example, convert some temperatures recorded in Fahrenheit to Celsius and keep them consistent with other temperature data formats. In addition, meta-information such as the form of the recorded time stamp also needs to be compared. For example, uniformly use the identification method of year-month-day hour:minute:second. Process all readings in the above way and merge the results to finally generate data in a unified format.

[0063] Based on the data in the unified format obtained above, call each of its fields one by one and compare with the observation batches corresponding to the time stamps to read the records of fields such as gas emissions, water quality indicators, and temperature at the same sampling moment. Arrange the data at the same time point or within the same sequence in the same group and determine whether there are missing or duplicate items. If a missing item is found, it is supplemented according to the previously sorted statistical extreme values or conventional distribution values. For duplicate items, the method of removing duplicates can be used to retain the earliest or latest record with the time stamp. After finishing the collation of each group of data, arrange gas emissions, water quality indicators, temperature, etc. in a fixed field order and summarize them. During this period, if invalid or logically conflicting data points are detected again, they can be reviewed by referring to the previously recorded scaling range. When all groups are completed and it is confirmed that the record fields are consistent, these data points are gathered together to form a standardized monitoring data set.

[0064] The steps to obtain the clustering result are as follows:

[0065] Based on the standardized monitoring data set, use fuzzy clustering to calculate the membership degree of each data point to the clustering center. The calculation formula is:

[0066]

[0067] Among them, M ij is the membership degree of data point i to cluster center j, x i is the data point, v j and v k are the cluster centers, C is the number of clusters, and m is the fuzzy coefficient;

[0068] According to the membership degree of each data point, determine which category it belongs to, and classify the data points with high membership degrees into the same category to obtain the clustering result.

[0069] Specifically, the advantage of the formula is that by simultaneously considering the relative distance between x i and different cluster centers v j , the membership degree of each data point to different clusters is measured, so that the same data point can be reasonably distinguished or merged under the condition of fuzzy clustering. Combining the measured values of each parameter, a more delicate clustering result can be obtained in the system.

[0070] The acquisition steps of parameter x i are as follows: directly use the i-th observation data in the standardized monitoring dataset obtained in the previous step as x i . Each observation data may include the values or coordinates of gas emissions, water quality indicators, and temperature. According to these values, several dimensions are formed in space. Through corresponding measurement schemes (such as actual measurement and recording of gas emissions in the reasonable range of 0 to 500 milligrams per cubic meter in the system, recording of pH in the range of 0 to 14, and segmented reading of temperature between -10°C and 90°C, etc.), all the recorded data are summarized to obtain the data point vector x i .

[0071] The acquisition steps of parameters v j and v k are as follows: regard each cluster center as a vector with a fixed position in the multi-dimensional space. The center vector is refined and recorded in the system from the previous clustering initialization or historical monitoring data. For example, v j and v k are obtained by pre-grouping and statistically calculating the central values or average values of gas emissions, water quality indicators, temperature, etc. in a large number of records. If the maximum value is detected to be approximately 480 milligrams per cubic meter and the minimum value is approximately 15 milligrams per cubic meter during the statistics of gas emissions, the central value is calculated in combination with this range during the clustering initialization, and the multi-dimensional components of the corresponding center vector are recorded.

[0072] The steps to obtain parameter C are as follows: Determine the number of clusters required according to the distribution of the previously monitored data. Refer to the degree of difference among all records in the standardized monitoring dataset. By comparing the clustering results under different C values multiple times, select the most suitable number of clusters C that can distinguish data categories. For example, if it is found in actual detection that 3 to 5 classes can better distinguish the differences, record C = 3 or C = 4 or C = 5. The specific value depends on the indicators of the pre-clustering evaluation.

[0073] The steps to obtain parameter m are as follows: Combine the fuzzy coefficient selection process of fuzzy clustering. Extract the available coefficient range by referring to relevant literature or in previous experimental comparisons. Usually set m between 1.5 and 3. For multi-dimensional environmental monitoring data, the specific value can be determined by comparing the iterative stability under different m values and recording the final convergence effect. For example, during the monitoring process, test m = 1.8, m = 2.0, m = 2.2 and select the value that can best improve the numerical stability for recording.

[0074] Calculation process:

[0075] First step, define the distance d ij = ∥x i - v j ∥. According to the coordinate differences in the standardized monitoring data, if the i-th record is x i = (20, 7, 35) (for example, the gas emission is 20 mg / m³, pH value is 7, and temperature is 35 °C), and the j-th cluster center is v j = (25, 6.5, 40), then according to the multi-dimensional Euclidean distance formula, we can get:

[0076]

[0077] Second step, calculate the numerator part of the membership degree: If v k corresponds to another cluster center (22, 7, 38), then Then the ratio of the numerator to the denominator is Third step, sum over all k from 1 to C, take the reciprocal, and perform the exponential operation according to . If m = 2, then Finally, raise each ratio to the power of 2, sum them up, take the reciprocal, record the result, and obtain M ij ,

[0078] This result indicates that for the i-th record compared with different cluster centers, when M ij is relatively higher, this record is more likely to belong to the j-th cluster. If this value is significantly smaller than the membership degrees corresponding to other centers, it does not belong to the j-th cluster. Thus, the classification determination of the data in the fuzzy clustering stage can be completed.

[0079] Based on the membership degree of each data point, after obtaining the value of M ij previously, read them one by one. Compare the membership degree result of each record with the corresponding cluster center. If the membership degree value of a certain record in a certain cluster is relatively large, then regard this record as a candidate point of this cluster. During this process, parallel comparisons are made for multi-dimensional observed values such as the gas emission amount, water quality index, and temperature of each data. For example, if the membership degree of a certain record with respect to the first cluster center reaches more than 0.8 while it is only between 0.2 and 0.4 for other cluster centers, it can be determined that it is closer at the first cluster center. At this time, it needs to be included in the attribution list of the first category and continue to check whether there are other records that also show high values for the first cluster center. If the membership degrees of some records do not significantly exceed the range of 0.5, they are included in the ambiguous state list. For the data in the ambiguous state list, it can be re-verified by the actual observed quantity or re-partitioned after performing local operations again. During this period, if it is found that the gas emission amount exceeds the upper limit of the previous record, such as 500 milligrams per cubic meter, or the water quality pH falls below 0 or above 14, it is also necessary to check whether there are abnormalities in the recording link or the measurement link. If it is confirmed as an abnormality, remove it or mark it again. When all records complete the membership degree test and complete the filing of the classification list, the clustering result is finally obtained.

[0080] The steps to obtain the similarity evaluation result are as follows:

[0081] Based on the clustering result, obtain the data of each category, including the cluster center and all data points belonging to each category, and obtain the data and center information of each category;

[0082] Based on the data and center information of each category, calculate the similarity between each data point and the center of the category to which it belongs. The calculation formula is:

[0083]

[0084] Among them, S ij represents the similarity between data point i and cluster center j, x i is the data point, c j is the cluster center, σ is the within-class variance, λ is the adjustment coefficient, n j is the number of data points of the jth cluster, and DN is the maximum number of data points;

[0085] According to the similarity of each data point, analyze the typicality in the category to which it belongs, and obtain the similarity evaluation result.

[0086] Specifically, based on the previously generated clustering results, each class label is read and the corresponding cluster center is extracted. The coordinate or numerical information is recorded item by item. At the same time, the data points that have been assigned to this class are retrieved and their observed values in terms of gas emissions, water quality indicators, temperature, etc. are summarized. When scanning a certain class, all the records in the data table are opened in sequence and the class numbers are checked. The observed data that matches the class number is temporarily stored in a list for integration. During the integration stage, the acquisition time or sequence position is checked to prevent the situation of repeated inclusion at the same observation time and conflicts between different observation times. If it is monitored that the gas emissions exceed the reasonable upper limit obtained from previous experience (for example, 500 milligrams per cubic meter), it is listed in the suspected abnormal category. If the water quality pH is found to be less than 0 or exceed 14, it is recorded in the remarks column in the same way. If the temperature value exceeds the range of -10°C to 90°C, it is classified into the abnormal statistics. These contents that deviate from the reasonable range ultimately need to be compared again. After confirming that they are indeed abnormal, a decision is made on whether to completely eliminate or reassign them. In this way, a data list corresponding to each class is obtained and a class file is established in combination with the existing cluster center information. The center vectors and available data points of all classes are all aggregated into materials for subsequent analysis and retrieval, and thus the data and center information of each class are obtained.

[0087] The benefit of the formula lies in combining the distance between the data point and the center vector and the class size factor, which can more fully measure the similarity degree of a certain observation record in the current cluster.

[0088] The steps to obtain the σ parameter are as follows: Based on the distance distribution between the data points and the center within each class obtained previously, the variance of all distances is statistically calculated and the square root is taken. After multiple acquisitions, a record sequence is formed. Performing a variance operation on this sequence obtains the value of σ. For example, if the statistical variance of a certain class of internal data is close to 25, then σ = 5;

[0089] The steps to obtain the λ parameter are as follows: After multiple rounds of detection of different clusters, combining the influence of the number of data points in each cluster on the similarity distribution, by comparing the results of multiple groups of experiments and recording the fluctuations at different λ values, the coefficient value that can best reflect the role of the cluster size is determined. For example, λ = 0.3;

[0090] n j The steps to obtain the parameter are as follows: Directly count the total number of data points in the j-th cluster. For example, if 120 observation records are accumulated under a certain class, then n j = 120;

[0091] The steps to obtain the DN parameter are as follows: Check the class with the largest amount of data among all clusters and record the number of data points in this class. It is also possible to scan the entire standardized data and take the total number of entries as a reference. For example, if the largest number of entries monitored is 500, then DN = 500;

[0092] x i and c j are respectively from the previously determined observation values and cluster centers. For example, x i =(35, 8, 42), c j =(30, 7, 40), and their respective components correspond to values such as gas emissions, water quality indicators, and temperature.

[0093] Calculation process:

[0094] In the first step, calculate ∥x i -c j ∥ 2 , where (35 - 30) 2 +(8 - 7) 2 +(42 - 40) 2

[0095] = 25 + 1 + 4 = 30;

[0096] In the second step, calculate

[0097] In the third step, calculate and multiply by λ = 0.3 to get 0.215 × 0.3 ≈ 0.0645, and add this value to 1.2 to get 1.2645;

[0098] In the fourth step, take the negative value of 1.2645, i.e., -1.2645 and calculate e -1.2645 ≈ 0.2828, add 1 and 0.2828 to get 1.2828, and then find its reciprocal to get 0.779;

[0099] This result indicates that when S ij ≈ 0.779, the data point x i is considered to have a moderate degree of similarity to the center c j . If the calculation result is higher than 0.7, it can be determined that the data point is closer to this category center. If the value is lower than 0.3, it shows a weak intra-class typicality. Thus, the performance of each data point in its respective category can be further quantitatively evaluated.

[0100] Open and read these observation points one by one according to the similarity values between each data point and the center of its category. Select entries with large differences in the similarity distribution and check their detailed records regarding gas emissions, water quality indicators, and temperature. During this process, the pH value of the water quality indicator can be compared within the range of 0 to 14, the gas emissions can be compared within the range of 0 to 500 milligrams per cubic meter, and the temperature can be compared within the range of -10°C to 90°C. If it is found that the data point records exceed these ranges, query the on-site measurement values in the previous monitoring log to confirm whether there are deviations in the sensor link or the data recording link. When it is confirmed that there are no deviations, mark them in the similarity list. For entries with particularly low similarity values, it is also possible to additionally compare whether the data is distributed in abnormal areas. After determining that it is abnormal or deviates too much from the overall distribution, add it to the suspected list. Finally, summarize the remaining data points that fit well with the center into the category feature reference file, and sort out several data entries with relatively high typicality according to the similarity in the record, so as to obtain the similarity evaluation result.

[0101] The steps to obtain the abnormal data identification result are as follows:

[0102] Based on the similarity evaluation result, calculate the threshold of the abnormal data point. The calculation formula is:

[0103]

[0104] Among them, S i is the similarity of data point i, is the average value of all similarities, EN is the total number of data points, min(S i ) is the minimum similarity, and T is the threshold;

[0105] Use the threshold to screen out data points with similarities lower than the threshold, and identify them as abnormal data to obtain the abnormal data identification result.

[0106] Specifically, the benefit of the formula lies in comprehensively using the dispersion degree of the overall similarity distribution and the minimum similarity value. The standard deviation reflects the fluctuation range of all data points, and at the same time combines the role of the minimum similarity in ln to make a threshold determination for the most deviated observations. The following is a separate explanation of each parameter:

[0107] S i The steps to obtain the parameter are as follows: According to the similarity evaluation result generated previously, read the similarity value of the i-th data point. For example, when detecting gas emissions, water quality indicators, or temperature, by combining the similarity formula or model inference, form a value between 0 and 1 (or an extended range that may have a small negative deviation) for each record, and record these values together in chronological order or other order to form {S 1 , S 2 ,…, SEN The set of {

[0108] The steps to obtain the parameter are as follows: Add all S i in sequence and then divide by the total number of data points EN to obtain the average similarity. If there are already EN similarity records, they can be accumulated one by one and then the division operation can be performed to obtain

[0109] The steps to obtain the EN parameter are as follows: Directly calculate the total number of all data point entries involved in the previous similarity evaluation results. For example, if 200 similarity information items are output in one monitoring, then EN = 200.

[0110] min(S i ) The steps to obtain the parameter are as follows: Select the minimum similarity value from {S 1 , S 2 , …, S EN}. If the minimum value obtained from multiple monitorings is slightly less than 0, record this value for substitution into the formula.

[0111] Calculation process:

[0112] First step, calculate For example, in one statistics, EN = 5 similarities are obtained as -0.8, 0.3, 0.6, 0.9, -0.2. Then their sum is -0.8 + 0.3 + 0.6 + 0.9 + (-0.2) = 0.8. Therefore

[0113] Second step, calculate the variance part Square the difference between each similarity and the average value and then accumulate. Among them:

[0114] (-0.8 - 0.16) 2 = 0.9216,

[0115] (0.3 - 0.16) 2 = 0.0196,

[0116] (0.6 - 0.16) 2 = 0.1936,

[0117] (0.9 - 0.16) 2 = 0.5476,

[0118] (-0.2 - 0.16) 2 = 0.1296

[0119] Add the above results to get 1.812;

[0120] In the third step, divide the accumulated value by EN - 1 = 5 - 1 = 4, obtaining 0.453. Take the square root of it to get

[0121] In the fourth step, record min(S i ) = -0.8, then min(S i ) + 1 = 0.2, and its reciprocal 1 / 0.2 = 5. Take the natural logarithm ln(5) ≈ 1.609. Finally, multiply 0.673 by 1.609 to get T ≈ 1.082;

[0122] This result indicates that the threshold obtained in this monitoring is approximately 1.082. When the similarity of a certain record is lower than 1.082, it can be determined as abnormal. In the above example, the similarities of all records are less than this value. Therefore, it can be further discriminated in subsequent steps whether all entries are regarded as abnormal or whether segmented comparison is still required. If higher or lower similarity results appear in some other application scenarios, this threshold can be used to distinguish the observed data points that exceed or do not reach a certain level, thereby generating subsequent abnormal data identification and judgment.

[0123] Using the obtained threshold, read the previous similarity distribution item by item and compare it with the threshold 1.082. When the similarity record is less than 1.082, temporarily store this data in the suspected abnormal list. While conducting a one-by-one review, check the observed values such as gas emissions, water quality indicators, and temperature again. For example, compare the gas emissions with the range of 0 to 500 milligrams per cubic meter, compare the pH with the standard range between 0 and 14, and compare the temperature with the reference range between -10°C and 90°C. If it is monitored that a certain record is not within the reasonable range and the similarity is also lower than 1.082, mark it in the suspected abnormal list. Subsequently, it can be confirmed whether it is a real abnormality through on-site investigation or further comparison of the sensor logs. After all records are processed, collect all the data identified as abnormal to form the abnormal data identification result.

[0124] The steps for obtaining the purified monitoring dataset are as follows:

[0125] Based on the abnormal data identification result, identify and list all the data points marked as abnormal to obtain the abnormal data point list;

[0126] From the standardized monitoring dataset, remove the abnormal data points in the abnormal data point list one by one to obtain the purified monitoring dataset.

[0127] Specifically, based on the previously obtained abnormal data recognition results, read the records marked as abnormal. First, retrieve all data entries with similarity lower than the corresponding threshold or already marked as abnormal observation values and integrate them into a temporary list. Then, check item by item within this list whether the gas emission data falls within the previously agreed range of 0 to 500 milligrams per cubic meter. By comparing with the on-site monitoring extreme values provided in the context, it can be known that this range is obtained based on long-term observation statistics and combined with existing industry data. In addition, conduct a full-range search for the water quality index pH within the range of 0 to 14 and check whether there are any cases less than 0 or greater than 14. When an out-of-range situation is found, review the timestamp and measurement environment of this data entry according to the previously saved monitoring records. If no detection process error is confirmed, keep it in the abnormal temporary list. For temperature, also conduct a point-by-point comparison with the effective range of -10°C to 90°C. As long as each record is outside this range, it will be retained in the abnormal temporary list for subsequent summarization. If the similarity of some data is extremely low and each value deviates too much from the established range, it can also be marked first. Finally, gather these confirmed records together to form a list of abnormal data points. After collection, organize the list numbers and match and proofread them with the existing record numbers in the standardized monitoring dataset. And mark all the entries determined to be abnormal in a separate list for convenient subsequent processing. In this way, all abnormal data point lists are identified and listed.

[0128] Open the entries with the same serial numbers as the abnormal data point list from the existing standardized monitoring dataset for comparison. Lock the exact positions of each abnormal data according to the coordinates or observed values recorded in the list. Then remove these suspicious entries item by item. During the removal process, check again whether there are duplicate abnormal records to avoid repeated cleaning. At the same time, keep the data established as abnormal in a special statistical list. During the removal process, also check whether the pH observed values within the range of 0 to 14 of the water quality index all conform to the previous clustering assignment logic. Re-check whether there are any potential abnormal omissions for the gas emission observations within the range of 0 to 500 milligrams per cubic meter and the temperature records within the range of -10°C to 90°C. If records that meet both the similarity deviation and multiple values violating the preset range are found during the removal stage, add them to the abnormal list in the same way and conduct additional removals. Finally, retain the remaining entries to form a new monitoring data collection. This process ensures that only more reliable observed values are used for subsequent clustering or weighted calculations, and finally a purified monitoring dataset is obtained.

[0129] The steps to obtain environmental monitoring results are as follows:

[0130] Based on the purified monitoring dataset, calculate the weight of each data source. The calculation formula is:

[0131]

[0132] Among them, W i represents the weight of data source i, d i represents the deviation value between data source i and the purified data set, τ is the deviation adjustment parameter, and N is the total number of data sources;

[0133] Weighted fusion is performed on each data source according to the weights, and the observation results of each data source are combined to obtain the environmental monitoring result.

[0134] Specifically, the advantage of the formula is that by introducing d i the deviation metric between and the purified monitoring data set, and combining τ to adjust the deviation, a larger weight can be assigned to the data source with a smaller deviation value during multi-source fusion, so that the overall weighted result can better reflect the main reliable data.

[0135] d i The steps for obtaining the parameter are as follows: Numerically compare the observation sequence of data source i with the corresponding records in the purified monitoring data set at the same time period or the same batch. For the gas emission amount, the difference between the two can be checked in the range of 0 to 500 milligrams per cubic meter. For the water quality index pH, the difference between the two is verified in the range of 0 to 14. For the temperature, the difference is checked in the range of -10°C to 90°C. Multiple differences are statistically analyzed and a comprehensive deviation value d i is obtained through methods such as mean square error or absolute difference. If the mean square error is used, a series of differences can be squared first, then averaged, and finally square-rooted to obtain d i .

[0136] The steps for obtaining the τ parameter are as follows: Comprehensively consider the deviation distribution between all data sources and the purified monitoring data set. By debugging different τ values and recording the convergence situation of the deviation distribution in the scenario, a regulation value that can balance the maximum and minimum deviations in actual statistics is finally selected. For example, through the distribution comparison of dozens of observations, it is found that setting τ between 2 and 5 can better match most common industrial monitoring data. Finally, the specific τ is screened out and its source is recorded.

[0137] The steps for obtaining the N parameter are as follows: Directly count and record the total number of all data sources. For example, if 5 sensors or 5 monitoring channels are deployed on-site, then N = 5.

[0138] Calculation process:

[0139] First step, read the deviation value d of data source i one by one i and divide it by τ. If the deviation calculation is performed on the gas emission amount, water quality index, and temperature in a single monitoring and d 1 = 1.2, d 2 = 3.1, d 3

[0140] = 2.6, and then determine τ = 2 on-site and substitute it into:

[0141] d 1 / τ = 1.2 / 2 = 0.6, d 2 / τ = 3.1 / 2 = 1.55, d 3 / τ = 2.6 / 2 = 1.3

[0142] In the second step, calculate respectively That is:

[0143] e -0.6 ≈ 0.5488, e -1.55 ≈ 0.212, e -1.3 ≈ 0.2707

[0144] In the third step, sum up these three items and perform normalization processing. The denominator is

[0145] Then:

[0146]

[0147] This result shows that data source 1 obtains a weight ratio of approximately 53.2% in this fusion, data source 2 is approximately 20.5%, and data source 3 is approximately 26.3%. If a certain data source has a larger deviation value, then its weight value is smaller, thus contributing a relatively lower influence degree in the subsequent fusion. On the contrary, the proportion is larger. The whole process can provide a quantifiable weight allocation for multi-source online monitoring fusion.

[0148] According to the previous weight values, open the observation sequences of each data source within the same monitoring period in turn and make records. By comparing the gas emission values at the same moment, segment scanning can be carried out within the range of 0 to 500 milligrams per cubic meter. Multiply the values of each monitoring source in each segment by their corresponding weights, and then accumulate to obtain the fused gas emission result. At the same time, for water quality indicators, view the observation records of each data source according to the range of 0 to 14, multiply them by the corresponding weights and summarize them at the same moment. For temperature observations, also compare the range of -10°C to 90°C and complete the numerical merging according to the same weighted fusion method. If there are a small number of records that are null values or abnormal values during the fusion process, first compare with the observations at adjacent time points to judge their rationality and eliminate them as appropriate. When all data sources at all moments have completed this weighted fusion process, arrange these weighted fused time series data into a new monitoring sequence for downstream links to consult, so as to finally obtain the environmental monitoring results.

[0149] The present invention provides a supervision system, including:

[0150] A data collection module that collects gas emissions, water quality indicators, and temperature readings from environmental monitoring sensors, synchronizes time and formats it, and generates an initial monitoring dataset;

[0151] A data standardization module that, based on the initial monitoring dataset, uniformly adjusts the data format and range, and performs data normalization to obtain a standardized monitoring dataset;

[0152] A data clustering module that uses a grouping strategy for the standardized monitoring dataset, divides environmental data into multiple categories according to similarity, and calculates the similarity of data points in each category based on the distance from the class center to obtain a clustering and similarity evaluation result;

[0153] An anomaly monitoring module that, based on the clustering and similarity evaluation results, sets a threshold to screen data points, identifies abnormal data points that deviate from the norm, separates the data points from the standardized monitoring dataset, and generates an abnormal data identification result;

[0154] A data fusion module that uses the purified data in the abnormal data identification result to summarize the data, assigns weights to each data source, and obtains the environmental monitoring result.

[0155] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for online environmental monitoring and supervision based on multi-source data fusion, characterized in that: The following steps are involved: Collect environmental data from sensors, including gas emissions, water quality indicators, and temperature readings, and generate an initial monitoring data set through time synchronization and data formatting; perform unified standardization processing based on the initial monitoring data set to generate a standardized monitoring data set; Based on the standardized monitoring data set, a fuzzy clustering algorithm is used to group the data, classify similar environmental data into the same category, and generate clustering results; Based on the clustering results, the similarity between the data points in each class and the cluster center is calculated, and the typicality of each data point is analyzed through the similarity to generate a similarity evaluation result; According to the similarity evaluation result, a threshold is set to filter abnormal data points and generate an abnormal data identification result; Based on the abnormal data identification result, the abnormal data is isolated from the standardized monitoring data set to generate a purified monitoring data set; Based on the purified monitoring data set, a weighted fusion algorithm is used to summarize the purified data, weights are assigned to data sources, and environmental monitoring results are obtained according to the weights of each data source.

2. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the initial monitoring data set are: Deploy environmental sensors to collect gas emissions, water quality indicators and temperature readings, and obtain raw environmental data through regular sampling; Based on the original environmental data, data alignment is performed to match the time of data collection by calibrating the timestamp, remove erroneous or illogical data points, and generate an initial monitoring data set.

3. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the standardized monitoring data set are: Based on the initial monitoring data set, scaling the values ​​to a uniform range, unifying the data format of each sensor, and generating data in a uniform format; Based on the data in the unified format, the data points are integrated to form a standardized monitoring data set.

4. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the clustering results are: Based on the standardized monitoring data set, fuzzy clustering is used to calculate the membership of each data point to the cluster center. The calculation formula is: Among them, M ij is the degree of membership of data point i to cluster center j, x i is the data point, v j and v k is the cluster center, C is the number of clusters, and m is the fuzzy coefficient; According to the membership degree of each data point, determine which category it belongs to, classify the data points with high membership degrees into the same category, and obtain the clustering result.

5. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the similarity evaluation result are: Based on the clustering results, the data of each category is obtained, including the cluster center and all data points belonging to each category, to obtain the data and center information of each category; Based on the data and center information of each category, the similarity between each data point and the center of the category to which it belongs is calculated using the following formula: Among them, S ij represents the similarity between data point i and cluster center j, x i is a data point, c j is the cluster center, σ is the intra-class variance, λ is the adjustment coefficient, n j is the number of data points in the jth cluster, and DN is the maximum number of data points; According to the similarity of each data point, its typicality in the category to which it belongs is analyzed to obtain the similarity evaluation result.

6. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the abnormal data identification result are: Based on the similarity evaluation result, the threshold of abnormal data points is calculated, and the calculation formula is: Among them, S i is the similarity of data point i, is the average of all similarities, EN is the total number of data points, min(S i ) is the minimum similarity, T is the threshold; The threshold is used to filter data points whose similarity is lower than the threshold and identify them as abnormal data, thereby obtaining an abnormal data identification result.

7. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the purified monitoring data set are: Based on the abnormal data identification result, identify and list all data points marked as abnormal to obtain an abnormal data point list; From the standardized monitoring data set, the abnormal data points in the abnormal data point list are removed one by one to obtain a purified monitoring data set.

8. The method for online environmental monitoring and supervision based on multi-source data fusion according to claim 1 is characterized in that: The steps for obtaining the environmental monitoring results are: Based on the purified monitoring data set, the weight of each data source is calculated using the following formula: Among them, W i represents the weight of data source i, d i represents the deviation value between data source i and the purified data set, τ is the deviation adjustment parameter, and N is the total number of data sources; The data sources are weighted and fused according to the weights, and the observation results of each data source are combined to obtain the environmental monitoring results.

9. The supervision system of the online environmental monitoring supervision method based on multi-source data fusion according to any one of claims 1 to 8 is characterized in that: include: The data collection module collects gas emissions, water quality indicators and temperature readings from environmental monitoring sensors, synchronizes time and formats them to generate an initial monitoring data set; The data standardization module uniformly adjusts the data format and range based on the initial monitoring data set, performs data normalization, and obtains a standardized monitoring data set; The data clustering module uses a grouping strategy for the standardized monitoring data set to divide the environmental data into multiple categories according to similarity. The data points in each category are calculated based on the distance to the class center to obtain clustering and similarity evaluation results; The anomaly monitoring module sets a threshold to screen data points based on the clustering and similarity evaluation results, identifies abnormal data points that deviate from the norm, separates data points from the standardized monitoring data set, and generates abnormal data identification results; The data fusion module uses the purified data from the abnormal data identification results to summarize the data, assign weights to each data source, and obtain environmental monitoring results.

Citation Information

Cited By

  • Fixed pollution source monitoring data analysis method and system and storage medium

    CN121117486A