Abnormal data processing system for air pollution monitoring

By using technical means of data acquisition, preliminary evaluation, data processing and concentration abnormality detection modules in air pollution monitoring, the problem of isolated forests ignoring the sensitivity of each tree in air pollution monitoring has been solved, and more efficient abnormal data detection is achieved.

CN119807917BActive Publication Date: 2025-05-23ZHONGJU (SHAANXI) ENG CONSULTING MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293272.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-23
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing isolated forest ignores the sensitivity of each isolated tree to abnormal situations in air pollution monitoring, resulting in low accuracy in abnormal data detection.

Method used

Through data acquisition, preliminary evaluation, data processing and concentration abnormality detection modules, the preliminary abnormality degree of concentration data of each gas at the monitoring point is obtained, and the detection impact weight of each segmented sequence is determined through sequence segmentation, curve fitting and normal distribution fitting, and each isolated tree has different decision weights to improve detection accuracy.

Benefits of technology

The abnormal detection accuracy of the concentration data of each gas at the monitoring point is improved, and abnormal phenomena in the gas concentration data can be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807917B_ABST
    Figure CN119807917B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing technology, and specifically to an abnormal data processing system for air pollution monitoring, the system comprising: obtaining a concentration time series sequence and a standard time series sequence of a gas; performing a preliminary evaluation on the concentration time series sequence to obtain a preliminary abnormality degree; sequentially segmenting and fitting the concentration time series sequence to determine a subsequence and a target curve; dividing a segmentation sequence and a segmentation matching sequence based on the target curve; analyzing the offset characteristics of the gas concentration change trend based on the change in the correlation between the segmentation sequence and the standard time series sequence over time to determine the trend time series offset, and determining the detection influence weight in combination with the normal distribution fitting result of the concentration time series sequence; obtaining the abnormal data in the concentration data of each gas at the monitoring point according to the detection influence weight and the preliminary abnormality degree. The present application improves the detection accuracy of abnormal data in concentration data by adaptively determining the decision weight of an isolated tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an abnormal data processing system for air pollution monitoring. Background Art

[0002] Air pollution monitoring mainly focuses on the concentration of pollutants in the atmosphere. Among them, the monitoring of gas concentration is of great significance for timely discovering pollution sources and taking effective measures to reduce pollutant emissions. The monitoring objects of gas concentration are extensive, mainly including sulfur dioxide from the combustion process of fossil fuels, nitrogen oxides from automobile exhaust and industrial emissions, ozone from photochemical reactions, and volatile organic compounds (VOCs).

[0003] At present, various types of sensors are usually used to monitor and process the gas concentration in the area in real time using technologies such as the Internet of Things and data processing, so as to find the pollution source in time. However, long-term detection will generate massive data, and the data collected by the sensor not only contains valuable data, but also contains a large amount of redundant data. At the same time, when monitoring the data collected by the sensor, it is necessary to analyze the data for a period of time to detect abnormal data. At present, isolation forests are often used in the anomaly detection processing of large-scale data. However, when extracting gas concentration data to train isolated trees, since the sensitivity of the training samples of each isolated tree to abnormal situations is different, the ability of each isolated tree to detect abnormalities is also different. However, the decision weights of isolated trees in the isolation forest are usually set to the same, ignoring the sensitivity of the training samples of each isolated tree to abnormal situations, and there is a problem of low accuracy in detecting abnormal data in concentration data. Summary of the invention

[0004] The present application provides an abnormal data processing system for air pollution monitoring to solve the problem that the existing isolation forest ignores the sensitivity of the training samples of each isolated tree to abnormal conditions, resulting in low accuracy of abnormal data detection.

[0005] The abnormal data processing system for air pollution monitoring of the present application adopts the following technical solutions:

[0006] An embodiment of the present application provides an abnormal data processing system for air pollution monitoring, the system comprising the following modules:

[0007] A data acquisition module is used to obtain the concentration time series and standard time series of each gas at the monitoring point;

[0008] A preliminary evaluation module, used for measuring the difference between the concentration time series and the standard time series of each gas, and taking the result of the difference measurement as the preliminary abnormality degree of the concentration time series;

[0009] A data processing module is used to sequentially segment the concentration time series of each gas, fit and determine subsequences and a target curve for each subsequence; based on the trough points on the target curve, the concentration time series of each gas and the standard time series are divided into segmentation sequences and segmentation matching sequences; based on the change of the correlation between the segmentation sequence of the concentration time series and the standard time series over time, the deviation characteristics of the gas concentration change trend are analyzed to determine the trend time series deviation of the segmentation sequence; based on the trend time series deviation and the normal distribution fitting result of the concentration time series, the detection influence weight of the segmentation sequence is determined;

[0010] The concentration anomaly detection module is used to obtain the abnormal data in the concentration data of each gas at the monitoring point according to the detection influence weight of the segmentation sequence and the preliminary abnormal degree of the concentration time series where the segmentation sequence is located, and complete the anomaly detection of the gas concentration data.

[0011] Preferably, the initial abnormality degree of the concentration time series is obtained by:

[0012] The concentration time series of each gas is measured against the standard time series by difference measurement, and the result of difference measurement is used as the preliminary abnormality degree of the concentration time series.

[0013] Preferably, the method of sequentially segmenting the concentration time series of each gas, fitting to determine the subsequences, and acquiring the target curve of each subsequence is:

[0014] Determine a sampling period according to the sampling frequency of the concentration time series of each gas, and divide the concentration time series into a plurality of subsequences according to the sampling period;

[0015] A rectangular coordinate system is established with the sampling time node as the horizontal coordinate and the gas concentration of each sampling as the vertical coordinate to obtain a time series scatter plot of each gas concentration data;

[0016] Each subsequence is taken as a target subsequence, and a fitting curve corresponding to data points in the time series scatter plot of the target subsequence is obtained by curve fitting as a target curve of the target subsequence.

[0017] Preferably, the method of obtaining the segmentation sequence and the segmentation matching sequence by dividing the concentration time series sequence of each gas and the standard time series sequence based on the trough point on the target curve is as follows:

[0018] Obtain all trough points on the target curve, use the trough points to segment the target curve, treat the single-peak curve where each trough point is located as a single-peak curve, and record the time node range corresponding to each single-peak curve as the segmentation time period;

[0019] The sequences consisting of the gas concentration data corresponding to each segmented time period in the concentration time series sequence and the standard time series sequence are recorded as segmented sequence and segmented matching sequence respectively.

[0020] Preferably, the trend time series deviation of the segmentation sequence is determined as follows:

[0021] Determine the standard period of the characteristic sequence based on the autocorrelation of the similarity measurement results between the segmentation sequence contained in the concentration time series sequence and the segmentation matching sequence in the standard time series sequence;

[0022] For each segmented sequence included in the concentration time series, the period of the sequence composed of all the similarity measurement results other than the similarity measurement results between each segmented sequence in the feature sequence and the segmented matching sequence is calculated as the control period, and the absolute value of the difference between the standard period and the control period is taken as the period difference;

[0023] Calculate the sum of the absolute values ​​of the difference between the number of single-peak curves in the target curve formed by the subsequence of each segmented sequence and the number of single-peak curves in the target curve formed by the adjacent subsequence as the adjacent difference;

[0024] The product of adjacent differences and period differences is taken as the trend time series deviation of the segmented sequence.

[0025] Preferably, the standard period of the characteristic sequence is determined as follows:

[0026] Calculate the similarity measurement results between each segmentation sequence included in the concentration time series sequence and the segmentation matching sequence in the standard time series sequence respectively, and use the sequence composed of all the similarity measurement results as the feature sequence;

[0027] The autocorrelation coefficients of the characteristic sequence at different phase differences are calculated respectively, and the period corresponding to the autocorrelation coefficient with the maximum significant score is taken as the standard period of the characteristic sequence.

[0028] Preferably, the method for determining the detection influence weight of the segmentation sequence includes:

[0029] Determine the overall fitting deviation of each segmented sequence based on the fitting results of the normal distribution fitting of the data points in each segmented sequence;

[0030] Perform linear normalization on the trend time series deviation of all segmented sequences, and use the product of the linear normalization result of the trend time series deviation of each segmented sequence and the overall fitting deviation as the numerator;

[0031] The sum of the standard deviation of the normal distribution fitting curve of each segmentation sequence and the preset parameters is used as the denominator;

[0032] The ratio of the numerator to the denominator is used as the detection influence weight of each segmentation sequence.

[0033] Preferably, the overall fitting deviation of each segmentation sequence is determined as follows:

[0034] Perform normal distribution fitting on the data points in each segmentation sequence to obtain a normal distribution fitting curve for each segmentation sequence;

[0035] The value of the data in each segmented sequence in the normal distribution fitting curve is recorded as the ideal value, the absolute value of the difference between each data in the segmented sequence and its corresponding ideal value is recorded as the difference degree, and the maximum value of all the difference degrees is recorded as the maximum difference degree;

[0036] The cumulative result of the ratio of the difference degree of each data to the maximum difference degree in each segmentation sequence is taken as the overall fitting deviation of each segmentation sequence.

[0037] Preferably, the abnormal data in the concentration data of each gas at the monitoring point is obtained by:

[0038] The decision weight of each isolated tree is obtained according to the detection influence weight of each segmentation sequence and the preliminary abnormality degree of the gas concentration time series;

[0039] For each gas at the monitoring point, the isolation forest model is trained using the concentration time series of each gas;

[0040] The concentration data of each gas to be detected at the monitoring point is input into the trained isolation forest model, and the anomaly score of the concentration data to be detected is obtained by combining the decision weights of all isolated trees. The concentration data with anomaly scores greater than the threshold are regarded as abnormal data.

[0041] Preferably, the decision weight of each isolated tree is obtained as follows:

[0042] The product of the detection influence weight of each segmented sequence included in the concentration time series sequence of each gas and the preliminary abnormality degree of the concentration time series sequence of each gas is used as the detection influence weight of each segmented sequence included in the concentration time series sequence of each gas;

[0043] The detection influence weight of each segmentation sequence is used as the sample weight of the concentration data in each segmentation sequence;

[0044] Determine the number of samples for training isolated trees, calculate the ratio of the sample weight of each sample for training each isolated tree to the sum of the sample weights of all concentration data in the segmentation sequence where each sample is located, and use the normalized result of the cumulative sum of the ratio on the number of samples for training isolated trees as the decision weight of each isolated tree.

[0045] The beneficial effects of the present application are: first, the preliminary abnormality degree of the concentration time series is determined by analyzing the difference between the concentration data of each gas at each monitoring point in the target area and the standard data, so as to achieve an overall evaluation of the concentration time series; second, the concentration time series and the standard time series are divided by sequence segmentation and curve fitting to obtain segmentation sequences and segmentation matching sequences, which are used for subsequent more refined local feature extraction of gas concentration data within each day; then, the trend time series deviation of the segmentation sequence is determined by analyzing the deviation characteristics of the gas concentration change trend based on the change of the correlation between the segmentation sequence and the standard time series of the concentration time series over time; the detection influence weight of the segmentation sequence is determined based on the trend time series deviation and the normal distribution fitting result of the concentration time series, and the influence of influencing factors such as the movement of pollution sources at the monitoring points on the temporal deviation of the concentration data is analyzed. The influence of the characteristics is used to evaluate the existence of abnormal phenomena in each segmentation sequence; then, the training weight of each segmentation sequence of the training isolated tree is obtained based on the detection influence weight of each segmentation sequence and the preliminary abnormality degree of the time series concentration sequence; then, the training weight of each segmentation sequence is used as the sample weight of each concentration data in each segmentation sequence. In this way, for the gas concentration at a single monitoring point, there will be no large fluctuations at multiple sampling moments within a single sampling cycle during the diffusion process of the gas, and the degree of influence on the detection model to extract data features should be relatively similar, avoiding the impact of excessive assignment on the model ability; finally, the decision weights of all isolated trees are used to detect abnormal data in the concentration data, and the decision weights are adaptively determined based on the different sensitivity of the concentration data of the training isolated tree to abnormal data, thereby improving the accuracy of anomaly detection in the concentration data of each gas at the monitoring point. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 This is a structural block diagram of the abnormal data processing system for air pollution monitoring of the present application. DETAILED DESCRIPTION

[0048] In order to further explain the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the abnormal data processing system for air pollution monitoring proposed in the present application, its specific implementation, structure, features and effects are described in detail as follows in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0049] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0050] The specific scheme of the abnormal data processing system for air pollution monitoring provided by the present application is described in detail below with reference to the accompanying drawings.

[0051] See also Figure 1 , which shows a structural block diagram of an abnormal data processing system for air pollution monitoring provided by an embodiment of the present application, the system includes the following modules:

[0052] The data acquisition module uses multiple sensors to collect a large amount of gas concentration data when monitoring the gas concentration in the target area for air pollution monitoring. When performing abnormal monitoring on the gas concentration data collected by the sensors, it is necessary to analyze the gas concentration data over a period of time. When faced with massive monitoring data, it is particularly important to compress and store the gas concentration data. Therefore, it is necessary to first obtain the gas concentration data collected by multiple sensors in the target area.

[0053] A variety of sensors are placed at monitoring points in the target area to collect gas concentration data of surrounding links in real time. The target area includes but is not limited to industrial parks, factories, and farming areas; the sensors include but are not limited to VOCs detection probes and hydrogen sulfide sensors. The monitoring objects of the embodiments of the present application include but are not limited to VOCs, particulate matter, ozone, hydrogen sulfide, carbon monoxide, nitric oxide, and hydrogen chloride. The VOCs detection probe is used to collect monitoring data of VOCs; the ozone monitoring sensor is used to collect monitoring data of ozone; the hydrogen sulfide sensor can be used to collect gases including but not limited to hydrogen sulfide, carbon monoxide, nitric oxide, and hydrogen chloride. The implementer can choose a suitable sensor according to the situation of the target area, and this application does not impose any special restrictions on this.

[0054] In the target area, a variety of sensors can be used to collect concentration data of polluted gases in the atmosphere in real time, realizing remote online monitoring and data analysis. However, long-term detection will generate massive data, and the data collected by the sensors not only contains valuable data, but also contains a large amount of redundant data. At the same time, when monitoring the data collected by the sensors, it is necessary to analyze the data for a period of time, and then firstly store the data collected by the sensors. Directly storing the data collected by the sensors in the database will bring great pressure to the server. In order to avoid taking up too much space, the massive time series data collected by the sensors needs to be compressed first.

[0055] In this application, the concentration data of each gas at each monitoring point in the target area is sampled and analyzed every day to obtain the concentration time series A of each gas at each sampling. , where n represents the number of samples per day, Represents the gas concentration data obtained by the nth sampling; for each gas, a set of time series data with normal gas concentration within a day is manually selected as the standard time series sequence for each gas .

[0056] It should be noted that when sampling and analyzing the concentration data of each gas monitored every day, in order to avoid too sparse sampling resulting in excessive difference in data distribution between the sampling results and the monitoring data, the sampling frequency in this application is not less than 10 Hz. Preferably, as an embodiment, sampling is performed with 30 Hz as the sampling frequency.

[0057] In the preliminary evaluation module, at each monitoring point in the target area, the sensor performs repeated periodic monitoring every day. Therefore, under normal circumstances, there is a large similarity between the time series data of the periodic ambient gas concentration collected by the sensor every day. The time series data without abnormal conditions is set as the standard data, and the concentration time series data obtained by sampling every day is analyzed for differences with the standard data. The greater the difference, the more likely it is that abnormal conditions such as excessive pollution and excessive emissions will occur at the monitoring point, and the more important the concentration time series data is.

[0058] Specifically, the concentration time series A of each gas is measured with the standard time series B by using the difference measurement method, and the result of the difference measurement is used as the preliminary abnormality degree of the concentration time series A. , the initial abnormality level The smaller the concentration time series A is, the lower the possibility of abnormal conditions such as excessive pollution and excessive emissions will occur. The larger the value is, the higher the probability that the concentration time series A will have abnormal conditions such as excessive pollution or excessive emissions. The difference measures between the sequences include but are not limited to Euclidean distance, DTW distance, value variance, and Flamenche distance. Preferably, as an embodiment of the present application, the DTW distance between the concentration time series A and the standard time series B is calculated as the preliminary abnormality degree of the concentration time series A. .

[0059] Data processing module, since the difference measurement of the concentration time series sequence A of each gas and the standard time series sequence B can only reflect the difference in data distribution within the two sequences, but since the pollution sources distributed around each monitoring point in the target area may be different, resulting in differences in the changing trends of the gas concentration monitoring data in different time periods, it is difficult to accurately detect abnormal concentration data by only obtaining the preliminary abnormality degree of the concentration time series sequence A by means of difference measurement. This application considers segmenting the concentration time series sequence of each gas, and further analyzing the abnormality of the concentration data by analyzing the changing trends of the gas concentration data in different time periods.

[0060] In the atmosphere, the diffusion effect of polluted gases is that a gas always diffuses from the pollution source to the surrounding area, that is, from the place with high concentration to the place with low concentration. The sensor at each monitoring point is fixed, and the concentration data of each gas collected by the sensor will change from low concentration to high concentration and then from high concentration to low concentration due to factors such as the change of the pollution source position within the sensor monitoring range in different time periods or the impact of the surrounding environment on the pollution source. The collected data is approximately normally distributed.

[0061] Firstly, the concentration data of each gas collected by each sensor at each monitoring point is divided into sampling periods, the sampling period is determined according to the sampling frequency, and the concentration data of each gas is divided into a plurality of subsequences according to the sampling period.

[0062] Secondly, a rectangular coordinate system is established with the sampling time node as the horizontal axis and the gas concentration of each sampling as the vertical axis to obtain a time series scatter plot of each gas concentration data within each day.

[0063] Afterwards, due to the uncertainty of the distribution of gas masses, there may be multiple gas masses in different spatial directions within the detection range of each sensor, resulting in multiple fluctuations in gas concentration data in each direction sequence, and the subsequences obtained by segmentation of the mutation points need to be further divided.

[0064] Specifically, any subsequence is denoted as a target subsequence, and the least squares method is used to perform curve fitting on the points of the target subsequence on the time series scatter plot of each gas, which is denoted as the target curve. The curve fitting method can reduce the influence of noise points in the scatter points on the curve fluctuation trend.

[0065] Furthermore, all trough points on the target curve are obtained, and the target curve is segmented using the trough points. The obtained single-peak curve is recorded as a single-peak curve; the time node range corresponding to each single-peak curve is recorded as a segmentation time period; the sequence composed of gas concentration data corresponding to each segmentation time period in the concentration time series sequence A and the standard time series sequence B is recorded as a segmentation sequence and a segmentation matching sequence, respectively. It should be noted that each segmentation sequence corresponds to a segmentation matching sequence.

[0066] It should be noted that this embodiment only provides a curve fitting method, that is, using the least squares method for curve fitting. On the premise that curve fitting can be achieved for the data points in the time series scatter plot, in other instances, the implementer can use other curve fitting methods to obtain the target curve.

[0067] Furthermore, when measuring the difference between the concentration time series A and the standard time series B of each gas, usually only the size of the data is considered, and the influence of the time attribute is ignored. In the concentration time series A and the standard time series B, there may be a situation where the concentration data has the same change trend in different time periods, which means that the gas concentration trend of the day has shifted in time in the concentration time series A. Therefore, according to the above process, this application analyzes the shift characteristics of the gas concentration change trend based on the change of the correlation between each segmented sequence of the concentration time series A and the standard time series B over time.

[0068] First, the similarity measurement results between each segmented sequence in the concentration time series A and its corresponding segmented matching sequence are calculated respectively, and the sequence composed of all the similarity measurement results in time order is recorded as a feature sequence T. The similarity measurement results include but are not limited to cosine similarity, Jaccard coefficient, intersection-over-union ratio, and preferably, as an embodiment of the present application, the Jaccard coefficient between each segmented sequence and its corresponding segmented matching sequence is calculated as the similarity measurement result.

[0069] Secondly, the autocorrelation coefficient of the feature sequence at different phase differences is calculated respectively, and the period corresponding to the autocorrelation coefficient of the maximum significant score is used as the standard period of the feature sequence. For any segmented sequence, the similarity measurement result between each segmented sequence and its corresponding segmented matching sequence is deleted from the feature sequence, and the period of the feature sequence after deletion is calculated as the control period of the segmented sequence. Among them, the autocorrelation coefficient is a commonly used technology in the field of data processing, and the specific process will not be repeated.

[0070] Here, the trend time series deviation is calculated to characterize the strength of the characteristic deviation of the concentration data change trend in each segmented sequence in the time dimension. The calculation formula for the trend time series deviation of the jth segmented sequence in the concentration time series sequence A is as follows:

[0071]

[0072] In the formula, is the trend time series deviation of the jth segmentation sequence in the concentration time series A, , are the absolute values ​​of the difference between the number of single-peak curves in the fitting curve formed by the subsequence where the j-th segmentation sequence is located and the number of single-peak curves in the fitting curve formed by the adjacent previous and next subsequences, t is the standard period of the characteristic sequence T, is the control period of the jth segmentation sequence in the concentration time series A.

[0073] Among them, adjacent difference It represents the difference in data distribution between the fitting curve corresponding to the segmentation time period of the jth segmentation sequence and the fitting curve corresponding to the adjacent subsequence. The larger the value of is, the more obvious the trend difference is in the subsequence of the jth segmentation sequence when the overall data distribution is relatively flat. It reflects the influence of the concentration data distribution between the j-th segmentation sequence and its segmentation matching sequence on the concentration data distribution in the concentration time series A and the standard time series B. The larger the value, the greater the influence of the concentration data distribution between the j-th segmentation sequence and its segmentation matching sequence on the period; that is, The larger the value of is, the greater the difference in the change trend between the concentration data in the jth segmentation sequence and the concentration data in the same time period in the standard time series B is, and the greater the possibility of the existence of the change trend deviation feature.

[0074] Furthermore, the trend time series deviation of each segmented sequence in the concentration time series sequence A is obtained respectively, and the trend time series deviation of all segmented sequences is linearly normalized.

[0075] Secondly, perform normal distribution fitting on the data points in each segmented sequence, obtain the normal distribution fitting curve of each segmented sequence, and obtain the standard deviation of each normal distribution fitting curve; obtain the value of the time node corresponding to each data in the segmented sequence in the normal distribution fitting curve, and record it as the ideal value; obtain the absolute value of the difference between each data in the segmented sequence and its corresponding ideal value, record it as the degree of difference, and record the maximum value of the degree of difference as the maximum degree of difference; it should be noted that normal distribution fitting is a prior art, and the embodiments of this application will not be repeated. Based on the analysis results of whether the concentration data in each segmented sequence conforms to the normal distribution, combined with the distribution trend of the concentration data in each segmented sequence, a comprehensive assessment is made of the degree of influence of the distribution characteristics of the concentration data in each segmented sequence on the actual abnormal concentration data in the detection concentration time series A.

[0076] Here, taking the jth segmentation sequence as an example, the detection influence weight of the jth segmentation sequence is calculated:

[0077]

[0078]

[0079] in, is the overall fitting deviation of the jth segmentation sequence in the concentration time series A, Indicates The first segmentation sequence The degree of difference in the data, Indicates The maximum difference between the segmented sequences, Indicates The number of data in the segmented sequence;

[0080] is the detection influence weight of the jth segmentation sequence in the concentration time series A, Indicates The standard deviation of the normal distribution fitting curve of the segmented sequence, is the normalized result of the trend time series deviation of the jth segmentation sequence in the concentration time series A. is a tuning constant used to prevent the denominator from being 0. The size of is a positive number not greater than 0.1. Preferably, in this embodiment, The value of is 0.001.

[0081] Under normal circumstances, the gas concentration data collected by the sensor at each monitoring point changes from low concentration to high concentration and then from high concentration to low concentration as the location of the pollution source within the sensor monitoring range changes or the impact of the surrounding environment on the pollution source. It is similar to a normal distribution. Therefore, the larger the standard deviation of the normal distribution fitting result of the j-th segmentation sequence, the flatter the normal distribution curve, and the smaller the standard deviation, the taller the normal distribution curve. The smaller the value of The greater the change in the gas concentration data corresponding to the segmentation sequence, the greater the change in the gas concentration at the monitoring point. The gas concentration data corresponding to each segmentation sequence is more important for anomaly detection; The larger the value, the The first segmentation sequence The higher the probability of mutation of data, The larger the value of , the greater the possibility that abnormal data exists in the concentration data of the j-th segmentation sequence in the concentration time series sequence A, and the greater the impact on the abnormality detection result.

[0082] The concentration anomaly detection module obtains the abnormal data in the concentration data of each gas at each monitoring point according to the detection influence weight of each segmentation sequence and the preliminary abnormal degree of the gas concentration time series, and completes the anomaly detection of the gas concentration data.

[0083] In the present application, through the data analysis results of the concentration data of various gases collected by sensors at various monitoring points in the target area, the detection model is trained using the data in different segmentation sequences to detect abnormal data in the concentration data of various gases.

[0084] Among them, when using data in different segmentation sequences to train the detection model, different data weights are assigned according to the detection influence weight of the segmentation sequence and the initial abnormality degree of the gas concentration time series, so as to improve the accuracy of anomaly detection when the detection model is facing the concentration data of various gases at different monitoring points in the target area.

[0085] Specifically, based on the preliminary abnormal degree of the concentration time series sequence A (B) of the b-th gas and the monitoring influence weight of each segmented sequence of the concentration time series sequence A (B), the training weight of each segmented sequence used for detection model training is determined. The training weight of the p-th segmented sequence contained in the concentration time series sequence A (B) is calculated as follows:

[0086]

[0087] In the formula, is the training weight of the pth segmentation sequence contained in the concentration time series A (B), is the initial abnormality degree of the concentration time series A (B), is the detection influence weight of the pth segmentation sequence contained in the concentration time series A (B).

[0088] Furthermore, the training weight of each segmentation sequence is used as the sample weight of each concentration data in each segmentation sequence. The reason for this assignment is that for the gas concentration at a single monitoring point, there will not be large fluctuations at multiple sampling moments within a single sampling cycle during the diffusion process of the gas, and the degree of influence on the detection model to extract data features should be relatively similar.

[0089] Secondly, for each gas at any monitoring point in the target area, the sample weight of each concentration data in each segmented sequence contained in the concentration time series of each gas is determined according to the above process. Secondly, M concentration data are randomly extracted from the concentration time series of each gas to train an isolated tree until the depth of the isolated tree reaches the preset depth L. K isolated trees form an isolated forest. The sizes of M, L, and K are taken as empirical values ​​of 128, 10, and 100 respectively. The training of the isolated forest is a well-known technology, and the specific process will not be repeated.

[0090] Specifically, for each isolated tree, the decision weight of each isolated tree when detecting the concentration data is determined based on the sample weight of the concentration data when training each isolated tree.

[0091] Taking the kth isolated tree as an example, the calculation formula of the decision weight of the kth isolated tree is as follows:

[0092]

[0093] In the formula, is the decision weight of the kth isolated tree, is the sample weight of the rth concentration data for training the kth isolated tree, It is the cumulative sum of the sample weights of all concentration data in the segmentation sequence where the rth concentration data of the kth isolated tree is located. is the number of concentration data for training the kth isolated tree, It is a normalization function that ensures that the sum of the decision weights of all isolated trees is 1.

[0094] According to the above steps, the decision weight of each isolated tree is determined respectively. Secondly, the concentration data of the gas to be detected is input into the isolated forest trained with the concentration data in the concentration time series of the gas, and the abnormal score of the input concentration data is determined based on the output and decision weight of all isolated trees. The concentration data with an abnormal score greater than the threshold of 0.7 is regarded as abnormal data, and the abnormal detection of the concentration data of each gas at any monitoring point in the target area is completed.

[0095] It should be noted that the above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

[0096] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

Claims

1. An abnormal data processing system for air pollution monitoring, characterized in that: The system includes the following modules: A data acquisition module is used to obtain the concentration time series and standard time series of each gas at the monitoring point; A preliminary evaluation module, used for measuring the difference between the concentration time series and the standard time series of each gas, and taking the result of the difference measurement as the preliminary abnormality degree of the concentration time series; A data processing module is used to sequentially segment the concentration time series of each gas, fit and determine subsequences and a target curve for each subsequence; based on the trough points on the target curve, the concentration time series of each gas and the standard time series are divided into segmentation sequences and segmentation matching sequences; based on the change of the correlation between the segmentation sequence of the concentration time series and the standard time series over time, the deviation characteristics of the gas concentration change trend are analyzed to determine the trend time series deviation of the segmentation sequence; based on the trend time series deviation and the normal distribution fitting result of the concentration time series, the detection influence weight of the segmentation sequence is determined; The concentration anomaly detection module is used to obtain the abnormal data in the concentration data of each gas at the monitoring point according to the detection influence weight of the segmentation sequence and the preliminary abnormal degree of the concentration time series where the segmentation sequence is located, and complete the abnormal detection of the gas concentration data; The method for obtaining abnormal data in the concentration data of each gas at the monitoring point is as follows: The decision weight of each isolated tree is obtained according to the detection influence weight of each segmentation sequence and the preliminary abnormality degree of the gas concentration time series; For each gas at the monitoring point, the isolation forest model is trained using the concentration time series of each gas; The concentration data of each gas to be detected at the monitoring point is input into the trained isolation forest model, and the anomaly score of the concentration data to be detected is obtained by combining the decision weights of all isolated trees. The concentration data with anomaly scores greater than the threshold are regarded as abnormal data.

2. The abnormal data processing system for air pollution monitoring according to claim 1, characterized in that: The method for obtaining the preliminary abnormal degree of the concentration time series is as follows: The concentration time series of each gas is measured against the standard time series by difference measurement, and the result of difference measurement is used as the preliminary abnormality degree of the concentration time series.

3. The abnormal data processing system for air pollution monitoring according to claim 1, characterized in that: The method of sequentially segmenting the concentration time series of each gas, fitting to determine the subsequences, and obtaining the target curve of each subsequence is as follows: Determine a sampling period according to the sampling frequency of the concentration time series of each gas, and divide the concentration time series into a plurality of subsequences according to the sampling period; A rectangular coordinate system is established with the sampling time node as the horizontal coordinate and the gas concentration of each sampling as the vertical coordinate to obtain a time series scatter plot of each gas concentration data; Each subsequence is taken as a target subsequence, and a fitting curve corresponding to data points in the time series scatter plot of the target subsequence is obtained by curve fitting as a target curve of the target subsequence.

4. The abnormal data processing system for air pollution monitoring according to claim 1, characterized in that: The method for obtaining the segmentation sequence and the segmentation matching sequence by dividing the concentration time series sequence of each gas and the standard time series sequence based on the trough point on the target curve is as follows: Obtain all trough points on the target curve, use the trough points to segment the target curve, treat the single-peak curve where each trough point is located as a single-peak curve, and record the time node range corresponding to each single-peak curve as the segmentation time period; The sequences consisting of the gas concentration data corresponding to each segmented time period in the concentration time series sequence and the standard time series sequence are recorded as segmented sequence and segmented matching sequence respectively.

5. The abnormal data processing system for air pollution monitoring according to claim 1, characterized in that: The trend time series deviation of the segmentation sequence is determined as follows: Determine the standard period of the characteristic sequence based on the autocorrelation of the similarity measurement results between the segmentation sequence contained in the concentration time series sequence and the segmentation matching sequence in the standard time series sequence; For each segmented sequence included in the concentration time series, the period of the sequence composed of all the similarity measurement results other than the similarity measurement results between each segmented sequence in the feature sequence and the segmented matching sequence is calculated as the control period, and the absolute value of the difference between the standard period and the control period is taken as the period difference; Calculate the sum of the absolute values ​​of the difference between the number of single-peak curves in the target curve formed by the subsequence of each segmented sequence and the number of single-peak curves in the target curve formed by the adjacent subsequence as the adjacent difference; The product of adjacent differences and period differences is taken as the trend time series deviation of the segmented sequence.

6. The abnormal data processing system for air pollution monitoring according to claim 5, characterized in that: The standard period of the characteristic sequence is determined as follows: Calculate the similarity measurement results between each segmentation sequence included in the concentration time series sequence and the segmentation matching sequence in the standard time series sequence respectively, and use the sequence composed of all the similarity measurement results as the feature sequence; The autocorrelation coefficients of the characteristic sequence at different phase differences are calculated respectively, and the period corresponding to the autocorrelation coefficient with the maximum significant score is taken as the standard period of the characteristic sequence.

7. The abnormal data processing system for air pollution monitoring according to claim 1, characterized in that: The method for determining the detection influence weight of the segmentation sequence includes: Determine the overall fitting deviation of each segmented sequence based on the fitting results of the normal distribution fitting of the data points in each segmented sequence; Perform linear normalization on the trend time series deviation of all segmented sequences, and use the product of the linear normalization result of the trend time series deviation of each segmented sequence and the overall fitting deviation as the numerator; The sum of the standard deviation of the normal distribution fitting curve of each segmentation sequence and the preset parameters is used as the denominator; The ratio of the numerator to the denominator is used as the detection influence weight of each segmentation sequence.

8. The abnormal data processing system for air pollution monitoring according to claim 7, characterized in that: The overall fitting deviation of each segmentation sequence is determined as follows: Perform normal distribution fitting on the data points in each segmentation sequence to obtain a normal distribution fitting curve for each segmentation sequence; The value of the data in each segmented sequence in the normal distribution fitting curve is recorded as the ideal value, the absolute value of the difference between each data in the segmented sequence and its corresponding ideal value is recorded as the difference degree, and the maximum value of all the difference degrees is recorded as the maximum difference degree; The cumulative result of the ratio of the difference degree of each data to the maximum difference degree in each segmentation sequence is taken as the overall fitting deviation of each segmentation sequence.

9. The abnormal data processing system for air pollution monitoring according to claim 8, characterized in that: The decision weight of each isolated tree is obtained as follows: The product of the detection influence weight of each segmented sequence included in the concentration time series sequence of each gas and the preliminary abnormality degree of the concentration time series sequence of each gas is used as the detection influence weight of each segmented sequence included in the concentration time series sequence of each gas; The detection influence weight of each segmentation sequence is used as the sample weight of the concentration data in each segmentation sequence; Determine the number of samples for training isolated trees, calculate the ratio of the sample weight of each sample for training each isolated tree to the sum of the sample weights of all concentration data in the segmentation sequence where each sample is located, and use the normalized result of the cumulative sum of the ratio on the number of samples for training isolated trees as the decision weight of each isolated tree.

Citation Information

Patent Citations

  • Data center fault detection system based on machine learning

    CN116955091A

  • Water quality index prediction method based on deep learning

    CN117522632A