SCR denitrification condition recognition method based on data preprocessing and sequence similarity judgment

By using data preprocessing and sequence similarity judgment methods in the steel production process, the problem of data quality problems and low accuracy in working conditions recognition is solved, data accuracy and recognition stability are improved, and the rational utilization of resources and cost optimization control are achieved.

CN119202543BActive Publication Date: 2025-05-06NANJING KUNLUN ENERGY STORAGE TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411681584.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-05-06
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively judge data quality problems and working conditions identification in the steel production process, which has low accuracy, resulting in waste of resources and increased costs.

Method used

The SCR denitrification condition recognition method based on data preprocessing and sequence similarity judgment is adopted, and data quality and recognition accuracy are improved by acquiring time sequence data, smoothing processing, preprocessing, sequence abnormality detection and similarity judgment.

Benefits of technology

Through data preprocessing and sequence similarity judgment, the accuracy and stability of data and identification similarity accuracy and stability are improved, and resource waste and cost increase are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202543B_ABST
    Figure CN119202543B_ABST
Patent Text Reader

Abstract

This invention proposes a method for identifying SCR denitrification operating conditions based on data preprocessing and sequence similarity judgment, comprising the following steps: S1: acquiring a time-series dataset; S2: smoothing the data in the time-series dataset; S3: preprocessing the data in the time-series dataset; S4: detecting sequence anomalies in the data in the time-series dataset; S5: judging the similarity of the time-series data sequences in the time-series dataset. By smoothing, preprocessing, and detecting sequence anomalies in the time-series data collected by the equipment, the accuracy of the collected time-series data is improved through judgment and quality monitoring. By creating a cumulative distance matrix, calculating the distance between elements in two time-series data sequences, filling the distance into the cumulative distance matrix, obtaining a regularized path based on the cumulative distance matrix, and judging the similarity of the time-series data sequences and whether they conform to the operating conditions through the distance of the regularized path, the similarity accuracy and stability of the operating condition identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an SCR denitration condition recognition method based on data preprocessing and sequence similarity judgment. Background Art

[0002] During the steel production process, the flue gas contains a large amount of pollutants, such as nitrogen oxides ( )、sulfur dioxide( ) and particulate matter; these pollutants pose a potential threat to air quality and the ecological environment; with the improvement of environmental protection and sustainable development awareness, as well as the formulation and implementation of relevant laws and regulations, steel companies have adopted a series of flue gas treatment measures, such as wet desulfurization, SCR denitrification, dust removal and other technologies; however, due to the complex working conditions of the production line and the instability of flue gas, steel companies usually use artificial control methods to consume excessive materials and energy to ensure that flue gas emissions can strictly meet the standards, resulting in waste of resources and increased costs;

[0003] To this end, some digital companies have actively innovated and developed a series of advanced industrial control algorithms. First, they use intelligent edge devices to read and collect data in the DCS / PLC system in real time to ensure the accuracy and timeliness of the data. Then, these algorithms perform in-depth analysis and precise calculations on the edge devices to obtain the most optimized control instructions. These instructions are then quickly transmitted to the equipment terminal through the DCS / PLC system, thereby achieving efficient regulation of the production line. Through this intelligent approach, steel companies can not only ensure that flue gas emissions are strictly met, but also significantly reduce excessive consumption of materials and energy, and achieve rational use of resources and optimized cost control.

[0004] Although industrial control algorithms have made significant progress and can replace manual optimization control to a certain extent, they still face many challenges in the actual application scenarios of steel enterprises. Due to the aging of steel plant equipment, relatively high failure rate, and complex and changeable production line conditions, it is difficult for algorithms to respond to various emergencies on site as flexibly as experienced operators. Although algorithms can perform accurate data analysis and calculations, their decision-making ability still has certain limitations when faced with nonlinear and uncertain problems. Specifically, there are the following disadvantages:

[0005] 1. Difficulty in judging data quality issues: Data quality is a key challenge in the process of identifying operating parameters. Due to factors such as sensor accuracy, equipment working environment, and data collection frequency, the collected data may be noisy, missing, or inaccurate. These low-quality data will directly affect the accuracy and stability of identification. However, there is currently a lack of effective methods to judge and monitor data quality in real time, making it difficult to ensure the accuracy, completeness, and timeliness of data.

[0006] 2. The accuracy of similarity in working condition identification is low: The existing working condition identification methods have limited accuracy and stability when facing complex and changeable production environments. Currently commonly used similarity measurement methods, such as correlation coefficient and Euclidean distance, have great limitations. For example, the correlation coefficient may not be accurate enough in some cases, and the Euclidean distance is easily affected by factors such as noise, resulting in unreliable identification results. Summary of the invention

[0007] The purpose of the present invention is to solve the shortcomings existing in the prior art and to propose an SCR denitration condition identification method based on data preprocessing and sequence similarity judgment.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The SCR denitration condition recognition method based on data preprocessing and sequence similarity judgment includes the following steps:

[0010] S1: Get time series data set;

[0011] Collecting time series data of each collection point of the on-site SCR denitration process, and saving the collected time series data into a time series data set; the SCR refers to selective catalytic reduction technology;

[0012] Each time series data in the time series data set is recorded as ( , );in, For time point, is the corresponding time series data value;

[0013] The time series data includes inlet CEMS related data, outlet CEMS related data, hot blast furnace related data, other process parameters, etc.; the CEMS refers to a continuous emission monitoring system;

[0014] The inlet CEMS related data includes the inlet CEMS flue gas content, inlet CEMS flue gas flow, inlet CEMS flue gas dust content, inlet CEMS flue gas temperature, inlet CEMS flue gas Content, outlet CEMS flue gas Content, etc. It refers to nitrogen oxides, It refers to oxygen;

[0015] The export CEMS related data includes export CEMS flue gas content, outlet CEMS flue gas flow, outlet CEMS flue gas dust content, etc.;

[0016] The hot blast stove related data include hot blast stove back-end temperature, hot blast stove furnace temperature, catalyst inlet temperature, etc.;

[0017] The other process parameters include the GGH original flue gas inlet pressure, the opening of the No. 1 main exhaust door, the opening of the No. 2 main exhaust door, etc.; the GGH refers to the flue gas heat exchanger.

[0018] S2: Smoothing the data in the time series data set;

[0019] The time series data in the time series data set are smoothed by the LOWESS algorithm, which is a non-parametric smoothing method that calculates the smoothing value of each data point by performing local regression on the data point;

[0020] The following sub-steps are included:

[0021] S21: Select smoothing parameters;

[0022] The engineer manually selects a smoothing parameter such as a window size h and sets a specific value of the window size h, where the window size h is used to specify the range of neighboring data points for performing local weighted regression near each time series data;

[0023] S22: Select local data points;

[0024] The time series data in the known time series data set is ( , ), and other data points are recorded as ( , );

[0025] For the time series data in the time series dataset ( , ), and then calculate with other data points ( , ), and the time difference is recorded as , the time difference ;

[0026] Filter out all time differences The corresponding other data points smaller than the window size h are regarded as local data points;

[0027] S23: Calculate the predicted value;

[0028] Use the weight function W(x) to calculate the time series data in sequence ( , ) and the corresponding local data point ( , )’s weight;

[0029] A weighted regression model is constructed according to the obtained weights, and the predicted value of each time series data in the time series data set is calculated in sequence through the weighted regression model; and the original value of the corresponding time series data in the time series data set is replaced by the predicted value;

[0030] The method of constructing a weighted regression model based on weights is a prior art, and the present invention merely applies this technology without making any innovation on this technology.

[0031] S3: preprocess the data in the time series dataset;

[0032] Filter unreasonable data in the time series data set by filtering the upper and lower limits of the slope and the upper and lower limits of the threshold of each time series data in the time series data set;

[0033] The following sub-steps are included:

[0034] S31: Determine whether the slope of the time series data is abnormal;

[0035] Set the upper and lower limits of the slope of the time series data in the time series data set;

[0036] For each time series data in the time series data set, calculate the slope between adjacent time points and record the slope as , ;

[0037] Will Compare with the set upper and lower limits of slope respectively. If If it is greater than the upper limit of the slope or less than the lower limit of the slope, the slope is abnormal;

[0038] if If the slope is between the upper and lower limits, the slope is normal.

[0039] S32: Remove unreasonable data from the time series data set;

[0040] In step S31, the time series data corresponding to the abnormal slope is unreasonable data, and the unreasonable data in the time series data set is deleted to complete the preprocessing of the time series data.

[0041] S4: Perform sequence anomaly detection on the data in the time series dataset;

[0042] Traversing the time series data in the time series data set by using the PersistAD algorithm to identify outliers, wherein the PersistAD algorithm is an adaptive algorithm based on data persistence and is used for optimization in data storage, processing, and transmission systems;

[0043] The following sub-steps are included:

[0044] S41: performing normalization processing;

[0045] The time series data in the time series data set are normalized by the minimum-maximum normalization method;

[0046] Specifically, the maximum and minimum values ​​corresponding to each time series data are found from the time series data set; and the time series data are normalized using the minimum-maximum normalization formula, which is as follows:

[0047] ;

[0048] in, is the normalized time series data, y is the original time series data, max(Y) is the maximum value of this type of time series data in the time series data set, and min(Y) is the minimum value of this type of time series data in the time series data set;

[0049] S42: Calculate the average value of the time series data within the time window;

[0050] Each type of time series data in the time series data set corresponds to a time series data sequence; a time window is set, and the time window size is window, and window is a positive integer;

[0051] Traverse each time series data sequence in the time series data set in turn according to the time window;

[0052] Specifically, for each time series data sequence in the time series data set, for each time point t, calculate the average value of the time series data sequence in the previous time window, that is, for the time series data sequence in the time series data set, look forward to window time series data from the current time point, and calculate the average value of this window time series data;

[0053] For example: Assume the time window is 5, and the time series data is the inlet CEMS flue gas Content, for each time point t, looking forward from the current time point to the CEMS flue gas The five time series data in the content time series data sequence are time windows. The five CEMS flue gas The average value of the content;

[0054] S43: Time series data sequence anomaly detection;

[0055] Set the threshold d according to actual needs;

[0056] For each time series data sequence in the time series data set, calculate the absolute value of the difference between the average value of the time series data corresponding to the current time point t and the average value of the time series data sequence in the time window corresponding to the previous time point. If the absolute value of the difference between the two average values ​​is greater than the threshold d, the time series data sequence is marked as abnormal at the current time point t; repeat the above steps until each time series data sequence in the time series data set is detected.

[0057] S5: Determine the similarity of the time series sequences in the time series dataset;

[0058] It includes the following sub-steps:

[0059] S51: Create a cumulative distance matrix;

[0060] Denote any two time series sequences in the time series dataset as time series A and time series B;

[0061] Time series , time series ;

[0062] Wherein, , ;

[0063] is the time point corresponding to the time series data in time series A, is the time series data value corresponding to the time series data ;

[0064] is the time point corresponding to the time series data in time series B, is the time series data value corresponding to the time series data ;

[0065] Create an n*m cumulative distance matrix D, where D(i, j) represents the minimum cumulative distance between the first i elements of time series A and the first j elements of time series B;

[0066] S52: Calculate the distance between elements;

[0067] Calculate the distance between any two elements in time series A and time series B in turn through the Euclidean distance formula, denoted as ;

[0068] ;

[0069] S53: Fill the cumulative distance matrix;

[0070] Save the distance calculated in step S52 to the cumulative distance matrix D;

[0071] Specifically, initialize the cumulative distance matrix, ;

[0072] For the elements in the first row and the first column, the value of the cumulative distance matrix is:

[0073] When 1 ≤ i < n, ;

[0074] When 1 ≤ j < m, ;

[0075] For other elements, the value of the cumulative distance matrix is:

[0076] ;

[0077] S54: Perform sequence similarity judgment;

[0078] Starting from the lower right corner D(m, n) of the matrix and backtracking to the upper left corner D(0, 0), the obtained warping path distance represents the similarity of the two time series. The greater the distance, the lower the similarity, and the smaller the distance, the higher the similarity;

[0079] Specifically, the form of the optimal warping path is ;

[0080] where the form of w is (i, j), i is the coordinate in time series A, and j is the coordinate in time series B; ; |A| and |B| respectively represent the number of rows of matrix A and the number of columns of matrix B; Starting from (1, 1) and ending at (|A|, |B|).

[0081] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0082] The SCR denitration working condition identification method based on data preprocessing and sequence similarity judgment proposed by the present invention performs smoothing processing, preprocessing, and sequence anomaly detection on the time series data collected by the device, judges and monitors the quality of the collected time series data, improves the accuracy of the data, and avoids problems such as missing and inaccurate time series data in the prior art;

[0083] The present invention creates a cumulative distance matrix, calculates the distances between elements in two time series data sequences in sequence, fills the calculated distances into the cumulative distance matrix, obtains a warping path based on the cumulative distance matrix, and judges the similarity of the time series data sequences and whether they conform to the working conditions through the warping path distance; improves the similarity accuracy and stability of working condition identification. Description of the Drawings

[0084] Figure 1 It is a step flow chart of the SCR denitration working condition identification method based on data preprocessing and sequence similarity judgment of the present invention. Detailed Embodiments

[0085] To further understand the purpose, structure, features, and functions of the present invention, the following is a detailed description in conjunction with the embodiments.

[0086] like Figure 1 As shown, the SCR denitration condition identification method based on data preprocessing and sequence similarity judgment includes the following steps:

[0087] S1: Get time series data set;

[0088] Collect the time series data of each collection point of the on-site SCR denitrification process, and save the collected time series data into the time series data set;

[0089] Each time series data in the time series data set is recorded as ( , );in, For time point, is the corresponding time series data value;

[0090] The time series data includes inlet CEMS related data, outlet CEMS related data, hot blast furnace related data, other process parameters, etc.;

[0091] The inlet CEMS related data includes the inlet CEMS flue gas content, inlet CEMS flue gas flow, inlet CEMS flue gas dust content, inlet CEMS flue gas temperature, inlet CEMS flue gas Content, outlet CEMS flue gas Content, etc.

[0092] The export CEMS related data includes export CEMS flue gas content, outlet CEMS flue gas flow, outlet CEMS flue gas dust content, etc.;

[0093] The hot blast stove related data include hot blast stove back-end temperature, hot blast stove furnace temperature, catalyst inlet temperature, etc.;

[0094] The other process parameters include the GGH raw flue gas inlet pressure, the opening of the No. 1 main exhaust door, the opening of the No. 2 main exhaust door, etc.

[0095] S2: Smoothing the data in the time series data set;

[0096] The time series data in the time series data set is smoothed using the LOWESS algorithm, which includes the following sub-steps:

[0097] S21: Select smoothing parameters;

[0098] The engineer manually selects a smoothing parameter such as a window size h and sets a specific value of the window size h, where the window size h is used to specify the range of neighboring data points for performing local weighted regression near each time series data;

[0099] S22: Select local data points;

[0100] The time series data in the known time series data set is ( , ), and other data points are recorded as ( , );

[0101] For the time series data in the time series dataset ( , ), and then calculate with other data points ( , ), and the time difference is recorded as , the time difference ;

[0102] Filter out all time differences The corresponding other data points smaller than the window size h are regarded as local data points;

[0103] S23: Calculate the predicted value;

[0104] Use the weight function W(x) to calculate the time series data in sequence ( , ) and the corresponding local data point ( , )’s weight;

[0105] The weight function W(x) satisfies the following conditions:

[0106] (1) When |x|<1, W(x)>0;

[0107] (2) W(-x) = W(x), that is, W(x) is a symmetric function;

[0108] (3) When x>= 0, W(x) is a non-increasing function;

[0109] (4) When |x|>= 1, W(x) = 0;

[0110] A weighted regression model is constructed according to the obtained weights, and the predicted value of each time series data in the time series data set is calculated in sequence through the weighted regression model; and the original value of the corresponding time series data in the time series data set is replaced by the predicted value;

[0111] The method of constructing a weighted regression model based on weights is a prior art, and the present invention merely applies this technology without making any innovation on this technology.

[0112] By smoothing the time series data and using the prediction of the weighted regression model to replace the original data, the noise can be effectively reduced, the prediction accuracy can be improved, and the stability of the model can be enhanced.

[0113] S3: preprocess the data in the time series dataset;

[0114] Filter unreasonable data in the time series data set by filtering the upper and lower limits of the slope and the upper and lower limits of the threshold of each time series data in the time series data set;

[0115] The following sub-steps are included:

[0116] S31: Determine whether the slope of the time series data is abnormal;

[0117] Set the upper and lower limits of the slope of the time series data in the time series data set;

[0118] For each time series data in the time series data set, calculate the slope between adjacent time points and record the slope as , ;

[0119] Will Compare with the set upper and lower limits of slope respectively. If If it is greater than the upper limit of the slope or less than the lower limit of the slope, the slope is abnormal;

[0120] if If the slope is between the upper and lower limits, the slope is normal.

[0121] S32: Remove unreasonable data from the time series data set;

[0122] In step S31, the time series data corresponding to the abnormal slope is unreasonable data, and the unreasonable data in the time series data set is deleted to complete the preprocessing of the time series data.

[0123] By judging the slope anomaly of time series data and removing unreasonable data, the quality of data can be significantly improved, the interference of anomalies on the model can be reduced, and the accuracy and stability of the model can be enhanced.

[0124] S4: Perform sequence anomaly detection on the data in the time series dataset;

[0125] The PersistAD algorithm is used to traverse the time series data in the time series data set to identify outliers, including the following sub-steps:

[0126] S41: performing normalization processing;

[0127] The time series data in the time series data set are normalized by the minimum-maximum normalization method;

[0128] Specifically, the maximum and minimum values ​​corresponding to each time series data are found from the time series data set; and the time series data are normalized using the minimum-maximum normalization formula, which is as follows:

[0129] ;

[0130] in, is the normalized time series data, y is the original time series data, max(Y) is the maximum value of this type of time series data in the time series data set, and min(Y) is the minimum value of this type of time series data in the time series data set;

[0131] S42: Calculate the average value of the time series data within the time window;

[0132] Each type of time series data in the time series data set corresponds to a time series data sequence; a time window is set, and the time window size is window, and window is a positive integer;

[0133] Traverse each time series data sequence in the time series data set in turn according to the time window;

[0134] Specifically, for each time series data sequence in the time series data set, for each time point t, calculate the average value of the time series data sequence in the previous time window, that is, for the time series data sequence in the time series data set, look forward to window time series data from the current time point, and calculate the average value of this window time series data;

[0135] For example: Assume the time window is 5, and the time series data is the inlet CEMS flue gas Content, for each time point t, looking forward from the current time point to the CEMS flue gas The five time series data in the content time series data sequence are time windows. The five CEMS flue gas The average value of the content;

[0136] S43: Time series data sequence anomaly detection;

[0137] Set the threshold d according to actual needs;

[0138] For each time series data sequence in the time series data set, calculate the absolute value of the difference between the average value of the time series data corresponding to the current time point t and the average value of the time series data sequence in the time window corresponding to the previous time point. If the absolute value of the difference between the two average values ​​is greater than the threshold d, the time series data sequence is marked as abnormal at the current time point t; repeat the above steps until each time series data sequence in the time series data set is detected.

[0139] S5: perform similarity judgment on the time series data sequences in the time series data set;

[0140] The following sub-steps are included:

[0141] S51: Create cumulative distance matrix;

[0142] Record any two time series data sequences in the time series dataset as time series A and time series B;

[0143] Time series , time series ;

[0144] Among them, , ;

[0145] In time series A, is the time point corresponding to the time series data is the time series data value corresponding to the time series data ;

[0146] In time series B, is the time point corresponding to the time series data is the time series data value corresponding to the time series data ;

[0147] Create an n*m cumulative distance matrix D, where D(i, j) represents the minimum cumulative distance between the first i elements of time series A and the first j elements of time series B;

[0148] S52: Calculate the distance between elements;

[0149] Calculate the distance between any two elements in time series A and time series B in turn through the Euclidean distance formula, denoted as ;

[0150] ;

[0151] S53: Fill in the cumulative distance matrix;

[0152] Save the distance calculated in step S52 to the cumulative distance matrix D;

[0153] Specifically, initialize the cumulative distance matrix, ;

[0154] For the elements in the first row and the first column, the value of the cumulative distance matrix is:

[0155] When 1 ≤ i < n, ;

[0156] When 1 ≤ j < m, ;

[0157] For other elements, the value of the cumulative distance matrix is:

[0158] ;

[0159] S54: perform sequence similarity determination;

[0160] Starting from the lower right corner of the matrix D(m, n), and tracing back to the upper left corner D(0, 0), the final regular path distance is expressed as the similarity of the two time series. The larger the distance, the lower the similarity, and the smaller the distance, the higher the similarity.

[0161] Specifically, the form of the optimal regularized path is ;

[0162] Among them, w is in the form of (i, j), i is the coordinate in time series A, and j is the coordinate in time series B; ; |A| and |B| represent the number of rows of matrix A and the number of columns of matrix B respectively; Starting from (1, 1), to Ends for (|A|, |B|);

[0163] Furthermore, h groups of time series data sequences with high similarity are marked and saved in the time series sample data set. The corresponding normalized path distance mean avg is calculated in sequence. The threshold coefficient l is set according to the specific situation. When a new data set is input, the normalized path is calculated with the h groups of time series sample data sets. When there is a corresponding set of regularized path distances , it is determined that this data set meets the working conditions.

[0164] The present invention has been described by the above-mentioned relevant embodiments, however, the above-mentioned embodiments are only examples for implementing the present invention. It must be pointed out that the disclosed embodiments do not limit the scope of the present invention. On the contrary, changes and modifications made without departing from the spirit and scope of the present invention are all within the scope of patent protection of the present invention.

Claims

1. A method for identifying SCR denitration conditions based on data preprocessing and sequence similarity judgment, characterized in that: The following steps are involved: S1: Get time series data set; Collect the time series data of each collection point of the on-site SCR denitrification process, and save the collected time series data into the time series data set; Each time series data in the time series data set is recorded as ( , );in, For time point, is the corresponding time series data value; S2: Smoothing the data in the time series data set; The time series data in the time series data set is smoothed using the LOWESS algorithm, which includes the following sub-steps: S21: Select smoothing parameters; S22: Select local data points; S23: Calculate the predicted value; S3: preprocess the data in the time series dataset; Filter unreasonable data in the time series data set by filtering the upper and lower limits of the slope and the upper and lower limits of the threshold of each time series data in the time series data set; The following sub-steps are included: S31: Determine whether the slope of the time series data is abnormal; S32: Remove unreasonable data from the time series data set; S4: Perform sequence anomaly detection on the data in the time series dataset; The PersistAD algorithm is used to traverse the time series data in the time series data set to identify outliers, including the following sub-steps: S41: performing normalization processing; S42: Calculate the average value of the time series data within the time window; S43: Time series data sequence anomaly detection; S5: perform similarity judgment on the time series data sequences in the time series data set; The following sub-steps are included: S51: Create cumulative distance matrix; S52: Calculate the distance between elements; S53: Fill in the cumulative distance matrix; S54: Perform sequence similarity determination.

2. The SCR denitration operating condition identification method based on data preprocessing and sequence similarity judgment according to claim 1 is characterized in that: The specific content of step S2 is as follows: S21: Select smoothing parameters; The engineer manually selects a smoothing parameter, namely, a window size h, and sets a specific value of the window size h, where the window size h is used to specify the range of neighboring data points for performing local weighted regression near each time series data; S22: Select local data points; The time series data in the known time series data set is ( , ), and other data points are recorded as ( , ); For the time series data in the time series dataset ( , ), and then calculate with other data points ( , ), and the time difference is recorded as , the time difference ; Filter out all time differences The corresponding other data points smaller than the window size h are regarded as local data points; S23: Calculate the predicted value; Use the weight function W(x) to calculate the time series data in sequence ( , ) and the corresponding local data point ( , )’s weight; A weighted regression model is constructed based on the obtained weights, and the predicted value of each time series data in the time series data set is calculated in sequence through the weighted regression model; And replace the original value of the corresponding time series data in the time series dataset with the predicted value.

3. The SCR denitration operating condition identification method based on data preprocessing and sequence similarity judgment according to claim 1 is characterized in that: The specific contents of step S3 are as follows: S31: Determine whether the slope of the time series data is abnormal; Set the upper and lower limits of the slope of the time series data in the time series data set; For each time series data in the time series data set, calculate the slope between adjacent time points and record the slope as , ; Will Compare with the set upper and lower limits of slope respectively. If If it is greater than the upper limit of the slope or less than the lower limit of the slope, the slope is abnormal; if If the slope is between the upper and lower limits, the slope is normal. S32: Remove unreasonable data from the time series data set; In step S31, the time series data corresponding to the abnormal slope is unreasonable data, and the unreasonable data in the time series data set is deleted to complete the preprocessing of the time series data.

4. The SCR denitration operating condition identification method based on data preprocessing and sequence similarity judgment according to claim 1 is characterized in that: The specific contents of step S4 are as follows: S41: performing normalization processing; The time series data in the time series data set are normalized by the minimum-maximum normalization method; Specifically, the maximum and minimum values ​​corresponding to each time series data are found from the time series data set; and the time series data are normalized using the minimum-maximum normalization formula, which is as follows: ; in, is the normalized time series data, y is the original time series data, max(Y) is the maximum value of this type of time series data in the time series data set, and min(Y) is the minimum value of this type of time series data in the time series data set; S42: Calculate the average value of the time series data within the time window; Each type of time series data in the time series data set corresponds to a time series data sequence; a time window is set, and the time window size is window, and window is a positive integer; Traverse each time series data sequence in the time series data set in turn according to the time window; Specifically, for each time series data sequence in the time series data set, for each time point t, calculate the average value of the time series data sequence in the previous time window, that is, for the time series data sequence in the time series data set, look forward to window time series data from the current time point, and calculate the average value of this window time series data; S43: Time series data sequence anomaly detection; Set the threshold d according to actual needs; For each time series data sequence in the time series data set, calculate the absolute value of the difference between the average value of the time series data corresponding to the current time point t and the average value of the time series data sequence in the time window corresponding to the previous time point. If the absolute value of the difference between the two average values ​​is greater than the threshold d, the time series data sequence is marked as abnormal at the current time point t; repeat the above steps until each time series data sequence in the time series data set is detected.

5. The SCR denitration operating condition identification method based on data preprocessing and sequence similarity judgment according to claim 1 is characterized in that: The specific content of step S5 is as follows: S51: Create cumulative distance matrix; Any two time series data sequences in the time series data set are recorded as time series A and time series B; Time Series , time series ; in, , ; For time series A, the time series data The corresponding time point, For time series data The corresponding time series data value; For time series B, the time series data The corresponding time point, For time series data The corresponding time series data value; Create an n*m cumulative distance matrix D, where D(i, j) represents the minimum cumulative distance between the first i elements of time series A and the first j elements of time series B; S52: Calculate the distance between elements; The distance between any two elements in time series A and time series B is calculated in sequence using the Euclidean distance formula, denoted as ; ; S53: Fill in the cumulative distance matrix; The distance calculated in step S52 is saved in the cumulative distance matrix D; Specifically, initialize the cumulative distance matrix, ; For the elements in the first row and first column, the value of the cumulative distance matrix is: When 1 ≤ i < n, ; When 1 ≤ j < m, ; For other elements, the value of the cumulative distance matrix is: ; S54: perform sequence similarity determination; Starting from the lower right corner of the matrix D(m, n) and tracing back to the upper left corner D(0, 0), the final regular path distance is expressed as the similarity of the two time series. The larger the distance, the lower the similarity, and the smaller the distance, the higher the similarity.

6. The SCR denitration operating condition identification method based on data preprocessing and sequence similarity judgment according to claim 5 is characterized in that: Specifically, the form of the optimal regularized path is ; Among them, w is in the form of (i, j), i is the coordinate in time series A, and j is the coordinate in time series B; ; |A| and |B| represent the number of rows of matrix A and the number of columns of matrix B respectively; Starting from (1, 1), to Ends for (|A|, |B|); Furthermore, h groups of time series data sequences with high similarity are marked and saved in the time series sample data set. The corresponding normalized path distance mean avg is calculated in sequence. The threshold coefficient l is set according to the specific situation. When a new data set is input, the normalized path is calculated with the h groups of time series sample data sets. When there is a corresponding set of regularized path distances , it is determined that this data set meets the working conditions.

Citation Information

Patent Citations

  • A fast searching method for time series stream data based on data characteristics

    CN109325060A

  • Landslide universal surface displacement monitoring data missing and abnormal value processing method

    CN112883075A