Multi-source coking time-series data splicing method based on similarity matching and dynamic fusion
By employing a similarity matching and dynamic fusion method, the problems of time offset and nonlinear distortion of multi-source time-series data in continuous coking production are solved, achieving high-precision, robust, and efficient data stitching, which is suitable for continuous coking production scenarios.
Patent Information
- Application Number
- CN202511167980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-20
AI Technical Summary
In the continuous coking process, the inconsistent time axes of multi-source time-series data, lack of process constraints, contradiction between computational efficiency and accuracy, and poor noise robustness lead to insufficient data splicing accuracy and high computational complexity, making it difficult to meet the needs of real-time processing.
A method based on similarity matching and dynamic fusion is adopted. Through data preprocessing, coarse matching, fine alignment and overlapping region fusion, combined with coking process characteristics, weighted Euclidean distance and DTW algorithm are used to select and align high similarity candidate time series vectors, and dynamic weighted fusion is used to process overlapping regions.
It significantly improves the accuracy and robustness of data splicing, reduces mean square error, increases processing speed, ensures that the alignment results conform to production logic, enhances the ability to resist noise interference, and is suitable for real-time production scenarios.
Smart Images

Figure CN120724394B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of continuous coking production data processing technology, specifically a method for stitching together multi-source coking time-series data based on similarity matching and dynamic fusion. Background Technology
[0002] In the continuous coking process, accurate time-series data splicing is crucial for process optimization and quality control. However, existing technologies have the following problems when processing coking time-series data: (1) Inconsistent time axis of multi-source data: Continuous coking involves multi-source sensor data such as furnace temperature, pressure, and coke pushing current. Due to the asynchronous sampling frequency of the equipment (e.g., furnace temperature sensor is 1Hz, pressure sensor is 0.5Hz), the data time axis is offset, and traditional similarity matching methods are difficult to handle this nonlinear time distortion. (2) Lack of process constraints: Continuous coking has strict process rules (e.g., temperature rises monotonically during coking), but existing DTW methods lack process constraints during alignment, which may generate alignment paths that do not conform to the production logic, such as incorrectly matching the heating stage with the cooling stage. (3) Contradiction between computational efficiency and accuracy: Although the pure DTW method can handle time offsets, its computational complexity is high (O(N²)), making it difficult to process large-scale coking time-series data (e.g., furnace temperature data for 72 consecutive hours) in real time; while the simple similarity matching method is efficient, its alignment accuracy is insufficient. (4) Poor noise robustness: The coking environment is harsh, and sensor data often contains impulse noise (such as instantaneous abnormal values caused by equipment vibration). Existing methods are easily affected by noise, leading to misjudgment of similarity and alignment deviation.
[0003] For example, patent CN202111633157.6 discloses a method for structured storage of coking data by production batch, which facilitates data management but fails to address the time alignment issue across batches. The similarity-based stitching method proposed in the literature "A Stitch in Time-Series" does not consider the strong process constraints of continuous coking production and lacks accuracy when processing highly nonlinear sequences such as furnace temperature. Therefore, there is an urgent need for a time-series data stitching method that combines the advantages of both methods and is adapted to the characteristics of the coking process. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a method for stitching multi-source coking time-series data based on similarity matching and dynamic fusion, which realizes high-precision integration of cross-batch, multi-source coking time-series data, significantly improves the accuracy and robustness of data stitching, provides high-quality data support for downstream applications such as coke quality prediction and process optimization, and is suitable for continuous coking production scenarios.
[0005] The technical solution of this invention is as follows:
[0006] The method for stitching together multi-source coking time-series data based on similarity matching and dynamic fusion includes the following steps:
[0007] (1) Data acquisition and preprocessing: Acquire the baseline time series vector and multiple candidate time series vectors and perform data preprocessing to obtain the preprocessed baseline time series vector and multiple candidate time series vectors; the baseline time series vector is the target coking time series vector to be analyzed, and the multiple candidate time series vectors are multiple historical coking time series vectors;
[0008] (2) Coarse matching: Calculate the volatility characteristics, level characteristics and coking process characteristics of the baseline time series vector and each candidate time series vector, construct the corresponding multi-dimensional feature vector, and then perform similarity matching between each candidate time series vector and the multi-dimensional feature vector corresponding to the baseline time series vector based on the weighted Euclidean distance, so as to select multiple highly similar candidate time series vectors from multiple candidate time series vectors.
[0009] (3) Fine alignment: Under process constraints, the optimal alignment path between each highly similar candidate time series vector and the reference time series vector is calculated and selected by the DTW algorithm. Then, the optimal alignment path with the shortest path length is selected from multiple optimal alignment paths. The highly similar candidate time series vector corresponding to the optimal alignment path with the shortest path length is the optimal candidate time series vector, thus realizing fine time alignment between the candidate time series vector and the reference time series vector.
[0010] (4) Overlapping region fusion and splicing: For the best candidate time series vector and the reference time series vector after fine alignment, the overlapping regions of the two are fused using a dynamic weighted fusion strategy, and the non-overlapping regions are directly spliced to generate a complete spliced time series vector.
[0011] The aforementioned baseline time series vector and each candidate time series vector include time series data of multiple coking features, including coking features of the coal blending section, coke oven section, gas section, and methanol section.
[0012] The coking characteristics of the coal blending section include the moisture content, volatile matter, sulfur content, ash content, caking index, and number of coals charged in the blended coal.
[0013] The coking characteristics of the coke oven section include the temperature of coke oven gas before preheating, the temperature of coke oven gas after preheating, the pressure at the inlet of the gas collecting pipe, the pressure at the outlet of the gas collecting pipe, the suction force on the fan side of the flue, the suction force on the coke side of the flue, the temperature on the fan side of the flue, the temperature on the coke side of the flue, and the coking time.
[0014] The coking characteristics of the gas section include the resistance of the primary cooler, the average gas temperature after the primary cooler, the resistance of the electrostatic precipitator, the gas pressure after the mist precipitator, the average oil temperature of the blower, the average oil pressure of the blower, the average gas pressure after the desulfurization tower, the gas temperature after the benzene washing tower, the gas pressure after the benzene washing tower, and the lean oil temperature of the benzene washing tower.
[0015] The coking characteristics of the methanol section include coal gas consumption, R104 catalyst bed temperature, compressor low-pressure cylinder inlet gas flow rate, compressor circulation section inlet flow rate, converter inlet coke gas temperature, converter coke oven gas and outlet converter gas pressure difference, converter outlet converter gas temperature, RT synthesis tower inlet temperature, and RP synthesis tower pressure difference.
[0016] The data preprocessing specifically includes data standardization, outlier removal, and missing value imputation.
[0017] The data standardization process described above uses Z-score standardization, and the calculation process is shown in the following formula (1):
[0018] (1);
[0019] In equation (1), Represents data after Z-score standardization. This represents the original data before Z-score standardization. The mean of the time series data for coking characteristics. The standard deviation of the time series data for coking characteristics;
[0020] The method for removing outliers is as follows: The missing values are filled using cubic spline interpolation.
[0021] The calculation process for the aforementioned volatility characteristics is shown in the following formula (2):
[0022] (2);
[0023] In equation (2), Representing volatility characteristics, each candidate time series vector or base time series vector contains multiple volatility characteristics corresponding to charring features in its multi-dimensional feature vector. This represents the eigenvalue at time j in the preprocessed baseline or candidate time series vector. This represents the eigenvalue at time j+1 in the preprocessed baseline or candidate time series vector. This represents the number of features used to calculate the volatility characteristics in a specific coking characteristic time series data within the preprocessed baseline or candidate time series vector.
[0024] The calculation process of the aforementioned level value feature is shown in the following formula (3):
[0025] (3);
[0026] In equation (3), The level value features represent the multi-dimensional feature vector corresponding to each candidate time series vector or baseline time series vector, which includes multiple level value features corresponding to charring features. This represents the eigenvalue at time j in the preprocessed baseline or candidate time series vector. This represents the number of features used to calculate the level value feature in a certain coking characteristic time series data within the preprocessed baseline or candidate time series vector.
[0027] The coking process characteristics include coking stage identification, furnace hole number matching degree, and coke pushing time deviation rate.
[0028] The weighted Euclidean distance is used to perform similarity matching between each candidate time series vector and the multi-dimensional feature vector corresponding to the baseline time series vector, as shown in the following equation (4):
[0029] (4);
[0030] In equation (4), Represents the weighted Euclidean distance. This represents the volatility characteristics in the multi-dimensional feature vector corresponding to the baseline time series vector. The volatility characteristics in the multi-dimensional feature vectors corresponding to the candidate time series vectors; The level value features represent the multi-dimensional feature vector corresponding to the baseline time series vector. The level value features in the multi-dimensional feature vector corresponding to the candidate time series vector; The coking process features are represented in the multi-dimensional feature vector corresponding to the baseline time series vector. The coking process features are represented in the multi-dimensional feature vectors corresponding to the candidate time series vectors. , , This represents the weighting coefficient, the value of which is dynamically adjusted according to the importance of coking characteristics in the continuous coking production stage.
[0031] The specific method for selecting multiple highly similar candidate time series vectors from multiple candidate time series vectors is: selecting the corresponding... The values are sorted from smallest to largest, and the top-ranked candidate time series vectors are selected as high-similarity candidate time series vectors.
[0032] The process constraints include time offset constraints, trend consistency constraints, and feature point constraints.
[0033] The time offset constraint is specifically shown in equation (5) below:
[0034] (5);
[0035] In equation (5), Represents any moment in the reference time series vector. Represents the candidate time series vectors with high similarity to The moment when alignment matching is performed. It is the duration of a single coking cycle. The maximum allowed time offset;
[0036] The trend consistency constraint is specifically shown in the following equation (6):
[0037] (6);
[0038] In equation (6), The temperature change rate is represented by the baseline time-series vector. The reference time vector represents the heating stage. The reference time vector represents the cooling stage; The time variable representing the reference time series vector, The time variable representing highly similar candidate time series vectors. The alignment path function represents the alignment path between the highly similar candidate time series vector and the baseline time series vector. Represents the slope of the alignment path;
[0039] The feature point constraint is to force a match between the temperature values of the start and end times of coking.
[0040] The specific steps for calculating and selecting the optimal alignment path using the DTW algorithm under process constraints are as follows:
[0041] a. Construct the local distance matrix D:
[0042] First, calculate the local distance between all time point pairs of the highly similar candidate time series vector and the baseline time series vector, forming a... The matrix D, The data length representing the baseline time series vector, The data length representing the candidate time series vectors with high similarity;
[0043] Points in matrix D The calculation process for the value is shown in the following formula (7):
[0044] (7);
[0045] In equation (7), The eigenvalue representing the i-th time point in the baseline time series vector The feature value of the j-th time point in the highly similar candidate time series vector Differences; Representative eigenvalues and eigenvalues The Euclidean distance;
[0046] b. Initialize the cumulative distance matrix C:
[0047] The cumulative distance matrix C is used to record the distance from the starting point. To each point The minimum cumulative distance is initialized as: The calculation process for the values of each point in the first row and first column of the cumulative distance matrix C is shown in the following formula (8):
[0048] (8);
[0049] In equation (8), This represents the penalty coefficient, with a value ranging from 0.1 to 0.3. and Represents the alignment path. This represents the alignment path corresponding to the first time point of the baseline time series vector and the j-th time point of the highly similar candidate time series vector. The process constraint is violated. When the process constraint condition is violated... Triggering penalty items When the process constraints are met, ; This represents the alignment path corresponding to the i-th time point of the baseline time series vector and the 1-th time point of the highly similar candidate time series vector. The process constraint is violated. When the process constraint condition is violated... Triggering penalty items When the process constraints are met, ;
[0050] c. Fill the cumulative distance matrix C:
[0051] For all points i>1 and j>1 in the cumulative distance matrix C, calculate the minimum cumulative distance using the following formula (9):
[0052] (9);
[0053] In equation (9), This indicates a process constraint violation indication for the upward downward path. When the upward downward path violates process constraints... When the process constraints are met, ; This indicates a violation of process constraints for a left-to-right shift path. When the process constraints are met, ; This indicates a violation of process constraints on the top-left diagonal shift path. When the top-left diagonal shift path violates process constraints... When the process constraints are met, ;
[0054] When calculating the cumulative distance matrix C, each point in the cumulative distance matrix C... The value of not only records the minimum cumulative distance from the starting point to the point, but also implicitly contains the optimal predecessor point to reach the point, that is, the point with the smallest cumulative distance in the three directions of left, top, and top left. By tracing back these predecessor points, the complete optimal alignment path can be restored.
[0055] The dynamic weighted fusion strategy is used for fusion, as shown in the following formula (10):
[0056] (10);
[0057] In equation (10), This represents the value of the i-th feature in the baseline time series vector. It is the value of the i-th feature in the optimal candidate time series vector. This represents the maximum value of the i-th feature in the historical data. This represents the minimum value of the i-th feature in the historical data. Sub-similarity representing the i-th feature; The dynamic weights represent the i-th feature; The similarity score represents the baseline time series vector and the optimal candidate time series vector; The eigenvalues representing the overlapping regions in the baseline time series vector. The feature values representing the overlapping region of the optimal candidate time series vectors. The feature value obtained by weighted fusion of overlapping regions represents the feature value.
[0058] Advantages of this invention:
[0059] (1) Improved accuracy: This invention solves the problems of time offset and nonlinear distortion of continuous coking production data through two-stage processing of coarse matching and fine alignment. The mean square error of the spliced sequence is reduced by 15-20% compared with the single method.
[0060] (2) Process adaptation: This invention introduces process constraints of continuous coking production process to ensure that the alignment result conforms to the production logic and avoids the alignment path from violating the production law. In the splicing of furnace temperature curves, the feature point matching accuracy rate reaches more than 95%.
[0061] (3) Efficiency optimization: This invention uses weighted Euclidean distance for similarity matching, which reduces the computational load of subsequent DTW algorithm for fine alignment. The processing speed is improved by 3-5 times compared to the DTW method, making it suitable for real-time production scenarios.
[0062] (4) Strong robustness: The present invention adopts a dynamic weighted fusion strategy for overlapping regions to enhance the anti-interference ability of noisy data and maintain stable performance even in data containing 10% noise. Attached Figure Description
[0063] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] See Figure 1 The method for stitching together multi-source coking time-series data based on similarity matching and dynamic fusion specifically includes the following steps:
[0066] (1) Data acquisition and preprocessing: Acquire the baseline time series vector and multiple candidate time series vectors and perform data preprocessing to obtain the preprocessed baseline time series vector and multiple candidate time series vectors; the baseline time series vector is the target coking time series vector to be analyzed, and the multiple candidate time series vectors are multiple historical coking time series vectors;
[0067] The aforementioned baseline time series vector and each candidate time series vector include time series data of multiple coking features, including coking features of the coal blending section, coke oven section, gas section, and methanol section.
[0068] The coking characteristics of the coal blending section include the moisture content, volatile matter, sulfur content, ash content, caking index, and number of coals charged in the blended coal.
[0069] The coking characteristics of the coke oven section include the temperature of coke oven gas before preheating, the temperature of coke oven gas after preheating, the pressure at the inlet of the gas collecting pipe, the pressure at the outlet of the gas collecting pipe, the suction force on the fan side of the flue, the suction force on the coke side of the flue, the temperature on the fan side of the flue, the temperature on the coke side of the flue, and the coking time.
[0070] The coking characteristics of the gas section include the resistance of the primary cooler, the average gas temperature after the primary cooler, the resistance of the electrostatic precipitator, the gas pressure after the mist precipitator, the average oil temperature of the blower, the average oil pressure of the blower, the average gas pressure after the desulfurization tower, the gas temperature after the benzene washing tower, the gas pressure after the benzene washing tower, and the lean oil temperature of the benzene washing tower.
[0071] The coking characteristics of the methanol section include coal gas consumption, R104 catalyst bed temperature, compressor low-pressure cylinder inlet gas flow rate, compressor circulation section inlet flow rate, converter inlet coke gas temperature, converter coke oven gas and outlet converter gas pressure difference, converter outlet converter gas temperature, RT synthesis tower inlet temperature, and RP synthesis tower pressure difference.
[0072] Data preprocessing specifically includes data standardization, outlier removal, and missing value imputation;
[0073] Data standardization is performed using Z-score standardization, and the calculation process is shown in the following formula (1):
[0074] (1);
[0075] In equation (1), Represents data after Z-score standardization. This represents the original data before Z-score standardization. The mean of the time series data for coking characteristics. The standard deviation of the time series data for coking characteristics;
[0076] Outlier removal uses Guidelines;
[0077] Missing values were filled using cubic spline interpolation.
[0078] (2) Coarse matching: Calculate the volatility characteristics, level characteristics, and coking process characteristics of the baseline time series vector and each candidate time series vector to construct the corresponding multi-dimensional feature vectors. Then, based on the weighted Euclidean distance, perform similarity matching between each candidate time series vector and the multi-dimensional feature vector corresponding to the baseline time series vector, thereby selecting multiple highly similar candidate time series vectors from multiple candidate time series vectors. The specific steps are as follows:
[0079] S21. Calculate the volatility characteristics, see the following formula (2):
[0080] (2);
[0081] In equation (2), Representing volatility characteristics, each candidate time series vector or base time series vector contains multiple volatility characteristics corresponding to charring features in its multi-dimensional feature vector. This represents the eigenvalue at time j in the preprocessed baseline or candidate time series vector. This represents the eigenvalue at time j+1 in the preprocessed baseline or candidate time series vector. This represents the number of features used to calculate the volatility characteristics in a specific coking characteristic time series data within the preprocessed baseline or candidate time series vector.
[0082] S22. Calculate the characteristics of the level value, as shown in the following formula (3):
[0083] (3);
[0084] In equation (3), The level value features represent the multi-dimensional feature vector corresponding to each candidate time series vector or baseline time series vector, which includes multiple level value features corresponding to charring features. This represents the eigenvalue at time j in the preprocessed baseline or candidate time series vector. This represents the number of features used to calculate the level value feature in a certain coking characteristic time series data within the preprocessed baseline or candidate time series vector.
[0085] S23. Coking process characteristics include coking stage identification, furnace hole number matching degree, and coking time deviation rate.
[0086] S24. The corresponding multi-dimensional feature vectors are constructed by the volatility characteristics, level value characteristics, and coking process characteristics of the baseline time series vector and each candidate time series vector.
[0087] S25. Based on the weighted Euclidean distance, perform similarity matching between each candidate time series vector and the multi-dimensional feature vector corresponding to the baseline time series vector, as shown in the following equation (4):
[0088] (4);
[0089] In equation (4), Represents the weighted Euclidean distance. This represents the volatility characteristics in the multi-dimensional feature vector corresponding to the baseline time series vector. The volatility characteristics in the multi-dimensional feature vectors corresponding to the candidate time series vectors; The level value features represent the multi-dimensional feature vector corresponding to the baseline time series vector. The level value features in the multi-dimensional feature vector corresponding to the candidate time series vector; The coking process features are represented in the multi-dimensional feature vector corresponding to the baseline time series vector. The coking process features are represented in the multi-dimensional feature vectors corresponding to the candidate time series vectors. , , This represents the weighting coefficient, the value of which is dynamically adjusted according to the importance of coking characteristics in the continuous coking production stage.
[0090] Multiple candidate time series vectors corresponding to The values are sorted from smallest to largest, and the top 20 candidate time series vectors are selected as high similarity candidate time series vectors.
[0091] (3) Fine alignment: Under process constraints, the optimal alignment path between each highly similar candidate time series vector and the reference time series vector is calculated and selected using the DTW algorithm. Then, the optimal alignment path with the shortest path length is selected from multiple optimal alignment paths. The highly similar candidate time series vector corresponding to the optimal alignment path with the shortest path length is the optimal candidate time series vector, thus achieving fine time alignment between the candidate time series vector and the reference time series vector. Specifically, it includes the following steps:
[0092] S31. Set process constraints: Process constraints include time offset constraints, trend consistency constraints, and feature point constraints.
[0093] The time offset constraint is shown in equation (5) below:
[0094] (5);
[0095] In equation (5), Represents any moment in the reference time series vector. Represents the candidate time series vectors with high similarity to The moment when alignment matching is performed. It is the duration of a single coking cycle. The maximum allowed time offset;
[0096] The trend consistency constraint is shown in equation (6) below:
[0097] (6);
[0098] In equation (6), The temperature change rate is represented by the baseline time-series vector. The reference time vector represents the heating stage. The reference time vector represents the cooling stage; The time variable representing the reference time series vector, The time variable representing highly similar candidate time series vectors. The alignment path function represents the alignment path between the highly similar candidate time series vector and the baseline time series vector. Represents the slope of the alignment path;
[0099] Feature point constraints are used to force matching of the temperature values at the start and end of coking.
[0100] S32. Construct the local distance matrix D:
[0101] First, calculate the local distance between all time point pairs of the highly similar candidate time series vector and the baseline time series vector, forming a... The matrix D, The data length representing the baseline time series vector, The data length representing the candidate time series vectors with high similarity;
[0102] Points in matrix D The calculation process for the value is shown in the following formula (7):
[0103] (7);
[0104] In equation (7), The eigenvalue representing the i-th time point in the baseline time series vector The feature value of the j-th time point in the highly similar candidate time series vector Differences; Representative eigenvalues and eigenvalues The Euclidean distance;
[0105] S33. Initialize the cumulative distance matrix C:
[0106] The cumulative distance matrix C is used to record the distance from the starting point. To each point The minimum cumulative distance is initialized as: The calculation process for the values of each point in the first row and first column of the cumulative distance matrix C is shown in the following formula (8):
[0107] (8);
[0108] In equation (8), This represents the penalty coefficient, with a value ranging from 0.2. and Represents the alignment path. This represents the alignment path corresponding to the first time point of the baseline time series vector and the j-th time point of the highly similar candidate time series vector. The process constraint is violated. When the process constraint condition is violated... Triggering penalty items When the process constraints are met, ; This represents the alignment path corresponding to the i-th time point of the baseline time series vector and the 1-th time point of the highly similar candidate time series vector. The process constraint is violated. When the process constraint condition is violated... Triggering penalty items When the process constraints are met, ;
[0109] S34, Fill the cumulative distance matrix C:
[0110] For all points i>1 and j>1 in the cumulative distance matrix C, calculate the minimum cumulative distance using the following formula (9):
[0111] (9);
[0112] In equation (9), This indicates a process constraint violation indication for the upward downward path. When the upward downward path violates process constraints... When the process constraints are met, ; This indicates a violation of process constraints for a left-to-right shift path. When the process constraints are met, ; This indicates a violation of process constraints on the top-left diagonal shift path. When the top-left diagonal shift path violates process constraints... When the process constraints are met, ;
[0113] When calculating the cumulative distance matrix C, each point in the cumulative distance matrix C... The value of not only records the minimum cumulative distance from the starting point to the point, but also implicitly contains the optimal predecessor point to reach the point, that is, the point with the smallest cumulative distance in the three directions of left, top, and top left. By tracing back these predecessor points, the complete optimal alignment path can be restored.
[0114] S35. Select the optimal alignment path with the smallest path length (the sum of elements on the optimal alignment path) from the optimal alignment paths of each highly similar candidate time series vector and the reference time series vector. The highly similar candidate time series vector corresponding to the optimal alignment path with the smallest path length is the optimal candidate time series vector, thus realizing fine time alignment between the candidate time series vector and the reference time series vector.
[0115] (4) Overlapping region fusion and splicing: For the best candidate time series vector and the reference time series vector after fine alignment, the overlapping region of the two is fused using a dynamic weighted fusion strategy to generate a complete spliced time series vector;
[0116] A dynamic weighted fusion strategy is adopted for fusion, as shown in the following formula (10):
[0117] (10);
[0118] In equation (10), This represents the value of the i-th feature in the baseline time series vector. It is the value of the i-th feature in the optimal candidate time series vector. This represents the maximum value of the i-th feature in the historical data. This represents the minimum value of the i-th feature in the historical data. Sub-similarity representing the i-th feature; The dynamic weight representing the i-th feature is dynamically adjusted according to the production batch, process stage, or equipment status. For example, in the later stage of coking, the weight of temperature will decrease, while the weight of pressure will increase. The similarity score represents the baseline time series vector and the optimal candidate time series vector; The eigenvalues representing the overlapping regions in the baseline time series vector. The feature values representing the overlapping region of the optimal candidate time series vectors. The feature values obtained by weighted fusion of overlapping regions;
[0119] Finally, the non-overlapping regions are directly spliced together, and the spliced non-overlapping regions and the weighted and fused overlapping regions form a complete spliced temporal vector. Example 1
[0120] (1) Data acquisition and preprocessing:
[0121] Temperature data of No. 3 and No. 2 coke ovens in a coking plant were collected (sampling frequency 1 time / second), including three production stages: coal charging stage (0-2h), coking stage (2-26h), and coking stage (26-28h). The data contains time offset (about 30 seconds) caused by sensor switching and local noise.
[0122] Data standardization processing: Z-score standardization was performed on the furnace temperature data to eliminate the influence of dimensions; then... Outliers (such as instantaneous high temperatures caused by sensor malfunctions) were removed according to the criteria. Finally, for the 20 seconds of missing data during the coal loading stage, cubic spline interpolation was used to fill in the missing values, and the data was filled based on the temperature rise trend before and after the process.
[0123] (2) Coarse matching: A window length of n=100 is used to construct a multi-dimensional feature vector composed of volatility features, level value features, and coking process features. Then, similarity matching is performed based on weighted Euclidean distance. , , The calculated weighted Euclidean distances are sorted from smallest to largest, and the top 30% of multiple candidate time series vectors are selected as high similarity candidate time series vectors.
[0124] (3) Fine alignment:
[0125] Process constraints include:
[0126] Time offset constraint: Set the maximum time difference to 280 seconds (1% of the 28-hour coking cycle);
[0127] Trend consistency constraint: During the coking stage (2-26h), the slope of the forced alignment path must be ≥0 (temperature shows an upward trend);
[0128] Feature point constraint: Force matching of temperature values at the start time (2h) and end time (26h) of coking;
[0129] Path calculation: The optimal alignment path is obtained through the DTW algorithm with process constraints, and the 30-second time offset is corrected.
[0130] (4) Overlapping region fusion and splicing: Overlapping regions (50 seconds) are fused using a dynamic weighted fusion strategy, with similarity scores... eigenvalues of weighted fusion Non-overlapping regions are directly spliced to generate a complete 28-hour furnace temperature time sequence.
[0131] Performance verification: The temperature trend of the most complete spliced time-series vector is continuous, and the error of feature points (such as the coking endpoint temperature at 26h) is ≤2℃. Compared with the single coarse matching (error 8℃) and the single fine alignment method (error 5℃), the accuracy is significantly improved.
[0132] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for stitching together multi-source coking time-series data based on similarity matching and dynamic fusion, characterized in that: Specifically, it includes the following steps: (1) Data acquisition and preprocessing: Acquire the baseline time series vector and multiple candidate time series vectors and perform data preprocessing to obtain the preprocessed baseline time series vector and multiple candidate time series vectors; the baseline time series vector is the target coking time series vector to be analyzed, and the multiple candidate time series vectors are multiple historical coking time series vectors; (2) Coarse matching: Calculate the volatility characteristics, level characteristics and coking process characteristics of the baseline time series vector and each candidate time series vector, construct the corresponding multi-dimensional feature vector, and then perform similarity matching between each candidate time series vector and the multi-dimensional feature vector corresponding to the baseline time series vector based on the weighted Euclidean distance, so as to select multiple highly similar candidate time series vectors from multiple candidate time series vectors. (3) Fine alignment: Under process constraints, the optimal alignment path between each highly similar candidate time series vector and the reference time series vector is calculated and selected by the DTW algorithm. Then, the optimal alignment path with the shortest path length is selected from multiple optimal alignment paths. The highly similar candidate time series vector corresponding to the optimal alignment path with the shortest path length is the optimal candidate time series vector, thus realizing fine time alignment between the candidate time series vector and the reference time series vector. (4) Overlapping region fusion and splicing: For the best candidate time series vector and the reference time series vector after fine alignment, the overlapping regions of the two are fused using a dynamic weighted fusion strategy, and the non-overlapping regions are directly spliced to generate a complete spliced time series vector.
2. The method for stitching multi-source coking time-series data according to claim 1, characterized in that: The aforementioned baseline time series vector and each candidate time series vector include time series data of multiple coking features, including coking features of the coal blending section, coke oven section, gas section, and methanol section. The coking characteristics of the coal blending section include the moisture content, volatile matter, sulfur content, ash content, caking index, and number of coals charged in the blended coal. The coking characteristics of the coke oven section include the temperature of coke oven gas before preheating, the temperature of coke oven gas after preheating, the pressure at the inlet of the gas collecting pipe, the pressure at the outlet of the gas collecting pipe, the suction force on the fan side of the flue, the suction force on the coke side of the flue, the temperature on the fan side of the flue, the temperature on the coke side of the flue, and the coking time. The coking characteristics of the gas section include the resistance of the primary cooler, the average gas temperature after the primary cooler, the resistance of the electrostatic precipitator, the gas pressure after the mist precipitator, the average oil temperature of the blower, the average oil pressure of the blower, the average gas pressure after the desulfurization tower, the gas temperature after the benzene washing tower, the gas pressure after the benzene washing tower, and the lean oil temperature of the benzene washing tower. The coking characteristics of the methanol section include coal gas consumption, R104 catalyst bed temperature, compressor low-pressure cylinder inlet gas flow rate, compressor circulation section inlet flow rate, converter inlet coke gas temperature, converter coke oven gas and outlet converter gas pressure difference, converter outlet converter gas temperature, RT synthesis tower inlet temperature, and RP synthesis tower pressure difference.
3. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The data preprocessing specifically includes data standardization, outlier removal, and missing value imputation. The data standardization process described above uses Z-score standardization, and the calculation process is shown in the following formula (1): (1); In equation (1), Represents data after Z-score standardization. This represents the original data before Z-score standardization. The mean of the time series data for coking characteristics. The standard deviation of the time series data for coking characteristics; The method for removing outliers is as follows: The missing values are filled using cubic spline interpolation.
4. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The calculation process for the aforementioned volatility characteristics is shown in the following formula (2): (2); In equation (2), Representing volatility characteristics, each candidate time series vector or base time series vector contains multiple volatility characteristics corresponding to charring features in its multi-dimensional feature vector. This represents the eigenvalue at time j in the preprocessed baseline or candidate time series vector. This represents the eigenvalue at time j+1 in the preprocessed baseline or candidate time series vector. This represents the number of features used to calculate the volatility characteristics in a specific coking characteristic time series data within the preprocessed baseline or candidate time series vector.
5. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The calculation process of the aforementioned level value feature is shown in the following formula (3): (3); In equation (3), The level value features represent the multi-dimensional feature vector corresponding to each candidate time series vector or baseline time series vector, which includes multiple level value features corresponding to charring features. This represents the eigenvalue at time j in the preprocessed baseline or candidate time series vector. This represents the number of features used to calculate the level value feature in a certain coking characteristic time series data within the preprocessed baseline or candidate time series vector.
6. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The coking process characteristics include coking stage identification, furnace hole number matching degree, and coke pushing time deviation rate.
7. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The weighted Euclidean distance is used to perform similarity matching between each candidate time series vector and the multi-dimensional feature vector corresponding to the baseline time series vector, as shown in the following equation (4): (4); In equation (4), Represents the weighted Euclidean distance. This represents the volatility characteristics in the multi-dimensional feature vector corresponding to the baseline time series vector. The volatility characteristics in the multi-dimensional feature vectors corresponding to the candidate time series vectors; The level value features represent the multi-dimensional feature vector corresponding to the baseline time series vector. The level value features in the multi-dimensional feature vector corresponding to the candidate time series vector; The coking process features are represented in the multi-dimensional feature vector corresponding to the baseline time series vector. The coking process features are represented in the multi-dimensional feature vectors corresponding to the candidate time series vectors. , , This represents the weighting coefficient, the value of which is dynamically adjusted according to the importance of coking characteristics in the continuous coking production stage. The specific method for selecting multiple highly similar candidate time series vectors from multiple candidate time series vectors is: selecting the corresponding... The values are sorted from smallest to largest, and the top-ranked candidate time series vectors are selected as high-similarity candidate time series vectors.
8. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The process constraints include time offset constraints, trend consistency constraints, and feature point constraints. The time offset constraint is specifically shown in equation (5) below: (5); In equation (5), Represents any moment in the reference time series vector. Represents the candidate time series vectors with high similarity to The moment when alignment matching is performed. It is the duration of a single coking cycle. The maximum allowed time offset; The trend consistency constraint is specifically shown in the following equation (6): (6); In equation (6), The temperature change rate is represented by the baseline time-series vector. The reference time vector represents the heating stage. The reference time vector represents the cooling stage; The time variable representing the reference time series vector, The time variable representing highly similar candidate time series vectors. The alignment path function represents the alignment path between the highly similar candidate time series vector and the baseline time series vector. Represents the slope of the alignment path; The feature point constraint is to force a match between the temperature values of the start and end times of coking.
9. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The specific steps for calculating and selecting the optimal alignment path using the DTW algorithm under process constraints are as follows: a. Construct the local distance matrix D: First, calculate the local distance between all time point pairs of the highly similar candidate time series vector and the baseline time series vector, forming a... The matrix D, The data length representing the baseline time series vector, The data length representing the candidate time series vectors with high similarity; Points in matrix D The calculation process for the value is shown in the following formula (7): (7); In equation (7), The eigenvalue representing the i-th time point in the baseline time series vector The feature value of the j-th time point in the highly similar candidate time series vector Differences; Representative eigenvalues and eigenvalues The Euclidean distance; b. Initialize the cumulative distance matrix C: The cumulative distance matrix C is used to record the distance from the starting point. To each point The minimum cumulative distance is initialized as: The calculation process for the values of each point in the first row and first column of the cumulative distance matrix C is shown in the following formula (8): (8); In equation (8), This represents the penalty coefficient, with a value ranging from 0.1 to 0.
3. and Represents the alignment path. This represents the alignment path corresponding to the first time point of the baseline time series vector and the j-th time point of the highly similar candidate time series vector. The process constraint is violated. When the process constraint condition is violated... Triggering penalty items When the process constraints are met, ; This represents the alignment path corresponding to the i-th time point of the baseline time series vector and the 1-th time point of the highly similar candidate time series vector. The process constraint is violated. When the process constraint condition is violated... Triggering penalty items When the process constraints are met, ; c. Fill the cumulative distance matrix C: For all points i>1 and j>1 in the cumulative distance matrix C, calculate the minimum cumulative distance using the following formula (9): (9); In equation (9), This indicates a process constraint violation indication for the upward downward path. When the upward downward path violates process constraints... When the process constraints are met, ; This indicates a violation of process constraints for a left-to-right shift path. When the process constraints are met, ; This indicates a violation of process constraints on the top-left diagonal shift path. When the top-left diagonal shift path violates process constraints... When the process constraints are met, ; When calculating the cumulative distance matrix C, each point in the cumulative distance matrix C... The value of not only records the minimum cumulative distance from the starting point to the point, but also implicitly contains the optimal predecessor point to reach the point, that is, the point with the smallest cumulative distance in the three directions of left, top, and top left. By tracing back these predecessor points, the complete optimal alignment path can be restored.
10. The method for stitching multi-source coking time-series data according to claim 2, characterized in that: The dynamic weighted fusion strategy is used for fusion, as shown in the following formula (10): (10); In equation (10), This represents the value of the i-th feature in the baseline time series vector. It is the value of the i-th feature in the optimal candidate time series vector. This represents the maximum value of the i-th feature in the historical data. This represents the minimum value of the i-th feature in the historical data. Sub-similarity representing the i-th feature; The dynamic weights represent the i-th feature; The similarity score represents the baseline time series vector and the optimal candidate time series vector; The eigenvalues representing the overlapping regions in the baseline time series vector. The feature values representing the overlapping region of the optimal candidate time series vectors. The feature value obtained by weighted fusion of overlapping regions represents the feature value.
Citation Information
Patent Citations
Coking production optimization method and device, electronic equipment and storage medium
CN114266412A
Coke quality parallel prediction method and device for simulating coking mechanism
CN116525013A
Intelligent smelting method and system for industrial silicon
CN116702028A