A production index prediction method and storage medium based on data dimensionality reduction clustering
By calculating the time-domain mutual information between input sequences and production indicators in the process industry, the model input dimension is reduced, and a production indicator prediction model that is suitable for different working conditions is established, which solves the problems of high model training complexity and insufficient prediction accuracy in the process industry, and achieves efficient production indicator prediction.
Patent Information
- Application Number
- CN202210894509.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-07-28
AI Technical Summary
The training complexity of production index prediction models in the existing process industry is high, and the prediction accuracy of traditional single soft measurement models is insufficient under different operating conditions.
By calculating the time domain mutual information between the input sequence and the production index sequence, the input sequence with the greatest correlation is selected, and dimensionality reduction and clustering are performed to establish a production index prediction model for different working conditions.
It effectively reduces the time and space complexity of model training, improves the accuracy of production indicator prediction and the ability to adapt to different working conditions.
Smart Images

Figure CN115330034B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of process industry data processing, and more specifically, relates to a production index prediction method and a storage medium based on data dimensionality reduction clustering. Background Art
[0002] The process industry refers to the industrial production process in which raw materials undergo a series of physical or chemical changes to obtain products. Common process industries include the cement industry, the petroleum industry, feed and pesticide production, etc. The influencing factors of the product quality in the process industry are relatively complex, including raw material components, the physical and chemical states of each sub-industrial system, etc. Accurately obtaining key production indicators such as production quality can specifically guide the adjustment of the working state of the industrial system, so as to achieve the purposes of saving raw materials, stabilizing product quality, reducing carbon emissions, etc. However, in the existing industrial scenarios, the key production indicators are often obtained through methods such as chemical analysis in the factory laboratory. Such methods are often unable to be carried out in real time, resulting in the lag of quality monitoring and working condition adjustment.
[0003] With the development of intelligent industry, soft measurement methods for production indicators have emerged in actual industrial scenarios. Specifically, this method analyzes data such as sensors and feedstock ratios in the production process and sends them into a prediction model to predict production indicators. However, in actual production, the amount of data obtained is large and contains a lot of redundant information. To address this problem, existing methods propose calculating the mutual information between the input information of the computational model candidate and the output information to be predicted, so as to measure the correlation between each candidate input information and the output information, and thereby screen out the candidate input information with a greater correlation with the output information for actual prediction. For example, the patent document with the application publication number "CN110570030A" discloses relevant data dimensionality reduction methods. This method can reduce the dimension of the model input to a certain extent. However, for a single input, long-time series data still needs to be collected. In the actual production index prediction process, the production index at a specific moment is only related to the sensor data within a limited time period. Therefore, after the existing method is used to perform dimensionality reduction processing on the model input data, there is still a certain amount of data redundancy, which makes the training complexity of the prediction model relatively high.
[0004] In addition, industrial production is affected by external conditions such as temperature and air pressure. At the same time, there are working states such as production increase and production decrease, and various working conditions will occur. Therefore, the traditional single soft measurement model is no longer applicable. To address this issue, in the patent document with the application publication number "CN113012766A", an adaptive soft measurement modeling method based on online selective integration is disclosed. First, it combines the advantages of K-means and KNN to construct diverse local regions, and at the same time establishes corresponding local models. Subsequently, probability analysis is used to eliminate redundant regions and corresponding local models. In addition, in the online prediction stage, the most recently obtained historical samples are used as the validation set to select the best candidate local model, and the model integration weights are determined, and then the adaptive fusion of local prediction results is realized. The generalization of this method has been improved, and it has a good effect on predicting production indicators under different working conditions. However, its model process still involves complex weight determination and model fusion processes, and the training difficulty and complexity of the prediction model are relatively high. In addition, in actual industrial production, different working conditions are often related to specific time series ranges, and this patent document does not fully consider this characteristic in the modeling process, and the prediction accuracy still needs to be further improved. Summary of the Invention
[0005] Aiming at the defects and improvement requirements of the prior art, the present invention provides a production index prediction method and a storage medium based on data dimensionality reduction clustering, aiming to reduce the training complexity of the production index prediction model.
[0006] To achieve the above object, according to one aspect of the present invention, a production index prediction method based on data dimensionality reduction clustering is provided, including: a prediction model establishment step; the prediction model establishment step includes:
[0007] Calculate the time-domain mutual information between each input sequence in the input sequence set and the production index sequence to be predicted under different time offsets; each input sequence corresponds to the detection result of a type of production data within a specified time period; the time-domain mutual information only contains the information related to the time series;
[0008] Integrate the time-domain mutual information corresponding to each input sequence respectively to obtain the total time-domain mutual information amount between each input sequence and the production index sequence, and screen out the m input sequences with the highest total time-domain mutual information amount as the target input sequences; m is a positive integer;
[0009] Obtain the dimensionality reduction features of each target input sequence at different times respectively, obtain the dimensionality reduction feature sequences at each time, and cluster the dimensionality reduction feature sequences at different times; for any input sequence, its dimensionality reduction feature is the weighted summation result of the input sequence in the time domain. During the weighted summation process, the greater the time-domain mutual information, the greater the weight value under the corresponding time offset, and the sum of all weight values is 1;
[0010] For each category obtained by clustering, use the dimensionality-reduced feature sequence and the known production index sequence therein to train a machine learning model, and obtain a production index prediction model corresponding to the time series of this category, which is used to predict the production index sequence according to the dimensionality-reduced feature sequence.
[0011] Further, for any input sequence x i (t), the time-domain mutual information I t (x i (t - Δt), y p (t)) is:
[0012]
[0013] where y p (t) represents the production index sequence; I(x i (t - Δt), y p (t)) represents the mutual information amount between the input sequence x i (t) and the production index sequence y p (t) at the time offset Δt, and its calculation formula is:
[0014] p(x) and p(y) respectively represent the marginal distributions of the input sequence x i (t) and the production index sequence y p (t), and p(x, y) represents the joint distribution of the input sequence x i (t) and the production index sequence y p (t) at the time offset Δt.
[0015] In some alternative embodiments, the total time-domain mutual information I i (t) between any input sequence x sumt (t) and the production index sequence is:
[0016] In some alternative embodiments, the total time-domain mutual information I i (t) between any input sequence x sumt (t) and the production index sequence is:
[0017] In some alternative embodiments, the dimensionality-reduced feature s i (t) of any input sequence x i (t) at time t is:
[0018] In some alternative embodiments, the dimensionality-reduced feature s i (t) of any input sequence xi (t) is:
[0019] Furthermore, the production index prediction method based on data dimensionality reduction and clustering provided by the present invention further includes: a prediction step;
[0020] The prediction step includes:
[0021] At the target moment to be predicted, obtain the dimensionality-reduced features of each target input sequence to obtain the dimensionality-reduced feature sequence at the target moment;
[0022] According to the time series range to which the target moment belongs, select the corresponding production index prediction model, input the dimensionality-reduced feature sequence at the target moment into the production index prediction model, and the production index prediction model outputs the prediction result of the production index sequence.
[0023] According to another aspect of the present invention, there is provided a computer-readable storage medium including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the production index prediction method based on data dimensionality reduction and clustering provided by the present invention.
[0024] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0025] (1) The present invention estimates the mutual information between each input and output of the prediction model under different time offsets to measure the correlation between different production data and the production index to be predicted, and screens out the part of the input sequence with the largest correlation with the production index. Thus, on the basis of ensuring the prediction effect of the model, the dimensionality of the model input can be effectively reduced; on this basis, the present invention further designs weights based on the estimated mutual information for weighted summation of the input sequence in the time domain, converting the selected input sequence into a one-dimensional value containing sequence information, further reducing the dimensionality of the model input. Generally speaking, the present invention performs a secondary dimensionality reduction process on the input of the model, effectively reducing the training difficulty of the production index prediction model and reducing the time complexity and space complexity of model training while ensuring the prediction effect of the model.
[0026] (2) On the basis of performing a secondary dimensionality reduction on the model input data, the present invention clusters the input data, and trains corresponding production index prediction models for different categories obtained by clustering respectively for predicting the production index under the corresponding time series range. Thus, appropriate models can be used for targeted prediction for different working conditions, further improving the prediction effect of the model. Description of the Drawings
[0027] Figure 1The variation relationship of the time-domain mutual information with respect to the time offset provided by the embodiments of the present invention;
[0028] Figure 2 In the production index prediction method based on data dimensionality reduction clustering provided by the embodiments of the present invention, it is a schematic diagram of the prediction model establishment step. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0030] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0031] In order to solve the problem in the existing process industry that due to data redundancy, the time and space complexity of the model training for predicting production indexes are high, the present invention provides a production index prediction method based on data dimensionality reduction clustering. The overall idea is as follows: using the mutual information between the model input and output to measure the correlation between production data and the production index to be predicted, and making full use of this information to perform sufficient dimensionality reduction processing on the model input data on the basis of ensuring the model prediction effect, thereby effectively reducing the training complexity of the model.
[0032] Before explaining the technical solution of the present invention in detail, the following brief introduction is given to the information related to mutual information.
[0033] For any two sequences X and Y, their marginal distributions p(x) and p(y), and their joint distribution p(x,y) are statistically analyzed. Then the mutual information amount I(x,y) of these two sequences can be calculated by the following formula:
[0034]
[0035] In the process industry, production indexes include salt content, etc., and production data includes sensor data such as temperature and oxygen concentration, as well as feed ratio data such as the calcium oxide content in the feed. Considering that there may be a certain time offset between the industrial process data input to the model and the production index sequence to be predicted during the production index prediction process, the present invention proposes a mutual information calculation method under the time offset Δt on the basis of the above mutual information amount. Specifically, for any input sequence x i (t), and the production index sequence y [[ID=?]] p It seems there is an incomplete tag in the original text for line 30. Please check and correct it if possible.(t) The mutual information I(x i (t - Δt), y p (t)) is as follows:
[0036]
[0037] At this time, p(x) and p(y) respectively represent the marginal distributions of the input sequence x i (t) and the production index sequence y p (t), and p(x, y) represents the joint distribution of the input sequence x i (t) and the production index sequence y p (t) at the time offset Δt.
[0038] The above mutual information I(x i (t - Δt), y p (t)) simultaneously contains the time-domain related mutual information between the input sequence x i (t) and the production index sequence y p (t), as well as other information that is time-domain independent; in order to be able to perform a more sufficient dimensionality reduction process on the production data input to the model, the present invention further analyzes the variation relationship of the mutual information between the input sequence and the production index sequence with the time offset, as Figure 1 shown, the results show that when the time offset between the input sequence and the production index sequence is large enough, the time-related mutual information in the mutual information I(x i (t - Δt), y p (t)) will tend to zero. Based on this characteristic, the present invention extracts the time-domain related mutual information from the mutual information I(x i (t - Δt), y p (t)) as the time-domain mutual information I t (x i (t - Δt), y p (t)), and the specific calculation formula is as follows:
[0039]
[0040] Since the time-domain mutual information calculated by the present invention only contains the time-sequence related information volume, therefore, based on this time-domain mutual information, the correlation between the input sequence and the output sequence of the production index prediction model in the time domain can be fully mined, and based on this, the input sequence is dimensionally reduced in the time domain to ensure that the prediction effect of the model is not affected while dimensionality reduction is performed.
[0041] The following are examples.
[0042] Example 1:
[0043] A production index prediction method based on data dimensionality reduction clustering. Specifically, in this embodiment, it is a product quality prediction method for the cement industrial scenario. Among them, the production index sequence to be predicted by soft measurement is defined as y p (t), and the input sequence set of the soft measurement prediction model is defined as X = {x1, x2,..., x n}, where n represents the total number of categories of production data. In this embodiment, n = 3. The selected production data specifically includes temperature x1, oxygen concentration x2, and the calcium oxide content x3 in the feedstock.
[0044] This embodiment includes: a step of establishing a prediction model; as Figure 2 shown, the step of establishing a prediction model specifically includes:
[0045] Calculate the time-domain mutual information between each input sequence in the input sequence set and the production index sequence to be predicted at different time offsets; each input sequence corresponds to the detection results of a certain type of production data within a specified time period; the time-domain mutual information only contains the information related to the time series and can be specifically calculated according to the above formulas (2) - (3).
[0046] After calculating the time-domain mutual information, integrate the time-domain mutual information corresponding to each input sequence respectively to obtain the total time-domain mutual information amount between each input sequence and the production index sequence, and select the m input sequences with the highest total time-domain mutual information amount as the target input sequences; m is a positive integer;
[0047] Based on the calculation expression of the time-domain mutual information, for any input sequence x i (t), the total time-domain mutual information amount I sumt (t) with the production index sequence is:
[0048]
[0049] To reduce the computational complexity, in this embodiment, based on the above formula (4), a fixed time difference T is set, and the total time-domain mutual information amount I i (t) between the input sequence x sumt (t) and the production index sequence is calculated according to the following formula:
[0050]
[0051] In the interval (-T, T), the quantities of the production data x1, x2, x3 sampled at the minute level and the production index sequence y p are greatly restricted, thereby effectively reducing the computational complexity of the dimensionality reduction features. The time difference T can be set in combination with statistical information and prior information.
[0052] Optionally, in this embodiment, after calculating the total time-domain mutual information between the input sequence and the production index sequence, only the input sequence with the largest total time-domain mutual information is selected as the target input sequence, that is, m = 1, and the selected target input sequence is denoted as x k ; in the subsequent model training process, only the selected target input sequence is used. This embodiment calculates the total time-domain mutual information between the input target sequence and the production index sequence using time-domain mutual information, which can accurately measure the correlation between the input sequence and the production index sequence in the time domain. By screening the input sequence with the largest total time-domain mutual information, the number of model inputs can be effectively reduced while ensuring the model prediction effect.
[0053] The dimensionality-reduced features of each target input sequence are obtained at different times to obtain the dimensionality-reduced feature sequences at each time, and the dimensionality-reduced feature sequences at different times are clustered; for any input sequence, its dimensionality-reduced feature is the weighted sum result of the input sequence in the time domain. During the weighted sum process, the larger the time-domain mutual information, the larger the weight value corresponding to the time offset, and the sum of all weight values is 1. The designed weight value ensures that the dimensionality-reduced feature is more strongly influenced by the data strongly correlated with the predicted production index sequence; optionally, for any input sequence x i (t), the dimensionality-reduced feature s i (t) at time t is:
[0054]
[0055] Based on the above formula (6), in order to further reduce the computational complexity, in this embodiment, a fixed time difference T is set, and the above dimensionality-reduced feature is approximately represented within the range (-T, T) as:
[0056]
[0057] This embodiment designs the weight value based on the estimated mutual information and performs weighted summation on the input sequence in the time domain, converting the selected input sequence into a one-dimensional value containing the entire sequence information, further reducing the input dimension of the model; that is to say, on the basis of reducing the number of model input sequences according to the total time-domain mutual information, this embodiment further compresses the selected input sequences in the time domain, realizing the secondary compression of the model input, and effectively reducing the input dimension of the model.
[0058] As an alternative implementation, in this embodiment, when clustering the dimensionality-reduced feature sequences at different times, the K-means clustering algorithm is specifically used; since in process industries, different operating conditions are often related to specific time series ranges. For example, in winter and summer, there are different operating conditions. Therefore, in this embodiment, the time series ranges corresponding to the multiple categories obtained by clustering often do not overlap with each other.
[0059] After clustering, for each category obtained by clustering, use the dimensionality-reduced feature sequence and the known production index sequence therein to train a machine learning model, and obtain a production index prediction model corresponding to the time series of this category, which is used to predict the production index sequence according to the dimensionality-reduced feature sequence.
[0060] Based on the above steps for establishing the prediction model, this embodiment further includes: a prediction step;
[0061] The prediction step includes:
[0062] At the target moment to be predicted, obtain the dimensionality-reduced features of each target input sequence to obtain the dimensionality-reduced feature sequence at the target moment;
[0063] According to the time series range to which the target moment belongs, select the corresponding production index prediction model, input the dimensionality-reduced feature sequence at the target moment into this production index prediction model, and the production index prediction model outputs the prediction result of the production index sequence.
[0064] Generally speaking, starting from the perspective of data mutual information, this embodiment can accurately capture the correlation between the input sequence and the production index sequence whether the correlation between the input sequence and the production index sequence is linear or non-linear; cluster the data under different working conditions, extract the working condition features of the data before the prediction model, and reduce the training and testing complexity of the prediction model, thereby significantly improving the accuracy of the soft measurement process.
[0065] Embodiment 2:
[0066] A computer-readable storage medium includes a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the production index prediction method based on data dimensionality reduction and clustering provided in Embodiment 1 above.
[0067] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A production index prediction method based on data dimensionality reduction clustering, characterized in that Including: Steps for establishing a prediction model; the steps for establishing the prediction model include: Calculating the time-domain mutual information between each input sequence in the input sequence set and the production index sequence to be predicted under different time offsets; each input sequence corresponds to the detection result of a class of production data within a specified time period; the time-domain mutual information only contains the information quantity related to time series; Integrating the time-domain mutual information corresponding to each input sequence respectively to obtain the total time-domain mutual information quantity between each input sequence and the production index sequence, and screening out the m input sequences with the highest total time-domain mutual information quantity as the target input sequences; m is a positive integer; Obtaining the dimensionality-reduced features of each target input sequence at different moments respectively to obtain the dimensionality-reduced feature sequences at each moment, and clustering the dimensionality-reduced feature sequences at different moments; for any input sequence, its dimensionality-reduced feature is the weighted summation result of the input sequence in the time domain. During the weighted summation process, the greater the time-domain mutual information, the greater the weight value at the corresponding time offset, and the sum of all weight values is 1; For each category obtained by clustering, using the dimensionality-reduced feature sequence therein and the known production index sequence to train a machine learning model to obtain the production index prediction model corresponding to the time series of this category, which is used to predict the production index sequence according to the dimensionality-reduced feature sequence; Any input sequence x i (t) and the time-domain mutual information I of the production index sequence at any time offset Δt t (x i (t - Δt), y p (t)) is as follows: Among them, y p (t) represents the production index sequence; I(x i (t - Δt), y p (t)) represents the mutual information between the input sequence x i (t) and the production index sequence y p (t) at the time offset Δt, and its calculation formula is: p(x) and p(y) respectively represent the marginal distributions of the input sequence x i (t) and the production index sequence y p (t), and p(x,y) represents the joint distribution of the input sequence x i (t) and the production index sequence y p (t) at a time shift of Δt; Any input sequence x i The dimensionality-reduced feature s i (t) at time t is: Alternatively, any input sequence x i (t)'s dimensionality-reduced feature s i (t) at time t is:
2. The production index prediction method based on data dimensionality reduction clustering according to claim 1, wherein Any input sequence x i (t) and the total time-domain mutual information I of the production index sequence sumt (t) is as follows:
3. The production index prediction method based on data dimensionality reduction clustering according to claim 1, wherein Any input sequence x i (t) and the total time-domain mutual information I of the production index sequence sumt (t) is as follows:
4. The production index prediction method based on data dimensionality reduction clustering according to any one of claims 1 to 3, characterized in that, Also including: Prediction steps; The prediction steps include: At the target moment to be predicted, obtaining the dimensionality-reduced features of each target input sequence to obtain the dimensionality-reduced feature sequence at the target moment; According to the time series range to which the target moment belongs, selecting the corresponding production index prediction model, inputting the dimensionality-reduced feature sequence at the target moment into the production index prediction model, and outputting the prediction result of the production index sequence by the production index prediction model.
5. A computer-readable storage medium, characterized in that, Including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the production index prediction method based on data dimensionality reduction and clustering according to any one of claims 1 to 4.
Citation Information
Patent Citations
Wind power cluster power interval prediction method and system based on deep learning
CN110570030A
Self-adaptive soft measurement modeling method based on online selective integration
CN113012766A
Methods and systems for inferred information propagation for aircraft prognostics
US20160292302A1
Data prediction system and data prediction method
US20190303783A1