Method for generating teacher data, method for generating coke quality prediction model, coke quality prediction method, system, and program

The machine learning-based teacher data generation method addresses the challenge of predicting coke quality by converting raw coal blending ratios into cluster-based ratios, significantly improving prediction accuracy by accounting for compatibility effects between raw coals.

JP2025085797APending Publication Date: 2025-06-05KANSAI COKE & CHEMICALS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025047153
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-09-01
Filing Date
2025-03-21
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing methods for predicting coke quality, such as those described in Patent Document 1, face challenges in accurately estimating coke quality due to compatibility issues between raw coals, leading to insufficient prediction accuracy.

Method used

A machine learning-based teacher data generation method that converts blending ratios of raw coals into cluster-based ratios, using dimension reduction and cluster analysis to improve prediction accuracy by handling compatibility effects.

Benefits of technology

The proposed method enhances the accuracy of coke quality prediction by addressing compatibility issues between raw coals, resulting in improved correlation coefficients and prediction performance compared to existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085797000001_ABST
    Figure 2025085797000001_ABST
Patent Text Reader

Abstract

To provide a technique that enables prediction of coke quality values based on blending ratios of multiple types of raw material coal, and improves prediction accuracy.SOLUTION: A method for generating teacher data used for machine learning of a prediction model for predicting coke quality, includes the steps of: acquiring measured data associating blending ratios of multiple types of raw material coal with actually obtained coke quality values; extracting, from the acquired measured data, multiple coke quality values for which combinations of blending ratios of the multiple types of raw material coal are consecutively identical; calculating, based on the extracted multiple quality values, a characteristic quality value serving as a representative value; converting into aggregated new measured data with the calculated representative value as a new quality value; and generating teacher data using blending ratios of the raw material coal in the new measured data as input and the coke quality value as output.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to machine learning techniques for predicting coke grade. [Background technology]

[0002] Coke is obtained by carbonizing a coal blend obtained by blending multiple brands of raw coal. In order to obtain coke that satisfies a desired quality (e.g., physical property values ​​such as coke strength DI), various methods for predicting the quality of coke have been proposed.

[0003] For example, Patent Document 1 describes that coke strength is calculated using machine learning based on the characteristics of blended coal and operating conditions. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 9-165579 Summary of the Invention [Problem to be solved by the invention]

[0005] In the method described in Patent Document 1, the properties of each of a plurality of types of raw coal are measured, and the product sum of the properties of the raw coal is calculated based on the blending ratio of the raw coal, and the product sum is regarded as the property of the coal blend. However, in blending, there is compatibility between the raw coals, and it is believed that the compatibility affects the properties of the coal blend. It is believed that there are cases in which the quality of coke cannot be accurately estimated by the product sum of various property values ​​(parameters) of the raw coal.

[0006] Therefore, the inventors of the present invention thought that by having a machine learning model learn the blending ratios of each of multiple types of raw coal, it would be possible to predict the quality of coke including the compatibility of the raw coal even if the blending ratio of the raw coal is unknown. In fact, the model learned the correspondence between the blending ratios of each of multiple types of raw coal and the quality value of the coke (coke strength). However, when comparing the prediction results with the actual measured values, it is difficult to say that sufficient accuracy has been obtained.

[0007] The present disclosure provides a technology that makes it possible to predict the quality value of coke based on the blending ratio of each of multiple types of raw coal, and improves the accuracy of the prediction. [Means for solving the problem]

[0008] The teacher data generation method disclosed herein is a method for generating teacher data to be used in machine learning of a predictive model for predicting the quality of coke, and includes the steps of acquiring actual measurement data correlating the blending ratios of each of the multiple types of raw coal with the quality value of the actually obtained coke, extracting from the acquired actual measurement data multiple quality values ​​of the coke that have the same consecutive combinations of blending ratios of each of the multiple types of raw coal, calculating a representative quality characteristic value based on the extracted multiple quality values, converting the calculated representative value into new aggregated actual measurement data as a new quality value, and generating teacher data that uses the blending ratios of each of the raw coals of the new actual measurement data as input and outputs the quality value of the coke. [Brief description of the drawings]

[0009] [Figure 1] FIG. 2 is a block diagram showing a teacher data generation system, a prediction model generation system, and a coke quality prediction system according to the present embodiment. [Diagram 2] 4 is a flowchart showing processing executed by the teacher data generation system of the present embodiment. [Diagram 3] 4 is a flowchart showing a process executed by the predictive model generation system of the present embodiment. [Figure 4]3 is a flowchart showing a process executed by the coke quality prediction system of the present embodiment. [Diagram 5] FIG. 13 is a graph showing prediction results and actual values ​​in Example 3 and Comparative Example 1. [Figure 6] FIG. 13 is a graph showing prediction results and actual values ​​in Example 2 and Comparative Example 1. [Figure 7] FIG. 13 is a diagram showing prediction results and actual values ​​in Examples 1 and 3. [Figure 8] FIG. 11 is a block diagram showing a system according to a third embodiment. [Figure 9] FIG. 13 is a diagram showing prediction results of pattern 1 (single learning), prediction results of pattern 2 (immediately preceding learning), and performance values. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] FIG. 1 is a block diagram showing a teacher data generation system 1, a prediction model generation system 2, and a coke quality prediction system 3 of this embodiment. As shown in FIG. 1, the coke quality prediction system 3 predicts the quality value of a coke obtained from a coal blend blended at the blending ratio of each raw coal (brand) represented by prediction target blending data D6. The coke quality prediction system 3 uses a prediction model 20 generated in advance by machine learning by the prediction model generation system 2. In this embodiment, coke strength DI is used as the quality value. The prediction target blending data D6 represents the blending ratio of each of multiple types of raw coal. The prediction model generation system 2 builds the prediction model 20 by learning it through machine learning using the teacher data D2. The teacher data generation system 1 generates teacher data D2 for machine learning the prediction model 20 based on the actual measurement data D1. The actual measurement data D1 represents the quality value of a coke produced from a certain coal blend and the blending ratio of each of multiple types of raw coal constituting the coal blend. The numerical value indicated by the actual measurement data D1 is based on actual measurement. Each system will be described below individually.

[0011] Each of the components constituting each of the systems 1 to 3 is realized in a computer including a CPU, RAM, ROM, non-volatile memory, an input / output interface, etc., by the CPU executing information processing in accordance with a program loaded from the ROM or non-volatile memory to the RAM.

[0012] <Teacher data generation system 1> As shown in FIG. 1, the teacher data generation system 1 includes an actual measurement data acquisition unit 10 and a teacher data generation unit 11.

[0013] The actual measurement data acquisition unit 10 acquires the actual measurement data D1 by input from a user or by reading data from an external storage or an internal storage. The actual measurement data D1 may be in any format as long as it is data that associates the blending ratio of each of a plurality of types of raw coal with the quality value of the coke actually obtained by the blending. The blending ratio may be the ratio itself or an amount that indirectly indicates the ratio. In this embodiment, the actual measurement data D1 is based on the operation results of the coke oven, and associates the blending ratio of each raw coal with the quality value of the coke obtained by the blending. Although the number of types (brands) of raw coal was L, L can be changed appropriately as long as it is multiple, depending on the convenience of implementation. In this embodiment, data for L=118 was used. In the example shown in FIG. 1, the actual measurement data D1 indicates that the quality value of the coke on April 1 is α1, the blending ratio of raw coal 1 is 3%, the blending ratio of raw coal 2 is 0%, and the blending ratio of raw coal 3 is 2%. Note that these numerical values ​​are examples for explanation.

[0014] The teacher data generating unit 11 generates teacher data D2 based on the actual measurement data D1 acquired by the actual measurement data acquiring unit 10. As a comparative example, when the actual measurement data D1 is directly used as the teacher data, the teacher data indicates that the blending ratios (L numerical values) of each of the multiple types of raw coal are input to the prediction model 20, and one output value that is the correct answer from the prediction model 20 is the grade value. That is, the number of input parameters is L (118), and the number of combinations of input values ​​is enormous. Perhaps for this reason, even if the prediction model generating system 2 trains the prediction model 20 using the teacher data, there is a limit to the improvement in accuracy. Details will be described below as Comparative Example 1.

[0015] Therefore, in the present disclosure, the teacher data generation system 1 is configured as follows.

[0016] As shown in FIG. 1, the teacher data generation system 1 further includes a raw coal property data acquisition unit 12, a dimension reduction processing unit 13, a cluster analysis unit 14, and a relationship identification unit 15.

[0017] The raw coal physical property data acquisition unit 12 shown in FIG. 1 acquires raw coal physical property data D3. In the raw coal physical property data D3, multiple (i types) physical property values ​​of each raw coal are associated with multiple types of raw coal (brands). There are various types of physical property values ​​of raw coal, but in this embodiment, the following 17 types of physical property values ​​are used for i=17 shown in FIG. 1. Of course, this is an example, and any physical property value can be used. In this embodiment, each of the multiple types (118 types) of raw coal has 17 types of physical property values, so the raw coal physical property data D3 can be expressed as a matrix of L×i, that is, 118×17. ASH: Ash content in coal FT: coal softening temperature MFT: Maximum fluidity temperature of coal logMF: Maximum fluidity of coal DT: coal resolidification temperature RF: Softening and melting temperature range (DT-FT) Ro: Average maximum reflectance of the vitrinite structure of the coal tissue TI: Amount of inactive components in coal (does not soften or melt) C(dry-base): The proportion of carbon in coal H(dry base): The proportion of hydrogen in coal N (dry base): The proportion of nitrogen in coal O(dry-base): Oxygen content in coal ORI: Reference) Evaluation of Coal Coking Properties by Pyrolysis Products (Journal of the Japan Institute of Energy, Vol. 73, No. 4, 1994, p. 279-289) HRI: Reference) Evaluation of coal coking properties using pyrolysis products (Journal of the Japan Institute of Energy, Vol. 73, No. 4, 1994, p. 279-289) dHI: Reference) Evaluation of coal coking properties using pyrolysis products (Journal of the Japan Institute of Energy, Vol. 73, No. 4, 1994, p. 279-289) VM (dry-ash-free): The amount of volatile matter in coal

[0018] The dimension reduction processing unit 13 shown in FIG. 1 performs a dimension reduction process on a plurality of physical property values ​​of raw coal, and calculates dimension-reduced feature data D4 representing the plurality of physical property values. In this embodiment, one type of raw coal is expressed by 17 types of physical property values. That is, the number of dimensions indicating the characteristics of raw coal is 17. The dimension reduction processing unit 13 reduces the number of dimensions indicating the characteristics of raw coal to a number smaller than 17 by the dimension reduction process. By the dimension reduction, the number of dimensions can be reduced while reproducing the original characteristics, and it is possible to reduce the calculation cost and avoid the curse of dimensionality (overlearning). As an algorithm used for the dimension reduction process, various algorithms such as principal component analysis, factor analysis, multifactor analysis, multiple correspondence analysis, t-SNE, and Autoencoder can be used. In this embodiment, principal component analysis (PCA) is used, and the first principal component, the second principal component, and the third principal component are adopted as the feature data. Of course, this is not limited to this, and various changes are possible. For example, in the principal component analysis, the feature data may be only the first principal component, or the first and second principal components. That is, a predetermined number (any number equal to or greater than 1) of principal components from the first order onwards are selected as feature data. The predetermined number can be set appropriately depending on the required accuracy. The dimension reduction processing unit 13 executes the above-mentioned dimension reduction processing for all raw coals for which there is physical property data in the raw coal physical property data D3.

[0019] In this embodiment, when a principal component analysis was performed on the above 17 types of physical property values, it was found that the first principal component alone could reproduce approximately 64.9% of the features, the first and second principal components alone could reproduce approximately 80.8% of the features, the first to third principal components alone could reproduce approximately 86.8% of the features, and the first to fourth principal components alone could reproduce approximately 92.1% of the features. These are just examples, and since they change depending on the combination of physical property values ​​used, they are merely examples for understanding the dimensionality reduction in this disclosure and do not limit the scope of rights.

[0020] The cluster analysis unit 14 shown in FIG. 1 classifies each of the multiple types (L=118 types in this embodiment) of raw coal into one of multiple clusters by performing cluster analysis on the feature data D4 calculated by the dimension reduction processing unit 13. In this embodiment, the number of initial clusters is set to 35, and the cluster analysis is performed using the x-means algorithm to obtain a calculation result of the number of clusters k=41, but this is not limited as long as it is a cluster analysis. For example, non-hierarchical cluster analysis such as k-means++, g-means, x-means, and xg-means may be used, or hierarchical cluster analysis, DBSCAN, agglomerative clustering, and partition clustering may be used. In this embodiment, the number of clusters k is smaller than the number L of raw coal types, but is not limited to this. The number of clusters may be the same as the number of raw coal types, or the number of clusters may be greater than the number of raw coal types. Note that raw coals that have defects in one or more of the 17 types of physical properties (five exception brands) are not grouped together with other raw coals and are placed in a single cluster. Therefore, the total number of clusters k is 41+5=46.

[0021] 1 identifies the correspondence between raw coal (brand) and clusters, based on the classification result of raw coal into clusters by cluster analysis unit 14. Data D5 indicating the correspondence is stored in memory unit 16.

[0022] The teacher data generating unit 11 shown in FIG. 1 generates teacher data D2 based on the actual measurement data D1 acquired by the actual measurement data acquiring unit 10 and the correspondence between raw coal and clusters stored in the storage unit 16. Specifically, the blending ratio of each raw coal in the actual measurement data D1 is converted into the blending ratio of each cluster. For example, as shown in FIG. 1, when cluster 2 corresponds to raw coal 1 (brand 1) and raw coal 3 (brand 3), the blending ratios of all raw coals corresponding to cluster 2 in the actual measurement data D1 are aggregated and converted into blending ratios on a cluster basis. Then, the teacher data generating unit 11 generates teacher data D2 in which the blending ratios of each cluster are converted and input to the prediction model 20, and the quality value of coke is output. In this embodiment, the number L of raw coal types is 118, but since the number of clusters is 46, the number of input parameters to the prediction model 20 can be reduced. In this embodiment, since the physical property values ​​differ depending on the lot even for the same raw coal, the physical property values ​​are measured for each lot. Therefore, approximately 1,500 lots were accurately classified (feature extracted) into clusters with a number of k = 46, which we believe has improved accuracy.

[0023] As shown in FIG. 1, the actual measurement data D1 is based on the actual measurement value of the coke produced in the coke oven, so that even if the blending ratio is the same, the quality value may include variation. For example, as shown in FIG. 1, the blending ratio is the same on April 1 and April 2, but the quality values ​​vary as α1 and α2. If the variation is small, the adverse effect on the prediction accuracy of the prediction model 20 is small, but if the variation is large, the prediction accuracy may be adversely affected. Therefore, in order to reduce this adverse effect, it is preferable to configure the actual measurement data acquisition unit 10 as follows. That is, it is preferable that the actual measurement data acquisition unit 10 has an identical blending extraction unit 10a, a representative value determination unit 10b, and a data aggregation unit 10c. The identical blending extraction unit 10a extracts multiple quality values ​​of coke having the same combination of blending ratios of multiple types of raw coal from the acquired actual measurement data. In the example of FIG. 1, since the blending ratio is the same on April 1 and April 2, multiple quality values ​​(α1, α2) related to the identical blending are extracted. The representative value determination unit 10b calculates a characteristic value of the quality to be the representative value based on the extracted multiple quality values. Examples of the characteristic value include the average value, the mode value, and the median value, but in this embodiment, the average value is used. It has been found that the average value is closer to the true value of the population than the mode value and the median value. The data aggregation unit 10c converts one calculated representative value into new actual measurement data aggregated as a new quality value. This makes it possible to reduce or eliminate the variation in data in which the same composition has different quality values. In addition, in order to further improve accuracy, after the representative value determination unit 10b calculates the representative value (average value), only the quality values ​​whose absolute value difference from the representative value (average value) is within a predetermined threshold value may be left, and the remaining quality values ​​may be recalculated, and the representative value may be converted into new actual measurement data aggregated as a new quality value. In addition, a data deletion unit may be provided that limits the actual measurement data used for machine learning to only data obtained when the same composition is continuously manufactured a predetermined number of times or more (for example, five or more measurements), and deletes data that has not been continuously manufactured for a predetermined period of time. Note that each of these units 10a to 10c may be omitted. The data deletion unit may delete data that has only one identical composition. This is because data that has only one identical composition cannot obtain a representative value (average value), and the quality value may not be a true value. Also, the acquired performance data D1 may be configured to include data on operations, and abnormal values ​​may be removed based on the operation data. Examples of data on operations include coke oven temperature data and measurement data on coal moisture. Data that are abnormal values ​​in a box-and-whisker plot for oven temperature and moisture over the entire operation period may be removed. Abnormal values ​​are in the range from the minimum value to the first quartile in the box-and-whisker plot (bottom 25%) and in the range from the third quartile to the maximum value (top 25%).

[0024] <Prediction model generation system 2> As shown in FIG. 1, the prediction model generation system 2 has a learning unit 21 that constructs a prediction model 20 by learning through machine learning using teacher data D2. As the prediction model 20, various models such as linear regression, regression tree, random forest, support vector machine, neural network, ensemble, etc. can be used as long as it is a supervised machine learning model. In this embodiment, two examples are used. The first example is a regression tree (M5P tree algorithm) that has a tree structure that performs case classification according to an input value and has a linear regression formula in the leaf. The prediction result will be described below as an example.

[0025] <Coke quality prediction system 3> The coke quality prediction system 3 shown in FIG. 1 includes a prediction target data acquisition unit 30, a conversion unit 31, and a prediction unit 32.

[0026] The prediction target data acquisition unit 30 acquires prediction target blending data D6 indicating the blending ratio of each of a plurality of raw coals constituting a blended coal, which is a raw material for coke whose quality value is to be predicted. The prediction target blending data D6 includes the blending ratio V 1 ,V 2 ,V 3 ,…,V L-1 ,V L Represents.

[0027] The conversion unit 31 converts the blending ratio of the acquired raw coal unit into the blending ratio of the cluster unit by using the correspondence relationship (data D5) that previously associates raw coal with the cluster. The conversion process by the conversion unit 31 is the same as the conversion process by the teacher data generation unit 11. The conversion unit 31 converts the blending ratio (X 1 ,X 2 ,X 3 ,…,X k-1 ,X k ) is converted into post-conversion blend data D7 representing the

[0028] The prediction unit 32 uses the prediction model 20 to calculate the mixture ratio (X 1 ,X 2 ,X 3 ,…,X k-1 ,X k ) to predict the quality value of the coke. As described above, the prediction model 20 is constructed in advance by the learning unit 21 of the prediction model generation system 2 through machine learning so as to output the quality value of the coke from the blending ratio of each cluster unit. By inputting the converted blending data D7 to the prediction model 20, the prediction model 20 outputs the quality value of the coke.

[0029] [How to generate training data] The teacher data generation method will be described with reference to Fig. 2. As shown in Fig. 2, steps ST100 to ST103 and steps ST104 to ST107 can be executed in parallel, and may be executed in parallel, or one may be executed first. The order is not limited. In step ST100, the teacher data generation system 1 acquires raw coal physical property data D3 in which multiple physical property values ​​of each raw coal are associated with each of multiple types (L types) of raw coal. In the next step ST101, the teacher data generation system 1 performs a dimension reduction process on the multiple physical property values ​​of each raw coal in the raw coal physical property data D3 for each of the multiple types (L types) of raw coal, and calculates dimension-reduced feature data D4 representing the multiple physical property values. In the next step ST102, the teacher data generation system 1 classifies the calculated feature data D4 into one of multiple clusters by performing cluster analysis. Next, in step ST103, the teacher data generation system 1 identifies the correspondence between raw coal and clusters based on the results of the cluster analysis. In the next step ST104, the teacher data generation system 1 acquires actual measurement data D1 in which the blending ratio of each of the multiple types (L types) of raw coal is associated with the grade value of the actually obtained coke.

[0030] In step ST105, the teacher data generation system 1 extracts multiple coke quality values ​​in which the combination of blending ratios of multiple types (L type) of raw coal is consecutively the same from the acquired actual measurement data D1. In the next step ST106, the teacher data generation system 1 calculates a representative quality characteristic value (average value in this embodiment) based on the multiple extracted quality values. In the next step ST107, the teacher data generation system 1 converts the calculated representative value into new aggregated actual measurement data as a new quality value. Note that steps ST105 to ST107 can be omitted.

[0031] In the next step ST108, the teacher data generation system 1 uses the correspondence between raw coal and clusters (data D5) to convert the blending ratios of each raw coal in the actual measurement data D1 into blending ratios of each cluster, and generates teacher data that takes the blending ratios of each converted cluster as input and outputs the coke quality value.

[0032] [How to generate a coke quality prediction model] A coke quality prediction model generation method will be described with reference to Fig. 3. As shown in Fig. 3, in step ST200, the above-mentioned teacher data generation process is executed to generate teacher data D2. In the next step ST201, the prediction model generation system 2 performs machine learning using the teacher data to generate a prediction model 20 that predicts the quality value of coke.

[0033] [Coke quality prediction method] The coke quality prediction method will be described with reference to Fig. 4. As shown in Fig. 4, in step ST300, the coke quality prediction system 3 calculates the blending ratio (V 1 ,V 2 ,V 3 ,…,V L-1 ,V L In the next step ST301, the coke quality prediction system 3 obtains the blending ratio (V 1 ,V 2 ,V 3 ,…,V L-1 ,V L ) to the cluster-by-cluster mix ratio (X 1 ,X 2 ,X 3 ,…,X k-1 ,X k In the next step ST302, the coke quality prediction system 3 converts the blending ratio (X 1 ,X 2 ,X 3 ,…,X k-1 ,X k ) to predict the coke quality value.

[0034] [Effects of this disclosure] In order to show the effects of the present disclosure, Comparative Example 1 and Examples 1, 2, and 3 are shown below. In the comparison between Comparative Example 1 and Example 3, the comparison between Comparative Example 1 and Example 2, and the comparison between Example 1 and Example 3, the same measured data D1 was used.

[0035] Example 1 [DI statistical processing (average) & stock clustering] The system configuration is shown in FIG. 1. In the teacher data generation process shown in FIG. 2, all of ST100 to 108 are executed. The number of brands L (number of types of raw coal) used is 118. The number of physical property values ​​j=17. The number of clusters k=46. The prediction model 20 is a regression tree (M5P tree algorithm). There are 46 input parameters to the prediction model 20. An averaging process is used to calculate the characteristic values ​​in the representative value determination unit 10b.

[0036] Example 2 [DI statistical processing not performed & stocks clustered] This configuration is the same as that of the first embodiment except that the identical blend extraction unit 10a, the representative value determination unit 10b, and the data aggregation unit 10c in the actual measurement data acquisition unit 10 are removed. In the teacher data generation process shown in FIG.

[0037] Example 3 [DI statistical processing (average) & stocks not clustered] The system configuration is shown in FIG. 8. As shown in FIG. 8, the raw coal physical property data acquisition unit 12, the dimension reduction processing unit 13, the cluster analysis unit 14, the relationship identification unit 15, the storage unit 16, and the conversion unit 31 are removed from the configuration of the first embodiment. That is, in the teacher data generation process shown in FIG. 2, ST100 to ST103 are not executed, but ST104 to ST107 are executed, and in ST108, teacher data is generated in which the blending ratio of each raw coal of the actual measurement data is input and the quality value of the coke is output (see teacher data D2 in FIG. 8). In the coke quality prediction process shown in FIG. 4, ST300 is executed, ST301 is not executed, and in ST302, the quality value of the coke is predicted from the blending ratio of the raw coal unit obtained using a prediction model constructed by machine learning so as to output the quality value of the coke from the blending ratio of the raw materials of the coke to be predicted from the blending ratio of the raw coal unit obtained. The number of brands L (the number of types of raw coal) used is 118. The prediction model 20 is a regression tree (M5P tree algorithm). There are 118 input parameters to the prediction model 20.

[0038] Comparison Example 1 [DI statistical processing not performed & stocks not clustered] This is a configuration in which the same blend extraction unit 10a, the representative value determination unit 10b, and the data aggregation unit 10c in the actual measurement data acquisition unit 10 shown in Fig. 8 are removed from Example 3. In the teacher data generation process shown in Fig. 2, ST100 to ST103 are not executed, and steps ST105, ST106, and ST107 are not executed. The rest is the same as Comparative Example 3.

[0039] FIG. 5 is a diagram showing the prediction results and actual values ​​of Example 3 and Comparative Example 1. The vertical axis indicates the quality value of the coke (coke strength), and the horizontal axis corresponds to the date of the measured data D1. The dates are the same in FIG. 5 and FIG. 7, but details are omitted. The correlation coefficient for the actual values ​​of Comparative Example 1 is 0.026. The correlation coefficient for the actual values ​​of Example 3 is 0.239. It can be seen that Example 3 has an improved correlation coefficient and improved prediction accuracy compared to Comparative Example 1. This makes it clear that the statistical processing of ST105, 106, and 107 in FIG. 2 is very effective in improving prediction accuracy.

[0040] FIG. 6 is a diagram showing the prediction results and actual values ​​of Example 2 and Comparative Example 1. The vertical axis is the same as FIG. 5. The horizontal axis is slightly longer than FIG. 5, but is almost the same. The correlation coefficient for the actual value of Example 2 is 0.379, and the correlation coefficient for the actual value of Comparative Example 1 is 0.239. It can be seen that Example 2 has a higher correlation coefficient than Comparative Example 1, and the prediction accuracy is improved. This makes it clear that the clustering process shown in ST100 to 103 in FIG. 2 is very effective in improving the prediction accuracy. Note that the slight difference in the period of the horizontal axis between FIG. 5 and FIG. 6 is due to the fact that averaging is not performed, resulting in data that is not reduced, and the actual data used for the prediction data is different. In other words, the data before averaging is performed is the same.

[0041] FIG. 7 is a diagram showing prediction results and actual values ​​in Examples 1 and 3. The vertical and horizontal axes are the same as in FIG. 5. The correlation coefficient for the actual measured values ​​in Example 1 is 0.379. The correlation coefficient for the actual measured values ​​in Example 3 is 0.239. It can be seen that Example 1 has an improved correlation coefficient and prediction accuracy compared to Example 3. This makes it clear that using the statistical processing in ST105, 106, and 107 in FIG. 2 in combination with the clustering processing shown in ST100 to 103 in FIG. 2 is effective in improving prediction accuracy.

[0042] [Prediction accuracy based on learning period and prediction time] The prediction results shown in Figures 5 to 7 use a model trained on data from time point 0, which is before time point 1, to just before time point 1 (the day before time point 1), and predict the grade values ​​from time point 1 to time point 2, or from time point 1 to time point 3. In other words, if time point 1 is April 1, 2018, the model is trained using data from time point 0 (for example, several years ago) to March 31, 2018, and the grade values ​​from time point 1 onwards are predicted without re-training.

[0043] Therefore, in order to investigate whether there is a difference in prediction accuracy depending on the relationship between the learning period and the prediction time point, learning was performed using the following two patterns 1 and 2, and the grade value was predicted from time point 4 to time point 5. For both patterns 1 and 2, the method of Example 1 [DI statistical processing (average) & brand clustering] was used, and the difference is the learning period. Time point 4 is later than time point 1, and time point 5 is later than time point 4. In pattern 1 (also referred to as one-time learning), data from time point 0 to time point 1 was learned once, and then the quality value from time point 4 to time point 5 was predicted without further learning. In pattern 2 (last-minute learning), before predicting the quality value at the prediction time point from time point 4 to time point 5, data from time point 0 to just before the prediction time point (including time points 0 and 1) was learned, and predictions were made after learning was completed. In pattern 2 (last-minute learning), the amount of data to learn increases as days pass.

[0044] FIG. 9 is a diagram showing the prediction results of pattern 1 (single learning), the prediction results of pattern 2 (last-minute learning), and the actual values. The vertical axis corresponds to the quality value of the coke (coke strength), and the horizontal axis corresponds to the date. The correlation coefficient of pattern 1 (single learning) with the actual values ​​is 0.111, and the correlation coefficient of pattern 2 (last-minute learning) with the actual values ​​is 0.379. It can be seen that pattern 2 (last-minute learning) has a higher correlation coefficient and prediction accuracy than pattern 1 (single learning). This is because the properties of raw coal may change depending on the lot, even if the brand is the same, and the blend of raw coal changes daily in coke production. Therefore, in pattern 1 (single learning), the prediction of unlearned areas increases as days pass, which is considered to have deteriorated the accuracy compared to pattern 2 (last-minute learning). Since pattern 2 (last-minute learning) learns up to the most recent time, there are fewer unlearned areas compared to pattern 1 (single learning), which is considered to have affected the high accuracy.

[0045] As described above, as in the present embodiment, the teacher data generation method is a method for generating teacher data D2 used in machine learning of a prediction model 20 that predicts the grade of coke, and may include the steps of acquiring raw coal physical property data D3 in which multiple physical property values ​​of each of multiple types of raw coal are associated with each of the multiple types of raw coal, performing a dimensionality reduction process on the multiple physical property values ​​of each of the multiple types of raw coal in the raw coal physical property data D3, and calculating dimension-reduced feature data D4 representing the multiple physical property values, classifying each of the multiple types of raw coal into one of multiple clusters by cluster analysis of the calculated feature data D4 and identifying the correspondence between the raw coal and the cluster, acquiring actual measurement data D1 in which the blending ratio of each of the multiple types of raw coal is associated with the grade value of the actually obtained coke, and using the correspondence between the raw coal and the cluster, converting the blending ratio of each of the raw coals in the actual measurement data D1 into the blending ratio of each cluster, and generating teacher data D2 in which the blending ratio of each converted cluster is used as an input and the grade value of the coke is used as an output.

[0046] The teacher data generation system 1 of this embodiment is a system for generating teacher data used in machine learning of a prediction model for predicting the quality of coke, and includes one or more processors that execute the above-mentioned teacher data generation method. The one or more processors realize an actual measurement data acquisition unit 10, a teacher data generation unit 11, a raw coal physical property data acquisition unit 12, a dimension reduction processing unit 13, a cluster analysis unit 14, and a relationship identification unit 15 in the system 1.

[0047] According to this teacher data generation method or system, it is possible to predict the quality value of coke based on the blending ratio of each of multiple raw coal types. Moreover, since the blending ratio of each of multiple raw coal types is converted into the blending ratio of each cluster classified based on the properties, it is possible to deal with cases where the properties are different even for the same brand, and it is possible to improve the prediction accuracy.

[0048] As in the teaching data generation method of the present embodiment, it is preferable that dimension reduction is performed by principal component analysis, and a predetermined number of principal components from the first rank onward are selected as feature data. In this way, it is preferable to use principal component analysis for dimension reduction.

[0049] As in the teaching data generation method of this embodiment, the step of acquiring actual measurement data may include the steps of extracting multiple coke quality values ​​having consecutively the same combination of blending ratios of multiple types of raw coal from the acquired actual measurement data D1, calculating a representative quality characteristic value based on the extracted multiple quality values, and converting the calculated representative value into new aggregated actual measurement data as a new quality value.

[0050] According to this method, it is possible to reduce or eliminate the variation in data, which occurs when the same compound has different quality values, and to improve prediction accuracy.

[0051] As in this embodiment, the teacher data generation method is a method for generating teacher data to be used in machine learning of a predictive model for predicting the quality of coke, and may include the steps of acquiring actual measurement data correlating the blending ratios of each of multiple types of raw coal with the quality value of the actually obtained coke, extracting from the acquired actual measurement data multiple quality values ​​of the coke that have the same consecutive combinations of blending ratios of each of the multiple types of raw coal, calculating a representative quality characteristic value based on the extracted multiple quality values, converting the calculated representative value into new aggregated actual measurement data as a new quality value, and generating teacher data that uses the blending ratios of each of the raw coals in the new actual measurement data as input and outputs the quality value of the coke. The teacher data generation system 1 of this embodiment is a system for generating teacher data used in machine learning of a prediction model for predicting the quality of coke, and includes one or more processors that execute the teacher data generation method. The one or more processors realize an actual measurement data acquisition unit 10 (having a same blend extraction unit 10a, a representative value determination unit 10b, and a data aggregation unit 10c) and a teacher data generation unit 11 in the system 1. According to this configuration, abnormal data contained in the actual measurement data and variations in quality values ​​can be suppressed, so that the training data used to train the prediction model becomes closer to the true value, thereby improving prediction accuracy.

[0052] The coke quality prediction model generation method of this embodiment includes a step of executing the above-mentioned teacher data generation method to generate teacher data D2, and a step of generating a prediction model 20 that predicts the coke quality value by performing machine learning using the generated teacher data D2. The coke quality prediction model generation system 2 of this embodiment includes a learning unit that generates a prediction model for predicting the quality value of the coke by performing machine learning using the training data generated by the training data generation method described above.

[0053] According to this coke quality prediction model generating method or system, it is possible to provide a prediction model 20 with improved prediction accuracy.

[0054] The coke quality prediction method of this embodiment includes the steps of acquiring the blending ratios of multiple types of raw coal representing the blend of raw materials for the coke to be predicted, converting the acquired blending ratios of the raw coal units into blending ratios of cluster units using a correspondence relationship previously associated with the raw coals and clusters, and predicting the coke quality value from the blending ratios of each converted cluster unit using a prediction model 20 constructed by machine learning so as to output a coke quality value from the blending ratios of each cluster unit. The coke quality prediction system 3 of this embodiment includes a prediction target data acquisition unit 30 that acquires the blending ratios of multiple types of raw coal that represent the blend of raw materials for the coke to be predicted, a conversion unit 31 that converts the acquired blending ratios of raw coal units into blending ratios of cluster units using a correspondence relationship previously established between the raw coals and clusters, and a prediction unit 32 that predicts the coke quality value from the blending ratios of each converted cluster unit using a prediction model 20 constructed by machine learning to output a coke quality value from the blending ratios of each cluster unit.

[0055] According to the method or system for generating a coke quality prediction model, it is possible to predict the quality value of coke based on the blending ratio of each of multiple types of raw material coal. Moreover, since the blending ratio of each of multiple types of raw material coal, which represents the blend of the raw materials of the coke to be predicted, is converted into the blending ratio of each cluster classified based on the properties, it is possible to deal with cases where the properties are different even for the same brand, and it is possible to improve the prediction accuracy.

[0056] The above method is executed by one or more processors. The program of this embodiment is a program that causes a computer (one or more processors) to execute the above method. Furthermore, a computer-readable temporary recording medium according to this embodiment stores the above program.

[0057] Although the embodiments of the present disclosure have been described above based on the drawings, the specific configurations should not be considered to be limited to these embodiments. The scope of the present disclosure is indicated not only by the description of the above-mentioned embodiments but also by the claims, and further includes all modifications within the meaning and scope equivalent to the claims.

[0058] For example, in the present embodiment, the grade value is the coke strength DI as an example, but the grade value is not limited to this, and any grade (physical property value) of the coke can be used.

[0059] The structures employed in each of the above embodiments can be employed in any of the other embodiments.

[0060] The specific configuration of each part is not limited to the above-described embodiment, and various modifications are possible without departing from the spirit of the present disclosure. [Explanation of symbols]

[0061] 1. Teacher data generation system 10. Actual data acquisition section 11 Teacher data generation unit 12. Coal property data acquisition section 13 Dimensionality reduction processing section 14 Cluster Analysis Section 15. Related Parties 2. Predictive model generation system 20 Predictive Models 21 Learning Department 30 Prediction target data acquisition section 31 Conversion section 32 Prediction Department 3 Coke quality prediction system

Claims

1. A method for generating training data to be used in machine learning of a prediction model for predicting a grade of coke, comprising: A step of acquiring actual measurement data correlating the blending ratios of each of the plurality of types of raw coal with the grade value of an actually obtained coke; extracting a plurality of quality values ​​of the coke in which the combinations of the blending ratios of the plurality of raw coals are consecutively the same from the acquired actual measurement data; calculating a representative quality characteristic value based on the extracted quality values; a step of converting the calculated representative value into new aggregated actual measurement data having a new grade value; A step of generating training data in which the blending ratios of each raw coal of the new measured data are input and the quality value of the coke is output; A method for generating teacher data, comprising:

2. A step of generating training data by executing the method according to claim 1; generating a prediction model for predicting a grade value of the coke by performing machine learning using the generated teacher data; The present invention relates to a method for generating a coke quality prediction model, the method comprising the steps of:

3. A step of acquiring a blend ratio of each of a plurality of types of raw coal representing a blend of raw materials of the coke to be predicted; A step of predicting a quality value of coke from the blending ratios of the respective types of raw coal obtained by using a prediction model constructed by machine learning using training data generated by the training data generation method according to claim 1 so as to output a quality value of coke from the blending ratios of the respective types of raw coal; The coke grade prediction method includes:

4. A system for generating training data used in machine learning of a prediction model for predicting the quality of coke, comprising: A training data generation system comprising one or more processors for executing the method of claim 1.

5. A coke quality prediction model generation system comprising: a learning unit that generates a prediction model for predicting a quality value of the coke by performing machine learning using training data generated by the method according to claim 1.

6. a prediction target data acquisition unit that acquires a blend ratio of each of a plurality of types of raw coal representing a blend of raw materials of a coke to be predicted; a prediction unit that predicts a quality value of coke from the blending ratios of the respective types of raw coal obtained by using a prediction model constructed by machine learning using training data generated by the training data generation method according to claim 1 so as to output a quality value of coke from the blending ratios of the respective types of raw coal; The coke quality prediction system includes:

7. A program for causing one or more processors to execute the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Coke quality prediction method and device and computer equipment

    CN110119783A

  • Apparatus for estimating coke strength

    JP1997165579A

  • Coke strength estimation device

    JP1997304376A

  • Apparatus for optimizing coal blend plan

    JP1998324876A

  • Device and method for preparing coal blending plan

    JP2001243301A