A method and system for predicting end point phosphorus content based on oxygen top blown converter steelmaking

CN122241170BActive Publication Date: 2026-09-18UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610146087.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-09-18
Estimated Expiration
2046-02-02

AI Technical Summary

Technical Problem

[0003]现有技术中,工业界尝试采用机器学习对多源异常、参数强耦合的氧气顶吹转炉炼钢法(Basic Oxygen Furnace,BOF)数据进行建模与多目标优化,以提升预测精度与稳健性,但工业噪声、样本分布偏斜及跨钢种/跨工况迁移仍是难点

Benefits of technology

[0031] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122241170B_ABST
    Figure CN122241170B_ABST
Patent Text Reader

Abstract

The application provides an endpoint phosphorus content prediction method and system based on oxygen top-blown converter steelmaking, and belongs to the field of steel smelting. The method collects multi-source process data in the historical production of the converter as features and corresponding endpoint phosphorus content data as targets, introduces mechanism information as additional features, pre-processes the data and unifies the caliber, then clusters to obtain three subsets, which are respectively medium-phosphorus, low-phosphorus and ultra-low-phosphorus subsets, and are divided into initial training sets and test sets; feature engineering is performed for feature screening to obtain formal training sets; N candidate endpoint phosphorus content prediction models are selected for each subset, and after training, verification and evaluation, the optimal prediction model of each subset is screened out, and testing and evaluation are performed; real-time multi-source process data of the current converter production is collected to determine which subset it belongs to, and the optimal prediction model corresponding to the selected subset is used to predict the endpoint phosphorus content. The application improves the prediction accuracy and precision of the endpoint phosphorus content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of iron and steel smelting, specifically relating to a method and system for predicting the final phosphorus content in oxygen top-blown converter steelmaking. Background Technology

[0002] Converter steelmaking uses molten iron, scrap steel, and ferroalloys as main raw materials. Through processes such as oxygen blowing, harmful elements like sulfur and phosphorus are removed from the molten iron, and carbon content and temperature are adjusted to ultimately produce steel that meets requirements, thus realizing the conversion of pig iron into steel. Rapid and efficient dephosphorization and decarburization are the core tasks of the converter. The final phosphorus content directly determines the compliance rate of electrical pure iron / high-purity iron, pipeline steel, etc., regarding inclusions and magnetic / mechanical properties. Therefore, the industry has long focused on mechanistic research and process optimization for "deeper and more stable dephosphorization," while employing various methods to predict the final phosphorus content of the converter.

[0003] In existing technologies, the industry has attempted to use machine learning to model and optimize data from the Basic Oxygen Furnace (BOF) steelmaking process, which has multiple anomalies and strong parameter coupling, in order to improve prediction accuracy and robustness. However, industrial noise, sample distribution skew, and migration across steel grades / operating conditions remain challenges. Among them, predictive control based on mechanism / thermodynamics can reproduce experimental or industrial average levels under "single operating conditions + stable furnace conditions," but its adaptability to industrial noise, non-equilibrium, and strong randomness is limited, and it is difficult to directly provide high-precision predictions for the "ultra-low phosphorus range" across steel grades. Black-box machine learning with a global single predictor uses the BOF process as a unified sample library and uses a single global model for endpoint phosphorus prediction and / or multi-objective optimization. It can filter out noise and improve the average error to a certain extent, but when the target value has the characteristics of "small value and concentrated distribution" (such as endpoint phosphorus in the tens of ppm range), the global model is easily "dragged" by the main peak sample, resulting in mismatch in the ultra-low phosphorus (ULP) range, and its interpretability is insufficient and its practical guidance is weak. It can be seen that the existing models, especially when facing the problem of industrial extreme dephosphorization, cannot achieve efficient, accurate, and interpretable endpoint phosphorus content prediction. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a method and system for predicting the final phosphorus content in oxygen top-blown converter steelmaking. Based on the Intelligent Dephosphorization Partitioning and Prediction Framework (iDePP), a domain division / clustering process is introduced into the final phosphorus content prediction model. Industrial samples are divided into three subsets: iDePP-MP, iDePP-LP, and iDePP-ULP. Feature selection, model optimization, and parameter tuning are performed within each subset to reduce model variance, improve interpretability, and enhance the accuracy, interpretability, and feasibility of final phosphorus content prediction.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:

[0006] In a first aspect, embodiments of the present invention provide a method for predicting the final phosphorus content in oxygen top-blown converter steelmaking, the method comprising the following steps:

[0007] Step S1: Collect multi-source process data from the historical production of the converter as features and corresponding endpoint phosphorus content data as targets to form raw data;

[0008] Step S2: Introduce the mechanism information as a supplementary feature into the original data to obtain the first dataset; preprocess the first dataset and standardize the criteria to obtain the second dataset;

[0009] Step S3: Perform iterative clustering on the second dataset to obtain three cluster subsets, which are designated as the medium phosphorus (MP) subset, the low phosphorus (LP) subset, and the ultra-low phosphorus (ULP) subset, respectively; and divide each subset into an initial training set and a test set.

[0010] Step S4: For the initial training set within each subset, feature engineering is used to filter features, resulting in the formal training set for each subset.

[0011] Step S5: Select N candidate endpoint phosphorus content prediction models for each subset; train the candidate models based on the formal training set of each subset, then use cross-validation to evaluate and optimize hyperparameters to confirm and select the optimal prediction model for each subset.

[0012] Step S6: Test the optimal prediction model using the test set of each subset, and evaluate the model after testing.

[0013] Step S7: Collect real-time multi-source process data of the current converter production and determine which subset (MP, LP, or ULP) the real-time multi-source process data belongs to; predict the final phosphorus content of the current converter based on the optimal prediction model corresponding to the selected subset.

[0014] As a preferred embodiment of the present invention, the multi-source process data in step S1 includes the amount of feed, the amount of air blown, the temperature of molten steel, the time, the furnace age, and the slag composition.

[0015] In a preferred embodiment of the present invention, the mechanism information includes slag system composition and alkalinity.

[0016] As a preferred embodiment of the present invention, the unification of the caliber mentioned in step S2 is to unify it to the throughput caliber based on the tonnage of molten steel.

[0017] In a preferred embodiment of the present invention, K-means++ unsupervised clustering is used when performing iterative clustering in step S3.

[0018] In a preferred embodiment of the present invention, the feature engineering in step S4 performs feature selection through correlation screening and recursive feature elimination cross-validation.

[0019] As a preferred embodiment of the present invention, the candidate models in step S5 include: Elastic Net, KNN, DT, RF, GB, XGBoost, LightGBM, SVM and NN models.

[0020] In a preferred embodiment of the present invention, the optimal model corresponding to the MP subset and LP subset after screening is the RF model, and the optimal model corresponding to the ULP subclass is the XGBoost model.

[0021] In a preferred embodiment of the present invention, step S6 includes evaluation metrics such as error metrics and engineering metrics; wherein, the error metrics include RMSE / MAE; the engineering metrics include hit rate (HR), and the preset HR threshold for each subset is as follows: HR MP ±0.002%, HR LP For ±0.0015%, HR ULP It is ±0.0006%.

[0022] Secondly, embodiments of the present invention also provide a system for predicting the final phosphorus content in oxygen top-blown converter steelmaking. The system includes: a data acquisition module, a mechanism supplementation module, a preprocessing module, a subset classification module, a feature selection module, a model selection module, and a final phosphorus content prediction module; wherein,

[0023] The data acquisition module is used to collect multi-source process data from the historical production of the converter as features and corresponding endpoint phosphorus content data as targets to form raw data.

[0024] The mechanism supplementation module is used to introduce mechanism information as supplementary features into the original data to obtain the first dataset;

[0025] The preprocessing module is used to preprocess the first dataset and standardize its definition to obtain the second dataset;

[0026] The subset classification module is used to perform iterative clustering on the second dataset to obtain three cluster subsets, which are respectively used as the medium phosphorus (MP) subset, the low phosphorus (LP) subset, and the ultra-low phosphorus (ULP) subset; and each subset is divided into an initial training set and a test set.

[0027] The feature filtering module is used to filter features from the initial training set within each subset through feature engineering, thereby obtaining the formal training set for each subset.

[0028] The model selection module is used to select N candidate endpoint phosphorus content prediction models for each subset; the candidate models are trained based on the formal training set of each subset, and then cross-validation is used for evaluation, hyperparameter optimization and confirmation to select the optimal prediction model for each subset; the optimal prediction model is tested using the test set of each subset, and the tested model is evaluated.

[0029] The endpoint phosphorus content prediction module is used to collect real-time multi-source process data of the current converter production and determine which subset (MP, LP, or ULP) the real-time multi-source process data belongs to; based on the optimal prediction model corresponding to the selected subset, the endpoint phosphorus content of the current converter is predicted.

[0030] The method and system for predicting the final phosphorus content in oxygen top-blown converter steelmaking provided in this invention embodiment employs feature engineering based on expert knowledge to collect a full database of converter steelmaking samples. All feed and gas flow rates are standardized to "throughput per ton of steel," constructing mechanism-derived features such as CaO, MgO, Al2O3, SiO2, and slag basicity, and performing anomaly handling / missing data imputation. Unsupervised domain segmentation (e.g., K-means++) is then performed, automatically obtaining MP, LP, and ULP subsets based on process / target distributions, rather than pre-labeling or simply coarsely classifying by chemical composition. Within each domain or subset, feature selection, algorithm optimization, hyperparameter search, and cross-validation are performed separately to obtain a subset-specific optimal model (rather than a single global predictor). Differentiated hit rate thresholds and evaluation windows (MP / LP / ULP) are set for different subsets. Different ± error bands) use hit rate (HR) as the core indicator for production decision-making, and run in parallel with RMSE / MAE to differentiate engineering evaluation criteria; based on the optimal model obtained from each subset, the final phosphorus content of the corresponding interval is predicted. Applying this invention to actual production, compared to a single global model, the MAE of the three subsets is reduced to 0.0011% / 0.0007% / 0.0004%, and the ULP-HR reaches 82.7% (±6 ppm), meeting the "hit" assessment of the extremely low P engineering window. On a 200 t converter, a finished product P of 0.0019% and a BOF endpoint of 0.0016% were obtained, with a model-production line deviation of approximately 4 ppm. SEM / EDS confirmed the scarcity of inclusions, demonstrating the method's transferable and scalable engineering usability. In addition to RMSE / MAE, establishing domain-differentiated HR thresholds (MP±0.002%, LP±0.0015%, ULP±0.0006%) accurately reflects the control quality in the extremely low P range. Through "automatic domain segmentation + subset specialization," the influence of mixed distributions and heterogeneous operating conditions is weakened, facilitating migration across steel grades / plants and extending to other data-scarce metallurgical processes, thus improving the robustness and scalability of the prediction results.

[0031] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a flowchart of the method for predicting the final phosphorus content in oxygen top-blown converter steelmaking, as described in an embodiment of the present invention. Detailed Implementation

[0034] After discovering the aforementioned problems, the inventors of this application conducted a detailed study on existing methods for predicting phosphorus content at the BOF endpoint. The study found that industrial data actually exhibits multimodality during the BOF process. Automatically dividing the raw data into domains yields three categories (medium / low / ultra-low P) that correspond to metallurgical classifications, with significant differences in the distribution of each category. If mixed modeling or only coarse classification based on composition is used, the prediction accuracy is difficult to achieve the target, revealing the inherent limitations of "whole-database modeling." Existing methods first perform unsupervised clustering on the entire database sample, then feed each cluster into predictors of the same or similar structure, separating multimodal data through "sample similarity" in the hope of reducing variance and overfitting. However, this method of applying different subsets to the same prediction model fails to achieve subset specialization, resulting in mutual influence of prediction results and preventing accurate prediction. The comparison results show that after automatically dividing the data into "medium / low / ultra-low P" categories and training within subsets, the MAE of the three categories decreased significantly compared to a single global predictor, and the HR of the ULP subset was greatly improved (reaching 82.7% within ±6 ppm), which proves the necessity of "mixed distributions should not be modeled across the entire database".

[0035] For different domains, the hit rate (HR) needs to have differentiated thresholds (e.g., stricter for ULP). This is a key indicator of concern in industrial performance evaluation, not just the root mean square error (RMSE) / mean absolute error (MAE). The endpoint P-value is small and concentrated, and a single model across the entire database will be influenced by the peak sample, making it difficult to account for the tail of ULP; therefore, MAE / RMSE alone is insufficient to reflect engineering value, and HR needs to be introduced as a hard indicator.

[0036] Meanwhile, within the ULP range, there exists a coupling window between the lime / limestone ratio, slag quantity, oxygen quantity, and blowing time. When O2 is sufficient, appropriately extending the oxygen blowing time and controlling the limestone proportion to a small ratio (e.g., 0-10, relative to the total lime content within the range of 40-60%), combined with a large slag quantity (≥40,000), yields the optimal contribution to dephosphorization, which also supports the effectiveness of the dual-slag method. This type of quantitative window requires the model to have interpretability, enabling the empirical rules to be applied to executable parameters. In contrast, windowed empirical control based on mechanism / slag system, which provides empirical windows and manually adjusts parameters around mechanistic quantities such as slag system (CaO-FeO-SiO2…), oxygen flow, and time, is suitable for average level control, but often struggles to provide precise hit criteria at the sample level when facing strong noise and skewed distribution in the extremely low P range.

[0037] It should be noted that the defects in the above-mentioned prior art solutions are all the result of the inventors' practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present invention in the following text should be the inventors' contributions to the present invention.

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can also be combined with each other.

[0039] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of this invention, the terms "first," "second," "third," "fourth," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0040] Based on the above in-depth analysis, this invention provides a method and system for predicting the final phosphorus content in oxygen top-blown converter steelmaking. For the first time, "automatic domain segmentation + subset-specific integrated learning" is applied to complex BOF control problems. This invention avoids the accuracy and stability bottlenecks of "full-database black box" when dealing with multimodal / strongly random / distributed BOF data. It achieves high-precision prediction and high hit rate (HR) in the ULP limit range, forming a universal solution that can be implemented and transferred to different steel plants and different steel grades.

[0041] like Figure 1 As shown, the method for predicting the final phosphorus content in oxygen top-blown converter steelmaking includes the following steps:

[0042] Step S1: Collect multi-source process data from the historical production of the converter as features and corresponding endpoint phosphorus content data as targets to form raw data.

[0043] In this step, the multi-source process data are all data related to phosphorus content, including feed rate, air blowing rate, molten steel temperature, time, furnace age, slag composition, etc.

[0044] Step S2: Introduce the mechanism information as a supplementary feature into the original data to obtain the first dataset; preprocess the first dataset and standardize the criteria to obtain the second dataset.

[0045] In this step, after introducing supplementary features into the original data, mechanism-related features are added to the existing features, making the consideration of phosphorus content-related factors more comprehensive. The mechanism information includes slag system composition, alkalinity, etc.; wherein, the slag system composition includes CaO / MgO / ... etc., for example, expressed as CaO%, MgO%, Al2O3%, SiO2%, etc.

[0046] The preprocessing includes outlier identification and cleaning, and missing value completion, to standardize specifications and reduce noise. Specifically, standardization refers to unifying the specifications to the throughput based on molten steel tones, thereby eliminating the impact of different converter specifications in the converter process.

[0047] Step S3: Perform iterative clustering on the second dataset to obtain three cluster subsets, which are designated as the medium phosphorus (MP) subset, the low phosphorus (LP) subset, and the ultra-low phosphorus (ULP) subset, respectively; and divide each subset into an initial training set and a test set.

[0048] In this step, K-means++ unsupervised clustering is used during iterative clustering. For the three subsets, metallurgical knowledge verification confirms that the three subsets are consistent with metallurgical classifications, resulting in the iDePP-MP / LP / ULP subsets. By dividing them into three distinct subsets, the estimation bias caused by the "mixed distribution" is eliminated, while simultaneously extracting the data foundation for subsequent subset-based modeling.

[0049] Regarding the verification of metallurgical knowledge, a common classification is adopted. The three subsets correspond highly to traditional metallurgical steel categories: iDePP-MP mainly covers structural steels such as Q235 and Q345, and the endpoint P is in the medium range; iDePP-LP mainly covers pipeline steels such as X42–X70, and the endpoint P is in the low range; iDePP-ULP mainly covers ultra-low phosphorus system steels such as FS and TLS, and the endpoint P is in the ultra-low range.

[0050] Step S4: For the initial training set within each subset, feature engineering is used to filter features to obtain the formal training set for each subset.

[0051] In this step, the feature engineering uses relevance screening and recursive feature elimination cross-validation (RFECV) to screen the corresponding features in each subset, and retains the screened features to obtain the formal training set.

[0052] Step S5: Select N candidate endpoint phosphorus content prediction models for each subset; train the candidate models based on the formal training set of each subset, then use cross-validation to evaluate and optimize hyperparameters to confirm the optimal prediction model for each subset.

[0053] In this step, candidate models are selected based on subsets, and the statistical-mechanistic differences between subsets are matched to obtain the optimal dedicated model with the lowest MAE / stable HR. In an executable embodiment, the candidate models include: ElasticNet, KNN, DT, RF, GB, XGBoost, LightGBM, SVM, NN, etc.; the cross-validation uses a 10-fold cross-validation model. After screening, RF is adapted to MP / LP, and XGBoost is adapted to ULP.

[0054] Step S6: Test the optimal prediction model using the test set of each subset, and evaluate the model after testing.

[0055] The evaluation indicators and standards for this step are as follows:

[0056] Error metrics include RMSE / MAE (10-fold and independent test set); engineering metrics include HR, with thresholds defined for each subset: MP ± 0.002%, LP ± 0.0015%, and ULP ± 0.0006%. Based on the test and evaluation results, and compared with existing global models, this embodiment divides the overall sample into three subsets and selects the optimal model for each subset. The MAE of the three subsets is reduced to 0.0011 / 0.0007 / 0.0004%, and the HR of the ULP subset reaches 82.7% (±6 ppm).

[0057] Step S7: Collect multi-source process data of the current converter production and determine which subset of MP, LP or ULP the multi-source process data belongs to; predict the final phosphorus content of the current converter according to the optimal prediction model corresponding to the selected subset.

[0058] Based on the same idea, this invention also provides a system for predicting the final phosphorus content in oxygen top-blown converter steelmaking. The system includes: a data acquisition module, a mechanism supplementation module, a preprocessing module, a subset classification module, a feature selection module, a model selection module, and a final phosphorus content prediction module.

[0059] The data acquisition module is used to collect multi-source process data from the historical production of the converter as features and corresponding endpoint phosphorus content data as targets to form raw data.

[0060] The mechanism supplementation module is used to introduce mechanism information as supplementary features into the original data to obtain the first dataset;

[0061] The preprocessing module is used to preprocess the first dataset and standardize its definition to obtain the second dataset;

[0062] The subset classification module is used to perform iterative clustering on the second dataset to obtain three cluster subsets, which are respectively used as the medium phosphorus (MP) subset, the low phosphorus (LP) subset, and the ultra-low phosphorus (ULP) subset; and each subset is divided into an initial training set and a test set.

[0063] The feature filtering module is used to filter features from the initial training set within each subset through feature engineering, thereby obtaining the formal training set for each subset.

[0064] The model selection module is used to select N candidate endpoint phosphorus content prediction models for each subset; the candidate models are trained based on the formal training set of each subset, and then cross-validation is used for evaluation, hyperparameter optimization and confirmation to select the optimal prediction model for each subset; the optimal prediction model is tested using the test set of each subset, and the tested model is evaluated.

[0065] The endpoint phosphorus content prediction module is used to collect real-time multi-source process data of the current converter production and determine which subset (MP, LP, or ULP) the real-time multi-source process data belongs to; based on the optimal prediction model corresponding to the selected subset, the endpoint phosphorus content of the current converter is predicted.

[0066] In this embodiment, each module is implemented using a processor, with additional memory added as needed for storage. The processor can be, but is not limited to, a microprocessor (MPU), a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, etc. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0067] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0068] It should also be noted that the endpoint phosphorus content prediction system for oxygen top-blown converter steelmaking described in this embodiment corresponds to the endpoint phosphorus content prediction method for oxygen top-blown converter steelmaking. The description and limitations of the method also apply to the system, and will not be repeated here.

[0069] The method and system for predicting the final phosphorus content in oxygen top-blown converter steelmaking, as described in this embodiment of the invention, were applied to a 200-ton converter. First, K-means++ was used to automatically segment 5102 industrial samples. Modeling and evaluation were performed for the MP, LP, and ULP intervals to obtain three subsets iDePP-MP, iDePP-LP, and iDePP-ULP that are consistent with metallurgical experience. Feature selection, model optimization, and parameter tuning were performed within each subset. Among nine candidate models, including Elastic Net, KNN, DT, RF, GB, XGBoost, LightGBM, SVM, and NN, the optimal model was selected for each subset. A 10-fold model was used for cross-validation. After selection, the optimal model for the MP and LP subsets was the RF model, and the optimal model for the ULP subset was the XGBoost model. Multi-source process data from the current converter production is collected, and if the current data characteristics belong to the low-phosphorus subset, then the current multi-source process data is used as the feature input to the RF model, outputting the predicted endpoint phosphorus content. In another production run, the collected multi-source process data belongs to the ULP subset. In this case, the corresponding multi-source process data is used as the feature input to the XGBoost model, outputting the predicted endpoint phosphorus content. Different prediction models are selected for different feature parameters, effectively improving the accuracy and precision of the prediction.

[0070] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed, and is not intended to limit the scope of the claimed invention, but merely to illustrate preferred embodiments of the invention. Those skilled in the art should understand that the scope of the invention is not limited to the specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

Claims

1. A method for predicting the final phosphorus content in oxygen top-blown converter steelmaking, characterized in that, The method includes the following steps: Step S1: Collect multi-source process data from the historical production of the converter as features and corresponding endpoint phosphorus content data as targets to form raw data; Step S2: Introduce the mechanism information as a supplementary feature into the original data to obtain the first dataset; preprocess the first dataset and standardize the criteria to obtain the second dataset; Step S3: Perform iterative clustering on the second dataset to obtain three cluster subsets, which are designated as the medium phosphorus (MP) subset, the low phosphorus (LP) subset, and the ultra-low phosphorus (ULP) subset, respectively; and divide each subset into an initial training set and a test set. Step S4: For the initial training set within each subset, feature engineering is used to filter features, resulting in the formal training set for each subset. Step S5: Select N candidate endpoint phosphorus content prediction models for each subset; train the candidate models based on the formal training set of each subset, then use cross-validation to evaluate and optimize hyperparameters to confirm and select the optimal prediction model for each subset. Step S6: Test the optimal prediction model using the test set of each subset, and evaluate the model after testing. Step S7: Collect real-time multi-source process data of the current converter production and determine which subset (MP, LP, or ULP) the real-time multi-source process data belongs to; predict the final phosphorus content of the current converter based on the optimal prediction model corresponding to the selected subset.

2. The method according to claim 1, characterized in that, The multi-source process data mentioned in step S1 includes the amount of feed, the amount of air blown, the temperature of molten steel, the time, the furnace age, and the slag composition.

3. The method according to claim 1, characterized in that, The mechanistic information includes the slag system composition and alkalinity.

4. The method according to claim 1, characterized in that, The standardization mentioned in step S2 refers to standardizing the throughput based on the tonnage of molten steel.

5. The method according to claim 1, characterized in that, In step S3, K-means++ unsupervised clustering is used when performing iterative clustering.

6. The method according to claim 1, characterized in that, The feature engineering described in step S4 involves feature selection through correlation screening and recursive feature elimination cross-validation.

7. The method according to claim 1, characterized in that, The candidate models in step S5 include: Elastic Net, KNN, DT, RF, GB, XGBoost, LightGBM, SVM, and NN models.

8. The method according to claim 7, characterized in that, After filtering, the optimal model for the MP subset and LP subset is the RF model, and the optimal model for the ULP subset is the XGBoost model.

9. The method according to claim 1, characterized in that, Step S6 evaluates metrics including error metrics and engineering metrics; the error metrics include RMSE / MAE; the engineering metrics include hit rate (HR), and preset HR thresholds for each subset are as follows: HR MP ±0.002%, HR LP For ±0.0015%, HR ULP It is ±0.0006%.

10. A system for predicting the final phosphorus content in oxygen top-blown converter steelmaking, characterized in that, The system includes: a data acquisition module, a mechanism supplementation module, a preprocessing module, a subset classification module, a feature selection module, a model selection module, and an endpoint phosphorus content prediction module; wherein, The data acquisition module is used to collect multi-source process data from the historical production of the converter as features and corresponding endpoint phosphorus content data as targets to form raw data. The mechanism supplementation module is used to introduce mechanism information as supplementary features into the original data to obtain the first dataset; The preprocessing module is used to preprocess the first dataset and standardize its definition to obtain the second dataset; The subset classification module is used to perform iterative clustering on the second dataset to obtain three cluster subsets, which are respectively used as the medium phosphorus (MP) subset, the low phosphorus (LP) subset, and the ultra-low phosphorus (ULP) subset; and each subset is divided into an initial training set and a test set. The feature filtering module is used to filter features from the initial training set within each subset through feature engineering, thereby obtaining the formal training set for each subset. The model selection module is used to select N candidate endpoint phosphorus content prediction models for each subset; the candidate models are trained based on the formal training set of each subset, and then cross-validation is used for evaluation, hyperparameter optimization and confirmation to select the optimal prediction model for each subset; the optimal prediction model is tested using the test set of each subset, and the tested model is evaluated. The endpoint phosphorus content prediction module is used to collect real-time multi-source process data of the current converter production and determine which subset (MP, LP, or ULP) the real-time multi-source process data belongs to; based on the optimal prediction model corresponding to the selected subset, the endpoint phosphorus content of the current converter is predicted.