Converter oxygen lance control method and system based on multi-modal large model

Through data processing and iterative optimization of multimodal large models, the problems of unclear reasoning logic and poor adaptability in converter oxygen lance control were solved, achieving high-precision and stable intelligent control.

CN122490219APending Publication Date: 2026-07-31UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2026-05-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing converter oxygen lance control technology relies on a black box model and lacks reasoning logic that matches the smelting process. This makes it difficult to trace and verify the control process, and the model has poor adaptability to complex and ever-changing on-site conditions. The control accuracy drops significantly when the operating conditions fluctuate, making it impossible to achieve high-precision intelligent control.

Method used

By adopting a multimodal large model, a structured decision sequence is generated by collecting and preprocessing historical multimodal converter data. Through error evaluation and iterative optimization, a converter oxygen lance control model is constructed to ensure that the control decisions have clear reasoning basis and logical links, thereby enhancing the model's adaptability to various operating conditions.

Benefits of technology

It improves the interpretability and traceability of oxygen lance control decisions, significantly enhances model control accuracy, and achieves stable and reliable high-precision intelligent control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490219A_ABST
    Figure CN122490219A_ABST
Patent Text Reader

Abstract

This invention provides a converter oxygen lance control method and system based on a multimodal large model, relating to the field of converter steelmaking control technology. The method includes: inputting historical multimodal input data into a baseline model when a first prompt word is provided; fine-tuning the baseline model according to a first structured decision sequence; inputting historical multimodal input data into a preheating model when a second prompt word is provided, and performing error evaluation on the second structured decision sequence; updating the second prompt word to obtain a third prompt word, and inputting the third prompt word and historical multimodal input data into the preheating model; fine-tuning the baseline model according to the third structured decision sequence to obtain a converter oxygen lance control model; and inputting preprocessed real-time multimodal converter data into the converter oxygen lance control model when a fourth prompt word is provided, and outputting converter oxygen lance control decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of converter steelmaking control technology, and in particular to a converter oxygen lance control method and system based on a multimodal large model. Background Technology

[0002] Converter steelmaking is a core process in modern steel production. As the core equipment in converter steelmaking, the oxygen lance's control directly determines the furnace reaction efficiency, final steel quality, smelting safety, and overall energy consumption during the blowing process. Improper oxygen lance height control can easily lead to problems such as insufficient molten pool stirring, prolonged smelting cycle, slag splashing, and accelerated furnace lining erosion. Therefore, achieving precise, intelligent, and adaptive control of the converter oxygen lance is a key technological requirement for improving the automation and intelligence level of converter steelmaking, reducing production costs, and ensuring continuous and stable production.

[0003] At present, converter oxygen lance control technology has gradually developed from traditional manual operation to automation and intelligence. The mainstream technical paths cover three categories: manual experience control, rule-based and static model-based automated control, and intelligent control based on traditional machine learning and deep learning. Among them, manual experience control relies on the operator to observe the furnace flame on site and manually adjust the oxygen lance height in combination with process parameters. Rule-based and model-based automated control achieves automatic adjustment of oxygen lance height by setting up process mapping rules and static models and combining sensor data. Traditional machine learning and deep learning solutions attempt to integrate furnace flame visual data and process parameter time series data to achieve end-to-end prediction of oxygen lance control actions through neural network models. All kinds of technical solutions have promoted the development of converter oxygen lance control at different stages.

[0004] However, existing control methods mostly rely on the output results of black box models and lack reasoning logic that matches the smelting process, making it difficult to trace and verify the control process. At the same time, the models are not adaptable to complex and ever-changing on-site conditions, and the control accuracy drops significantly when the conditions fluctuate, making it impossible to stably achieve high-precision intelligent oxygen lance control. Summary of the Invention

[0005] To address the technical problems of existing control methods relying heavily on black-box model outputs and lacking reasoning logic that matches the smelting process, making it difficult to trace and verify the control process, and the poor adaptability of the models to complex and changing on-site conditions, resulting in a significant decrease in control accuracy when conditions fluctuate, thus failing to stably achieve high-precision intelligent oxygen lance control, this invention provides a converter oxygen lance control method and system based on a multimodal large model.

[0006] The technical solutions provided by the embodiments of the present invention are as follows: The first aspect of this invention provides a converter oxygen lance control method based on a multimodal large model, comprising: S1: Collect historical multimodal converter data; S2: Preprocess the historical multimodal converter data to obtain historical multimodal input data; S3: Given the first prompt word, input the historical multimodal input data into the baseline model and output the first structured decision sequence; S4: Based on the first structured decision sequence, fine-tune the baseline model to obtain the preheating model; S5: When a second prompt word is provided, input the historical multimodal input data into the preheating model, output the second structured decision sequence, and perform error evaluation on the second structured decision sequence to obtain the error evaluation result; S6: Based on the error assessment results, update the second prompt word to obtain the third prompt word, and input the third prompt word and historical multimodal input data into the warm-up model to output the third structured decision sequence; S7: Based on the third structured decision sequence, the baseline model is fine-tuned to obtain the converter oxygen lance control model; S8: Acquire real-time multimodal converter data; S9: When a fourth prompt is provided, the preprocessed real-time multimodal converter data is input into the converter oxygen lance control model, and the converter oxygen lance control decision is output.

[0007] A second aspect of the present invention provides a converter oxygen lance control system based on a multimodal large model, comprising: processor; The memory stores computer-readable instructions, which, when executed by the processor, implement the converter oxygen lance control method based on a multimodal large model as described in the first aspect.

[0008] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the converter oxygen lance control method based on a multimodal large model as described in the first aspect.

[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, by inputting training data into the basic model and generating a structured decision sequence under the guidance of prompts containing real gun control actions, the control decisions have clear reasoning basis and logical links, effectively improving the interpretability and traceability of the decisions. At the same time, a training dataset is constructed based on historical multimodal converter data, and the preheating model is iteratively optimized according to the error evaluation results, enhancing the model's adaptability to various working conditions, significantly improving the model's control accuracy, and realizing stable and reliable high-precision intelligent control of the oxygen lance. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a converter oxygen lance control method based on a multimodal large model, provided in an embodiment of the present invention.

[0012] Figure 2 This is a schematic diagram of a converter oxygen lance control system based on a multimodal large model, provided for an embodiment of the present invention. Detailed Implementation

[0013] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0014] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0015] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0016] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0017] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0018] Reference manual attached Figure 1 The diagram shows a flowchart of a converter oxygen lance control method based on a multimodal large model provided by an embodiment of the present invention.

[0019] This invention provides a converter oxygen lance control method based on a multimodal large-scale model. This method can be implemented by a converter oxygen lance control device based on a multimodal large-scale model, which can be a terminal or a server. The processing flow of the converter oxygen lance control method based on a multimodal large-scale model may include the following steps: S1: Collect historical multimodal converter data.

[0020] Optionally, historical multimodal converter data specifically includes: visual modal input data and text modal input data.

[0021] The visual modal input data specifically includes: keyframe sequences of video clips of the furnace flame.

[0022] The text modal input data specifically includes: historical oxygen lance height sequence, blowing time, process parameter summary, and structured prompt information.

[0023] For example, taking data from 80 heats of real converters collected by a steel plant, the original sampling frequency of the oxygen lance height sensor is 4 seconds. To improve the alignment accuracy and sample density between the flame image and the time-series data, linear interpolation is first applied to the oxygen lance height sequence, resampling it to a 1-second frequency. Then, the corresponding flame video frames per second are extracted as visual input. To avoid data leakage due to the introduction of future information through interpolation, no interpolation is performed within the 3-second interval before predictive control.

[0024] S2: Preprocess the historical multimodal converter data to obtain historical multimodal input data.

[0025] In this embodiment of the invention, a unified preprocessing operation is performed on historical multimodal converter data, which can remove redundant information and abnormal interference content in the original data, standardize the format and feature expression of various modal data, eliminate the differences between different types of operating condition data, and improve the overall data quality and information effectiveness.

[0026] S3: Given the first prompt word, input the historical multimodal input data into the baseline model and output the first structured decision sequence.

[0027] It should be noted that the first prompt word is based on the initial prompt word template. After filling in the corresponding historical multimodal data information, additional content related to real gun control actions is added to form a complete prompt content. This integrates working condition information and standard gun control references and is used in the model training phase with truth value guidance.

[0028] Optionally, the baseline model specifically includes a visual encoding module and a language generation module.

[0029] Optionally, the visual encoding module is specifically used to: perform unified feature encoding on the image sequence.

[0030] The language generation module is specifically used to: fuse visual features and structured text prompts based on the autoregressive language model structure to obtain a structured decision sequence.

[0031] Among them, the autoregressive language model structure is the core architecture of the language generation module in the multimodal large model. It follows the temporal dependency logic of sequence generation. In the process of generating structured decision sequences, it outputs semantic units one by one. The prediction of the current semantic unit in each generation step is based on all the semantic units generated previously, the visual modal features extracted by the visual encoding module, and the structured text modal input features. By utilizing the temporal correlation of the generated information, it ensures that the structured decision sequence of converter oxygen lance control output by the model is logically coherent and consistent, which fits the process reasoning logic of converter steelmaking and realizes the orderly mapping from multimodal input to structured decision output.

[0032] The first structured decision sequence is as follows:

[0033] in, t Indicates time, Y t express t The first structured decision sequence at time step. r t express t The intermediate reasoning process at each moment, s t express t The refining stage of time, c t express t The furnace condition or abnormal condition at any given time. d t express t Constant control of direction, a t express t The oxygen lance control actions at all times.

[0034] In this embodiment of the invention, under the condition of incorporating real lance control action prompts, a basic model with a visual encoding unit and a text generation unit is used to generate structured decision content. On the one hand, the visual encoding unit performs unified feature extraction on the continuous flame image at the furnace opening, making full use of image information to reflect the actual working conditions inside the furnace. On the other hand, the visual features and structured text prompts are integrated in an orderly generation manner, sequentially outputting complete decision content including the reasoning process, smelting stage, furnace condition, operation direction, and specific actions. This allows the model to stably learn the correspondence between multi-dimensional data and oxygen lance control under the supervision of real operation information, while ensuring that the decision logic is coherent and conforms to the actual steelmaking process. This provides a standard and reliable reference result for subsequent model optimization, effectively improving the model's learning efficiency and the interpretability of control decisions.

[0035] S4: Based on the first structured decision sequence, fine-tune the baseline model to obtain the preheating model.

[0036] In this embodiment of the invention, the parameter optimization of the basic model based on the structured decision sequence generated by the model enables the model to gradually learn and solidify the mapping law from multimodal operating condition information to standardized control decisions under the supervision of real operating data. At the same time, it stably grasps the reasoning logic that conforms to the steelmaking process, enabling the model to initially have the ability to independently generate reasonable and coherent oxygen lance control decisions. This provides a reliable basic model for subsequent iterative optimization without real prompts, and significantly improves the convergence speed of subsequent model training and the final control effect.

[0037] S5: When a second prompt word is provided, input the historical multimodal input data into the preheating model, output the second structured decision sequence, and perform error evaluation on the second structured decision sequence to obtain the error evaluation result.

[0038] It should be noted that the second prompt word is generated based on a unified initial prompt word template and is populated with content information from historical multimodal data of the current batch. No real gun control action information is added. It is the basic prompt text for model autonomous reasoning and error evaluation without truth value guidance.

[0039] In this embodiment of the invention, based on the configuration of the second prompt word, the preprocessed historical multimodal input data is sent to the preheating model to generate the corresponding second structured decision sequence and conduct error evaluation. This can objectively test the autonomous decision-making level of the model under given guidance conditions, accurately identify the deviations and defects in the gun control reasoning process, and comprehensively quantify the matching degree between the model output results and actual smelting requirements. This provides an objective basis for subsequent targeted adjustment of prompt words and optimization of model reasoning logic, ensuring the directionality and effectiveness of subsequent model iteration optimization.

[0040] S6: Based on the error assessment results, update the second prompt word to obtain the third prompt word, and input the third prompt word and historical multimodal input data into the warm-up model to output the third structured decision sequence.

[0041] In one possible implementation, S6 specifically includes sub-steps S601 to S605: S601: Based on the error assessment results, the second prompt word is updated using different strategies to obtain the third prompt word.

[0042] Optionally, the error assessment results specifically include: a first error assessment result, a second error assessment result, and a third error assessment result.

[0043] When the error assessment result is the first error assessment result, the second structured decision sequence is saved as the third structured decision sequence.

[0044] When the error assessment result is the second error assessment result, the second prompt word is updated through a local correction strategy to prompt the warm-up model to adjust the manipulation intensity of the second structured decision sequence.

[0045] The local correction strategy is a prompt word update strategy used when the model output result is judged as Correctable. This strategy is suitable for scenarios where the model's overall judgment of the smelting stage and furnace condition is generally correct, and only the value of the oxygen lance control action has a moderate error compared with the true value, requiring only a fine adjustment of the control force. Specifically, guiding statements are added to the prompt words input to the preheating model to clearly indicate that the model's overall judgment of the current converter smelting condition has no significant deviation and there is no need to re-infer the smelting stage and furnace condition. Only a slight correction is needed to the specific adjustment force under the oxygen lance control direction based on the original inference result. This guides the model to focus on optimizing the control force to generate reasonable oxygen lance control actions, ensuring the efficiency of model inference and enabling precise fine-tuning of control actions based on error feedback, thereby improving the accuracy of the model's lance control decision.

[0046] When the error assessment result is the third error assessment result, the second prompt word is updated through the rethinking strategy to indicate that the previous judgment of the warm-up model was wrong.

[0047] The rethinking strategy is a prompt word update strategy used when the model output result is determined to be Invalid. It is applicable to scenarios where the oxygen lance control action output by the model has a significant error compared with the true value and there is a fundamental error in the judgment of the current converter smelting condition. Specifically, it adds a clear guiding statement to the prompt words input to the preheating model, indicating that the model's previous reasoning results regarding the smelting stage and furnace condition are incorrect. The original decision path needs to be discarded, and the current smelting stage, furnace condition or abnormal state needs to be re-judged by combining the visual features of the furnace flame and process parameters such as oxygen lance height and blowing time from the multimodal input. Then, based on the new condition judgment results, a new lance control decision path that conforms to the process logic is generated, thereby correcting the model's erroneous judgment and ensuring that the final output oxygen lance control action matches the actual smelting requirements.

[0048] It should be noted that the model output results are categorized into three states: Success, Correctable, and Invalid. If the result is Success, it indicates that the error between the analyzed oxygen lance control action and the actual control action is less than a set threshold. In this case, the current training data processing is complete, the result is saved, and the model proceeds directly to the next training data iteration. If the result is Correctable, it indicates that the error between the analyzed oxygen lance control action and the actual control action is within a moderate range. The model's overall judgment of the current smelting condition is generally correct, requiring only minor adjustments to the oxygen lance control force. In this case, a local correction strategy is used to update the input prompts, guiding the model to fine-tune the control force based on the original inference. If the result is Invalid, it indicates that the error between the analyzed oxygen lance control action and the actual control action is significant, indicating an error in the model's judgment of the current smelting condition. In this case, a rethinking strategy is used to update the input prompts, guiding the model to discard the previous inference results, re-infer the smelting stage and furnace condition, and generate a new control decision path.

[0049] S602: Input historical multimodal input data and third prompt words into the warm-up model and output the decision results.

[0050] S603: Analyze the decision results.

[0051] S604: Repeat steps S601 to S603 until parsing is successful.

[0052] It should be noted that the analysis is considered successful when the decision obtained by the model inference is consistent with the gun control action.

[0053] S605: Based on the parsing result, determine whether the verification was successful. If yes, save the decision result as the third structured sequence. Otherwise, determine whether the maximum number of verifications has been reached. If yes, inject the actual gun control action into the third prompt word, and input it along with the historical multimodal input data into the preheating model to obtain the third structured sequence. Otherwise, return to step S601.

[0054] In this embodiment of the invention, based on the error assessment results, different optimization strategies are adopted to update the prompt content. Depending on the severity of the deviation, different adjustment methods are adopted, such as retaining valid results, fine-tuning the control intensity locally, and comprehensively re-deducing the working conditions. Combined with the iterative analysis and maximum verification limit mechanism, new structured decision content is dynamically generated. This can efficiently complete detailed optimization when the working condition judgment is correct, and thoroughly correct the reasoning logic when the working condition perception is biased. It continuously corrects the model output deviation, takes into account the reasoning efficiency and decision correction effect, and continuously optimizes the model's ability to judge the smelting working conditions and the level of gun control decision generation, providing high-quality decision sample support for subsequent in-depth model optimization.

[0055] S7: Based on the third structured decision sequence, the baseline model is fine-tuned to obtain the converter oxygen lance control model.

[0056] In one possible implementation, S7 specifically includes sub-steps S701 and S702: S801: Construct a training sample set based on the third structured decision sequence.

[0057] S802: Input the training sample set into the baseline model, and obtain the converter oxygen lance control model through the supervised fine-tuning algorithm.

[0058] Among them, the supervised fine-tuning algorithm constructs a training sample set with the adjustment results after iterative correction as the supervision signal, inputs multimodal converter data and prompt information into the model, constructs the optimization target with the deviation between the decision sequence output by the model and the actual lance control action, continuously updates the model parameters during training, so that the model gradually reduces the prediction error, accurately learns the mapping law between multimodal working condition characteristics and oxygen lance control actions, and then solidifies the reasonable lance control reasoning logic, and finally obtains a converter oxygen lance control model with higher control accuracy and stronger working condition adaptability.

[0059] The supervised fine-tuning algorithm employs an autoregressive conditional generation framework to learn from multimodal inputs. The mapping relationship to the structured output sequence is determined, and the negative log-likelihood is used as the optimization objective.

[0060] in, C t express t Multimodal input data at any given time. express t Visual modal input data at any given time. express t Text modal input data at any given time. L This represents the total number of the smallest semantic units in the structured output sequence. θ Represents training parameters, N This represents the total number of training samples. n Indicates the index of the training sample. i The index representing the smallest semantic unit of the structured output sequence. P Represents conditional probability. express t The first time in the structured output sequence i The smallest semantic unit, express t The first time in the structured output sequence i The smallest semantic unit before the smallest semantic unit.

[0061] In this embodiment of the invention, a training sample set is constructed based on the third structured decision sequence after error correction. A supervised fine-tuning algorithm is used to optimize the preheating model with negative log-likelihood as the optimization objective. This allows the model to fully learn the mapping relationship between multimodal inputs and structured lance control decisions under the guidance of precise supervision signals. By continuously iterating and optimizing the model parameters, the deviation between the output results and the actual actions is further reduced, and the reasoning path that conforms to the process logic is solidified. This effectively improves the control accuracy and generalization ability of the model, and finally obtains a stable and reliable converter oxygen lance control model that can be directly used in actual production.

[0062] S8: Acquire real-time multimodal converter data.

[0063] S9: When a fourth prompt is provided, the preprocessed real-time multimodal converter data is input into the converter oxygen lance control model, and the converter oxygen lance control decision is output.

[0064] It should be noted that the fourth prompt word and the second prompt word share the same initial prompt word template. The difference lies in the independent prompt content generated by filling in different historical multimodal data. Similarly, it does not contain real gun control actions. Due to the differences in input working condition data, the specific content of the fourth prompt word and the second prompt word are different from each other, adapting to the model inference needs of different samples.

[0065] Reference manual attached Figure 2 The diagram shows a schematic of the structure of a converter oxygen lance control system based on a multimodal large model provided by the present invention.

[0066] The present invention also provides a converter oxygen lance control system 20 based on a multimodal large model, applied to the above-mentioned converter oxygen lance control method based on a multimodal large model, comprising: Processor 201.

[0067] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, they implement the converter oxygen lance control method based on a multimodal large model as described in the method embodiment.

[0068] The converter oxygen lance control system 20 based on a multimodal large model provided by the present invention can execute the above-mentioned converter oxygen lance control method based on a multimodal large model and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.

[0069] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0070] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0071] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0072] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0073] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0074] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0075] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0076] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0077] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0079] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0080] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0081] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the converter oxygen lance control method based on a multimodal large model as described in the method embodiment.

[0082] The present invention provides a computer-readable storage medium that can implement the steps and effects of the converter oxygen lance control method based on a multimodal large model in the above-described method embodiments. To avoid repetition, the present invention will not elaborate further.

[0083] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0084] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.

[0085] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0086] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0087] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A converter oxygen lance control method based on a multi-modal large model, characterized by, include: S1: Collect historical multimodal converter data; S2: Preprocess the historical multimodal converter data to obtain historical multimodal input data; S3: When the first prompt word is provided, the historical multimodal input data is input into the baseline model, and the first structured decision sequence is output; S4: Based on the first structured decision sequence, fine-tune the baseline model to obtain the preheating model; S5: When a second prompt word is provided, the historical multimodal input data is input into the preheating model, a second structured decision sequence is output, and an error assessment is performed on the second structured decision sequence to obtain the error assessment result; S6: Based on the error evaluation result, update the second prompt word to obtain the third prompt word, and input the third prompt word and the historical multimodal input data into the preheating model to output the third structured decision sequence; S7: Based on the third structured decision sequence, the baseline model is fine-tuned to obtain the converter oxygen lance control model; S8: Acquire real-time multimodal converter data; S9: When a fourth prompt is provided, the preprocessed real-time multimodal converter data is input into the converter oxygen lance control model, and the converter oxygen lance control decision is output.

2. The multi-modal large model-based converter oxygen lance control method according to claim 1, characterized in that, The historical multimodal converter data specifically includes: visual modal input data and text modal input data; The visual modal input data specifically includes: a keyframe sequence of a video clip of the furnace flame; The text modal input data specifically includes: historical oxygen lance height sequence, blowing time, process parameter summary, and structured prompt information. 3.The multi-modal large model based converter oxygen lance control method according to claim 1, wherein, The baseline model specifically includes a visual encoding module and a language generation module.

4. The multi-modal large model-based converter oxygen lance control method according to claim 3, characterized in that, The visual encoding module is specifically used for: performing unified feature encoding on image sequences; The language generation module is specifically used to: fuse visual features and structured text prompts based on an autoregressive language model structure to obtain the structured decision sequence.

5. The multi-modal large model-based converter oxygen lance control method according to claim 1, wherein, S6 specifically includes: S601: Based on the error assessment result, the second prompt word is updated using different strategies to obtain the third prompt word; S602: Input the historical multimodal input data and the third prompt word into the preheating model, and output the decision result; S603: Analyze the decision results; S604: Repeat steps S601 to S603 until parsing is successful; S605: Based on the parsing result, determine whether the verification was successful; if yes, save the decision result as the third structured sequence; otherwise, determine whether the maximum number of verifications has been reached; if yes, inject the real control action into the third prompt word, and input it together with the historical multimodal input data into the preheating model to obtain the third structured sequence; otherwise, return to step S601.

6. The multi-modal large model-based converter oxygen lance control method according to claim 5, characterized in that, The error assessment results specifically include: a first error assessment result, a second error assessment result, and a third error assessment result; When the error evaluation result is the first error evaluation result, the second structured decision sequence is saved as the third structured sequence; When the error assessment result is the second error assessment result, the second prompt word is updated through a local correction strategy to prompt the preheating model to adjust the manipulation intensity of the second structured decision sequence; When the error assessment result is the third error assessment result, the second prompt word is updated through the rethinking strategy to indicate that the preheating model made an error in the previous judgment.

7. The converter oxygen lance control method based on a multimodal large model according to claim 1, characterized in that, Specifically, S7 includes: S701: Construct a training sample set based on the third structured decision sequence; S702: Input the training sample set into the baseline model, and obtain the converter oxygen lance control model through a supervised fine-tuning algorithm.

8. A converter oxygen lance control system based on a multimodal large model, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the converter oxygen lance control method based on a multimodal large model as described in any one of claims 1 to 7.

9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the converter oxygen lance control method based on a multimodal large model as described in any one of claims 1 to 7.