A blast furnace opening oxygen lance intelligent regulation and control system based on combustion state in blast furnace
Patent Information
- Application Number
- CN202610769144.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-31
- Publication Date
- 2026-08-21
AI Technical Summary
该独立评估机制在面对高炉开炉特有的退化场景时暴露出严重缺陷:当炉渣喷溅等瞬态事件同时污染两路传感器时,双路置信度同时跌至谷底导致融合层陷入信息真空,系统丧失对燃烧状态的基本感知能力;当仅单路受独立干扰源影响时,系统无法利用健康路的状态交叉验证退化成因,丧失了对退化路施加强化惩罚以确保健康路主导融合决策的修正机会
[0007] Compared with existing technologies, this application provides an intelligent oxygen lance control system based on the combustion state within a blast furnace. This system constructs a spatiotemporal synchronization and deep feature processing link for multimodal acoustic and optical sensing data. It introduces a modal anti-interference confidence assessment mechanism based on harsh operating conditions to dynamically quantify the reliability of each modal signal. Confidence weights drive the adaptive fusion of cross-modal spatiotemporal features to obtain a robust representation of the combustion dynamics state in the swirl zone. This representation is then input into a deep reinforcement learning policy network for policy reasoning in the continuous action space, ultimately parsing it into physically executable multidimensional oxygen lance control commands. This approach significantly improves the reliability and control continuity of combustion state sensing under unsteady blast furnace conditions, achieving robust real-time sensing of the combustion dynamics state in the blast furnace swirl zone and multidimensional coordinated optimal adjustment of oxygen lance flow rate, airflow distribution, and insertion depth.
Smart Images

Figure CN122609773A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control, and more specifically, to an intelligent control system for oxygen lances based on the combustion state inside a blast furnace. Background Technology
[0002] Blast furnace start-up is the most technically challenging and safety-risk non-steady-state operation stage in the ironmaking process. The combustion state of coke and oxygen-enriched hot blast in the vortex zone directly determines the rate of heat accumulation in the hearth and the initial gas flow distribution pattern. Precise control of parameters such as oxygen lance flow rate, gas flow distribution, and insertion depth is crucial for establishing stable initial furnace conditions. Traditional oxygen lance operation during start-up relies heavily on the operator's experience-based judgment and manual intervention. Faced with the rapidly changing combustion field in the vortex zone, manual response exhibits significant time lag and subjectivity, making it difficult to achieve real-time and accurate perception of combustion dynamics and optimal coordinated adjustment of oxygen lance parameters. Therefore, there is an urgent engineering need to develop an intelligent oxygen lance control scheme based on the combustion state within the blast furnace.
[0003] In existing technologies, some solutions attempt to introduce single-modal visual or acoustic sensors to monitor the combustion state in the blast furnace swirl zone online, and combine this with rule-based control or simple models to assist in adjusting oxygen lance parameters. However, the operating environment during blast furnace start-up is extremely harsh, with frequent interference factors such as slag splashing, dust obstruction, and mechanical vibration. Single-modal sensing solutions are prone to signal degradation or even complete failure under extreme conditions. Although some improved solutions adopt a multimodal fusion strategy of acoustic and optical sensors, they still use an independent scoring paradigm for each modality in the modal quality assessment stage. The visual channel scores based solely on the information entropy of its own feature map, and the acoustic channel scores based solely on the signal-to-noise ratio of its own spectrum. There is no information interaction between the two assessment paths. This independent evaluation mechanism reveals serious flaws when faced with the degradation scenarios unique to blast furnace start-up: when transient events such as slag splashing simultaneously contaminate both sensors, the confidence levels of both channels plummet to their lowest points, causing the fusion layer to fall into an information vacuum, and the system loses its basic ability to perceive the combustion state; when only one channel is affected by an independent interference source, the system cannot use the state of the healthy channel to cross-verify the cause of degradation, thus losing the opportunity to apply enhanced penalties to the degraded channel to ensure that the healthy channel dominates the fusion decision.
[0004] Therefore, there is a need for an optimized intelligent control system for oxygen lances based on the combustion state inside the blast furnace. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides an intelligent control system for oxygen lances based on the combustion state within the blast furnace.
[0006] According to one aspect of this application, a smart control system for oxygen lances based on the combustion state inside a blast furnace is provided, comprising: The multimodal data synchronization and characterization module is used to perform spatiotemporal synchronization and feature extraction on the acquired acoustic and optical multimodal sensing data to obtain the visual spatiotemporal feature tensor of the vent and the acoustic spectrum feature matrix of the swirling zone. The acoustic and optical multimodal sensing data includes high frame rate image sequences of vent combustion and broadband acoustic oscillation signals of the direct-blowing pipe. The modal immunity confidence assessment module is used to perform modal immunity confidence assessment based on harsh working conditions on the visual spatiotemporal feature tensor of the wind vent and the acoustic spectrum feature matrix of the swirling zone to obtain the visual modal confidence weight and the acoustic modal confidence weight. The acoustic-optical cross-modal spatiotemporal feature adaptive fusion module is used to perform acoustic-optical cross-modal spatiotemporal feature adaptive fusion of the visual modal confidence weight and the acoustic modal confidence weight on the air vent visual spatiotemporal feature tensor and the acoustic spectrum feature matrix of the swirl zone to obtain the combustion dynamics state characterization vector of the swirl zone. The strategy reasoning module is used to input the combustion dynamics state representation vector of the swirl zone as the current environmental observation into the pre-trained deep reinforcement learning policy network to perform strategy reasoning in order to obtain the continuous control action vector of the oxygen lance. The physical control command parsing module is used to parse the oxygen lance continuous control action vector to obtain the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value. The multi-dimensional coordinated adjustment module is used to send the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value to the programmable logic controller to drive the oxygen lance to perform multi-dimensional coordinated adjustment during furnace start-up.
[0007] Compared with existing technologies, this application provides an intelligent oxygen lance control system based on the combustion state within a blast furnace. This system constructs a spatiotemporal synchronization and deep feature processing link for multimodal acoustic and optical sensing data. It introduces a modal anti-interference confidence assessment mechanism based on harsh operating conditions to dynamically quantify the reliability of each modal signal. Confidence weights drive the adaptive fusion of cross-modal spatiotemporal features to obtain a robust representation of the combustion dynamics state in the swirl zone. This representation is then input into a deep reinforcement learning policy network for policy reasoning in the continuous action space, ultimately parsing it into physically executable multidimensional oxygen lance control commands. This approach significantly improves the reliability and control continuity of combustion state sensing under unsteady blast furnace conditions, achieving robust real-time sensing of the combustion dynamics state in the blast furnace swirl zone and multidimensional coordinated optimal adjustment of oxygen lance flow rate, airflow distribution, and insertion depth. Attached Figure Description
[0008] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 This is a block diagram of an intelligent control system for oxygen lances based on the combustion state inside a blast furnace, according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow of an intelligent oxygen lance control system based on the combustion state inside a blast furnace, according to an embodiment of this application. Figure 3 This is a block diagram of the modal anti-interference confidence assessment module in the intelligent control system for oxygen lance opening based on combustion state in blast furnace, according to an embodiment of this application. Figure 4 This is a block diagram of the strategy reasoning module in the intelligent control system for oxygen lance based on the combustion state inside the blast furnace, according to an embodiment of this application. Detailed Implementation
[0010] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0011] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0012] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.
[0013] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0014] The technical solution of this application proposes an intelligent control system for oxygen lances based on the combustion state inside the blast furnace. Figure 1 This is a block diagram of an intelligent control system for oxygen lances based on the combustion state inside a blast furnace, according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow of an intelligent oxygen lance control system based on the combustion state inside a blast furnace, according to an embodiment of this application. Figure 1 and Figure 2 As shown, the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace according to this application includes: a multimodal data synchronization and characterization module 310, used to perform spatiotemporal synchronization and feature extraction on the collected acoustic-optical multimodal sensing data to obtain the tuyere visual spatiotemporal feature tensor and the cyclone acoustic spectrum feature matrix, wherein the acoustic-optical multimodal sensing data includes a high frame rate image sequence of tuyere combustion and a broadband acoustic oscillation signal of the direct-blowing pipe; a modal anti-interference confidence assessment module 320, used to perform modal anti-interference confidence assessment based on harsh working conditions on the tuyere visual spatiotemporal feature tensor and the cyclone acoustic spectrum feature matrix to obtain the visual modal confidence weight and the acoustic modal confidence weight; and an acoustic-optical cross-modal spatiotemporal feature adaptive fusion module 330, used to perform adaptive fusion based on the visual modal confidence weight and the acoustic modal confidence weight. The system employs an adaptive fusion of the visual spatiotemporal feature tensor of the tuyeres and the acoustic spectral feature matrix of the swirl zone to obtain the combustion dynamics state representation vector of the swirl zone. A strategy reasoning module 340 uses this swirl zone combustion dynamics state representation vector as input to a pre-trained deep reinforcement learning strategy network for strategy reasoning to obtain the continuous control action vector of the oxygen lance. A physical control command parsing module 350 parses the continuous control action vector of the oxygen lance to obtain the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value. A multi-dimensional collaborative adjustment module 360 sends the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value to a programmable logic controller to drive the oxygen lance to perform multi-dimensional collaborative adjustment.
[0015] Specifically, the multimodal data synchronization and characterization module 310 is used to perform spatiotemporal synchronization and feature extraction on the acquired acoustic-optical multimodal sensing data to obtain the tuyere visual spatiotemporal feature tensor and the cyclone zone acoustic spectral feature matrix. The acoustic-optical multimodal sensing data includes a high-frame-rate image sequence of tuyere combustion and a broadband acoustic oscillation signal from the direct-blown pipe. It should be understood that during the blast furnace start-up stage, the combustion process of coke and oxygen-enriched hot air in the cyclone zone exhibits highly unsteady characteristics, and the spatial morphology and acoustic oscillation characteristics of the combustion field continuously and drastically change on a millisecond-level timescale. The high-frame-rate image sequence of tuyere combustion and the broadband acoustic oscillation signal from the direct-blown pipe are acquired by different types of sensors with their own independent sampling clocks, and the two signals naturally have asynchronous deviations in time reference. If the acoustic-optical multimodal sensing data is not strictly spatiotemporally synchronized, a time misalignment will occur between the visual frames and acoustic frames in the subsequent fusion process, resulting in the destruction of cross-modal feature correlation and an inability to accurately reflect the true combustion dynamics state of the cyclone zone at the same physical moment. Furthermore, both the original image sequence and the acoustic oscillation signal are high-dimensional redundant data. Directly inputting them into subsequent modules would bring a huge computational burden and make it difficult to effectively extract the essential features of the combustion state. Therefore, in the technical solution of this application, deep feature extraction is performed on the aligned acoustic-optical raw data streams to compress the high-dimensional raw data into a compact feature representation with physical semantics, providing high-quality input for subsequent modal confidence assessment and cross-modal fusion.
[0016] In practice, firstly, based on the microsecond-level global clock stamp of the blast furnace distributed control system, the video frame timestamps of the high frame rate image sequence of tuyere combustion and the sampling clock stamps of the broadband acoustic oscillation signal of the direct-blown pipe are resampled and interpolated to align to obtain the aligned acoustic-optical raw data stream. Here, it can be understood that the blast furnace distributed control system maintains a microsecond-level global clock synchronization system covering all sensor nodes in the plant. Both the tuyere endoscope high-speed camera and the direct-blown pipe acoustic sensor obtain timestamps from this global clock when acquiring data. However, the video frame rate of the high-speed camera and the sampling frequency of the acoustic sensor are usually inconsistent. For example, the visual channel acquires image sequences at a fixed frame rate of hundreds of frames per second, while the acoustic channel continuously acquires oscillation signals at a sampling rate much higher than the visual frame rate. To eliminate the time reference deviation and sampling rate difference between the two signals, the visual frame timestamp sequence is used as a unified time reference for resampling and interpolation processing of the acoustic oscillation signal. Specifically, for each visual frame timestamp, the two nearest neighbor sampling points are located in the continuous sampling sequence of the acoustic signal, and the instantaneous value of the acoustic signal corresponding to that visual frame moment is calculated using a linear interpolation method. Specifically, the relative position of the visual frame moment on the acoustic sampling time axis is used as the interpolation weight. The two nearest neighbor acoustic sampling values are weighted and summed according to the inverse ratio of their time distance. That is, the proportion of the time interval between the current visual frame moment and the previous acoustic sampling moment to the total interval between the two acoustic sampling moments is used as the weight of the subsequent sampling value, and the remaining proportion is used as the weight of the previous sampling value. The weighted sum of the two yields the interpolated acoustic signal result corresponding to that visual frame moment. By performing the above interpolation operation on each visual frame timestamp, an acoustic signal sequence strictly aligned frame-by-frame with the visual frame sequence on the time axis is finally obtained. The two data streams together constitute the aligned acoustic-optical raw data stream. This process ensures that each pair of visual-acoustic data frames in the subsequent feature extraction and fusion operations accurately corresponds to the cyclotron combustion state at the same physical moment.
[0017] Next, a 3D convolutional neural network is used to extract joint features of spatial morphology and temporal evolution from the image components in the aligned acoustic-optical raw data stream to obtain the spatiotemporal feature tensor of the wind vent. Specifically, the image components in the aligned acoustic-optical raw data stream are a continuous sequence of video frames, and their data form is a multi-frame two-dimensional image arranged along the time axis. Each frame records the spatial morphology, brightness distribution, and color features of the burning flame in the wind vent swirling area at that moment. The convolutional kernel of the 3D convolutional neural network slides simultaneously in the spatial dimension (height and width of the image) and the temporal dimension (frame sequence direction), enabling it to capture both the spatial morphological features (such as flame outline, brightness gradient, and regional texture) and temporal evolution features (such as flame pulsation frequency, morphological change rate, and brightness fluctuation trend) of the burning flame in a single convolutional operation. By stacking multiple 3D convolutional layers and using pooling operations to gradually expand the receptive field, the network can gradually abstract from the local texture and short-term motion features at the lower level to the global combustion morphology and long-term evolution trend features at the higher level. The final output of the wind vent visual spatiotemporal feature tensor encodes the joint feature information of the combustion flame in the swirling zone in two dimensions: spatial morphology and temporal evolution, in the form of a compact multidimensional array. Its dimensions include batch size, number of feature channels, compressed time steps, and compressed spatial height and width.
[0018] Furthermore, time-frequency domain power spectral density mapping is performed on the acoustic components in the aligned acoustic-optical raw data stream to obtain the acoustic spectral feature matrix of the cyclotron region. Here, it should be understood that the acoustic components in the aligned acoustic-optical raw data stream are a one-dimensional time-domain oscillatory signal sequence aligned frame-by-frame with the visual frame sequence, recording the broadband acoustic oscillations generated by the airflow and combustion process within the direct-blowing pipe. Although the time-domain signal contains rich combustion state information, the dynamic changes in its frequency components over time are difficult to reveal through simple time-domain analysis. Therefore, it is transformed to the time-frequency joint domain to obtain the energy distribution characteristics of each frequency component evolving over time. This application uses short-time Fourier transform to perform time-frequency domain power spectral density mapping on the acoustic components. Specifically, a sliding Hanning window function is applied to the aligned acoustic signal for frame-by-frame processing. A discrete Fourier transform is performed on each frame of the windowed signal to obtain the spectral information within that time period, and then the power spectral density of each frequency component is calculated. In this process, the continuous acoustic signal is first divided into multiple overlapping short-time frames according to a fixed window length and frame shift step size. Each frame's signal is multiplied by a Hanning window function to reduce spectral leakage. Then, a Discrete Fourier Transform is performed on the windowed signal of each frame, decomposing the time-domain signal into complex amplitude representations of different frequency components. Finally, the square of the modulus of the complex amplitude of each frequency component is normalized by dividing by the window length, yielding the power spectral density value at each frequency point of that frame. The power spectral density values of all time frames are arranged along the time and frequency axes to obtain the final cyclotron acoustic spectral characteristic matrix, with the number of rows equal to the total number of frames and the number of columns equal to the number of frequency resolution points. Each row of this matrix corresponds to the energy distribution of each frequency component within a time frame, and each column corresponds to the energy evolution trajectory of a specific frequency component over time. Thus, the power spectral density distribution characteristics of cyclotron combustion acoustic oscillations in the time-frequency joint domain are completely encoded in two-dimensional matrix form.
[0019] Specifically, the modal anti-interference confidence assessment module 320 is used to perform modal anti-interference confidence assessment based on harsh operating conditions on the tuyeres visual spatiotemporal feature tensor and the cyclone acoustic spectrum feature matrix to obtain visual modal confidence weights and acoustic modal confidence weights. It should be understood that during the blast furnace start-up stage, the operating environment in the cyclone zone is extremely harsh, with frequent and unpredictable interference factors such as slag splashing, dust obstruction, and mechanical vibration. Although the tuyeres visual spatiotemporal feature tensor and the cyclone acoustic spectrum feature matrix encode the spatial morphological temporal evolution characteristics of the combustion flame and the time-frequency energy distribution characteristics of acoustic oscillations, respectively, the reliability of the two modal signals is not constant under extreme operating conditions. When slag splashing obstructs the camera lens, a large amount of invalid information will be mixed into the visual feature tensor; when mechanical vibration introduces broadband noise, the combustion feature signals in the acoustic spectrum feature matrix will be submerged by noise. If the two modal features are fused indiscriminately with equal weights in the subsequent cross-modal fusion stage, noise and invalid information from the degraded modes will severely contaminate the fusion result, leading to distortion of the combustion dynamics representation vector in the swirl zone and consequently causing erroneous decisions in the oxygen lance control strategy. Therefore, in the technical solution of this application, the current quality state of each modal signal is evaluated in real time before fusion, and the reliability of each mode is dynamically quantified. This provides a confidence weight basis for subsequent adaptive fusion, ensuring that the fusion process can automatically suppress the contribution of degraded modes and strengthen the dominant position of healthy modes.
[0020] Figure 3 This is a block diagram of the modal anti-interference confidence assessment module in the intelligent control system for oxygen lance based on combustion state in a blast furnace, according to an embodiment of this application. Figure 3 As shown, in the first embodiment of this application, the modal anti-interference confidence assessment module 320 includes: a signal quality sensing unit 321, used to perform signal quality sensing on the visual spatiotemporal feature tensor of the wind vent and the acoustic spectrum feature matrix of the swirling zone to obtain a visual feature quality evaluation score and an acoustic signal-to-noise ratio evaluation score; a modal confidence nonlinear mapping unit 322, used to perform modal confidence nonlinear mapping on the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score respectively based on a preset modal sensitivity adjustment coefficient to obtain a visual mapping confidence score and an acoustic mapping confidence score; and a competitive normalization unit 323, used to perform competitive normalization on the visual mapping confidence score and the acoustic mapping confidence score to obtain a visual modal confidence weight and an acoustic modal confidence weight.
[0021] Specifically, the signal quality sensing unit 321 is used to perform signal quality sensing on the visual spatiotemporal feature tensor of the vent and the acoustic spectral feature matrix of the swirling zone to obtain a visual feature quality evaluation score and an acoustic signal-to-noise ratio evaluation score. Specifically, for the visual channel, the signal quality sensing unit quantifies the information richness of the visual spatiotemporal feature tensor of the vent based on information entropy theory. Specifically, it calculates the local spatial information entropy along the spatial dimension of the visual spatiotemporal feature tensor of the vent, quantifying the richness and uniformity of effective information distribution in the current feature map. When the visual signal is in a healthy state, the feature map contains rich information on flame morphology, brightness gradient, and texture details, and the feature value distribution of each pixel region exhibits high diversity, corresponding to a high information entropy value. When the visual signal degrades due to slag obstruction or dust coverage, the feature values of large areas in the feature map tend to be uniform or saturated, effective information is lost, and the corresponding information entropy value decreases significantly. In this process, the spatial feature map is obtained by global average pooling of the visual spatiotemporal feature tensor along the channel dimension. The pixel values of this spatial feature map are normalized to form a probability distribution, and then the Shannon information entropy of this probability distribution is calculated as the visual feature quality evaluation score. The higher the visual feature quality evaluation score, the more uniform the feature value distribution and the richer the information in each spatial location of the feature map; the lower the visual feature quality evaluation score, the more uniform the feature map becomes due to degradation or information loss.
[0022] For the acoustic channel, the signal quality sensing unit quantifies the signal purity of the acoustic spectrum feature matrix in the blast furnace vortex zone based on the frequency domain signal-to-noise ratio (SNR). Specifically, the acoustic spectrum feature matrix in the blast furnace vortex zone is divided into the main frequency combustion energy band and the background environmental noise band along the frequency dimension. The ratio of signal energy in the main frequency band to noise energy in the noise band is calculated as the acoustic signal SNR evaluation score. Under normal combustion conditions in the blast furnace vortex zone, there is a characteristic main frequency energy concentration band in the acoustic spectrum generated by coke combustion and airflow pulsation, whose energy is significantly higher than the background noise. When the acoustic signal degrades due to mechanical vibration or external interference, the noise energy increases over a wide frequency range, and the energy contrast between the main frequency band and the noise band decreases. In this process, the visual spatiotemporal feature tensor of the tuyeres is globally averaged and pooled along the channel dimension to obtain a spatial feature map. The pixel values of this spatial feature map are normalized to form a probability distribution, and then the Shannon information entropy of this probability distribution is calculated as the visual feature quality evaluation score. This process can be expressed by the formula: in, As a visual feature quality evaluation score, For the spatial feature map, the first Normalized eigenvalues of each spatial location, This represents the spatial resolution of the feature map (i.e., the total number of spatial locations). This represents a logarithmic operation with base 2. The more uniform the distribution of feature values at each spatial location in the feature map and the richer the information, the higher the visual feature quality score; conversely, the lower the visual feature quality score, the more uniform the feature map becomes due to degradation or information loss.
[0023] Specifically, the modal confidence nonlinear mapping unit 322 is used to perform modal confidence nonlinear mapping on the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score respectively, based on a preset modal sensitivity adjustment coefficient, to obtain visual mapping confidence scores and acoustic mapping confidence scores. Here, it should be understood that although the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score quantify the quality status of the two modal signals respectively, their original numerical ranges and dimensions are different, and the relationship between the quality evaluation score and the actual reliability is not a simple linear one—when the signal quality degrades from high to low, the rate of decrease in reliability should exhibit a nonlinear accelerating characteristic, that is, slight degradation has a small impact on confidence, while severe degradation should lead to a sharp drop in confidence. Therefore, in the technical solution of this application, the original quality evaluation score is transformed into a confidence score with a unified dimension and conforming to the aforementioned nonlinear attenuation characteristics through a nonlinear mapping function.
[0024] Specifically, a Sigmoid-type nonlinear mapping function with a modal sensitivity adjustment coefficient as a parameter is used to transform the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score, respectively. For the visual channel, the calculation process of the visual mapping confidence score can be expressed by the formula: in, The confidence score for visual mapping ranges from 0 to 1. As a visual feature quality evaluation score, This is the visual modality sensitivity adjustment coefficient (a preset hyperparameter that controls the steepness of the mapping curve). Visual quality threshold (a preset hyperparameter representing the center point of the mapping curve, i.e., the quality assessment score corresponding to a confidence score of 0.5). When The larger the value, the steeper the mapping curve near the threshold, meaning that small changes in the quality assessment score near the threshold will cause drastic changes in the confidence score, and the system is more sensitive to visual degradation; when The smaller the value, the flatter the mapping curve, and the milder the system's response to visual degradation.
[0025] For the acoustic channel, the calculation process of the acoustic mapping confidence score can be expressed by the formula: in, The acoustic mapping confidence score ranges from 0 to 1. The signal-to-noise ratio (SNR) score is used to evaluate acoustic signals. This is the acoustic modal sensitivity adjustment coefficient (a preset hyperparameter that controls the steepness of the mapping curve). The acoustic quality threshold (a preset hyperparameter that represents the center point of the mapping curve, i.e., the signal-to-noise ratio evaluation score corresponding to a confidence score of 0.5).
[0026] It is worth noting that the visual modal sensitivity adjustment coefficient and the acoustic modal sensitivity adjustment coefficient can be set independently according to the actual operating conditions of the blast furnace, so as to adapt to the differentiated response requirements of the two modal signals under different degradation modes.
[0027] Specifically, the competitive normalization unit 323 is used to competitively normalize the visual mapping confidence score and the acoustic mapping confidence score to obtain the visual modality confidence weight and the acoustic modality confidence weight. Considering that although the visual mapping confidence score and the acoustic mapping confidence score have been mapped to a unified range of 0 to 1, their sum is not necessarily equal to 1, and therefore cannot be directly used as weighting coefficients in subsequent fusion stages, the technical solution of this application places the confidence scores of the two modalities under the same competitive framework through competitive normalization. This ensures that a relative increase in the confidence score of one modality is accompanied by a relative decrease in the weight of the other modality, thereby achieving dynamic competition and resource reallocation between modalities. Specifically, the Softmax normalization function is used to perform competitive normalization processing on the visual mapping confidence score and the acoustic mapping confidence score.
[0028] The calculation process for the visual modality confidence weight can be expressed by the following formula: in, For visual modality confidence weights, For visual mapping confidence scores; The calculation process of acoustic modal confidence weights can be expressed by the following formula: in, For acoustic modal confidence weights, For acoustic mapping confidence scores.
[0029] In particular, after competitive normalization by Softmax, The principle holds true consistently, with both weights strictly greater than 0 and less than 1. When the confidence score of the visual mapping is significantly higher than that of the acoustic mapping, the confidence weight of the visual modality approaches 1 while that of the acoustic modality approaches 0, indicating that the visual modality is more reliable at the current moment, and subsequent fusion should prioritize visual features; conversely, the opposite is also true. When the confidence scores of the two mappings are similar, the two weights approach 0.5 each, indicating that the reliability of the two modalities is comparable, and subsequent fusion should give them approximately equal attention. This competitive normalization mechanism ensures that the sum of the fusion weights is always 1 under any condition, avoiding the problems of weight overflow or insufficiency. At the same time, the amplification effect of the exponential function enhances the competitive contrast between the two modalities, enabling two modalities with small quality differences to produce discriminative weight allocations.
[0030] However, research has found that in the real physical scenario of blast furnace start-up, the combustion aerodynamics of the swirl zone determines that the visual degradation event and the acoustic degradation event are not statistically independent random processes, but rather there is a temporal causal coupling relationship driven by the same physical source.
[0031] Specifically, when a large-scale slag splash occurs in the swirling zone, the high-temperature splashes simultaneously block the lens of the endoscopic camera at the tuyeres in a very short time (causing a large-area brightness change and information loss in the visual feature map) and impact the wall of the direct-blowing pipe (generating broadband impact noise superimposed on the normal combustion acoustic signal). At this time, the degradation of the two sensors exhibits highly synchronized characteristics on the time axis.
[0032] However, another type of degradation scenario exhibits completely different temporal characteristics—when only slow lens dust accumulation (progressive visual degradation) occurs without airflow disturbance, the acoustic signal remains clean and unaffected; conversely, when mechanical vibrations such as the start and stop of an external hydraulic pump introduce acoustic noise, the visual signal is also unaffected. These two types of degradation scenarios are completely asynchronous in time.
[0033] In the first embodiment, the calculation of the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score adopts completely independent evaluation paths—the visual channel is scored only based on the information entropy of its own feature map, and the acoustic channel is scored only based on the signal-to-noise ratio of its own spectrum. There is no information interaction between the two evaluation paths. This independent evaluation paradigm exposes serious adaptive defects when facing the two types of degradation scenarios mentioned above: when slag splashing causes both signals to degrade simultaneously due to the same physical event, the independent evaluation mechanism marks both as extremely low quality. The subsequent fusion layer falls into an information vacuum because the confidence of both channels drops to the bottom at the same time, losing the basic perception ability of the combustion state in the vortex zone; when only one channel degrades, the independent evaluation mechanism cannot use the health status of the other channel to cross-verify the cause of degradation, losing the opportunity for diagnostic correction to impose a stronger penalty on the degraded channel to strengthen the dominance of the healthy channel. Ultimately, the essential limitation of the first embodiment lies in treating degradation only as noise that needs to be quantified, while ignoring the physical causal information carried by the degradation mode itself.
[0034] In view of the above-mentioned technical defects, this application further proposes a second embodiment.
[0035] Specifically, firstly, by calculating the local spatial entropy and frequency signal-to-noise ratio frame by frame along the time axis, the degradation trend sequence of the visual spatiotemporal feature tensor of the wind vent and the acoustic spectral feature matrix of the swirl zone is extracted to obtain the visual degradation time-series curve and the acoustic degradation time-series curve, respectively. Before evaluating the quality of the two modal signals, it is necessary to obtain the dynamic trajectory of the degradation degree over time, rather than relying solely on a static snapshot at a single time point. In this process, the local spatial entropy of the visual spatiotemporal feature tensor of the wind vent is calculated frame by frame along the time axis, quantifying the degree of effective information loss caused by dust or slag obstruction in each frame. The entropy values of each frame are arranged in time sequence to form a visual degradation time-series curve reflecting the fluctuation of visual quality over time.
[0036] Simultaneously, the amplitude ratio of the main frequency combustion energy band to the background environmental noise band is calculated frame by frame along the time axis of the acoustic spectrum feature matrix of the cyclotron region. The degree of interference of the acoustic signal in each frame is quantified, and the signal-to-noise ratio values of each frame are arranged in time sequence to form an acoustic degradation time sequence curve that reflects the acoustic quality fluctuation over time.
[0037] In the blast furnace start-up scenario, the visual degradation time-series curve can capture the sudden drop in brightness caused by slag splashing and the gradual information attenuation caused by lens ash accumulation. The acoustic degradation time-series curve can capture the pulse-like signal-to-noise ratio drop caused by pipe wall impact and the periodic noise superposition caused by mechanical vibration. The construction of these two time-series curves transforms the originally isolated single-frame quality assessment into a dynamic degradation trajectory description with temporal context, laying the data foundation for subsequent cross-modal degradation correlation analysis.
[0038] Next, within a sliding window, the visual and acoustic degradation time-series curves are quantized for temporal synchronicity, and the instantaneous value of the current frame is extracted to obtain the cross-modal degradation synchronicity coefficient and the original visual and acoustic quality scores. In other words, after obtaining two degradation time-series curves, their co-variation trend in the temporal dimension is quantified to determine whether the current degradation event is driven by the same physical source. During this process, within a preset sliding time window, the Pearson correlation coefficient is calculated for the visual and acoustic degradation time-series curves to obtain the cross-modal degradation synchronicity coefficient. This process can be expressed by the formula: in, This is the cross-modal degradation synchronization coefficient, with a value range of [value range missing]. ; The value of the local spatial information entropy of the visual degradation time-series curve in the t-th frame within the sliding window; Let be the frequency domain signal-to-noise ratio value of the acoustic degradation time-series curve in the t-th frame within the sliding window; This represents the arithmetic mean of the information entropy of each frame in the visual degradation timeline curve within the sliding window; This is the arithmetic mean of the signal-to-noise ratio of each frame of the acoustic degradation time-series curve within the sliding window; The length of the number of frames contained in the sliding time window.
[0039] It should be noted that when large-scale slag splashing occurs in the swirl zone, the splashes almost simultaneously obscure the lens and impact the tube wall. The visual degradation timeline and the acoustic degradation timeline show a highly positively correlated synchronous decline trend within the window. A value approaching positive 1 indicates homogeneous degradation—meaning the degradation of both sensors is caused simultaneously by the same combustion anomaly. However, when only slow lens ash accumulation occurs (visual degradation alone) or only mechanical vibration introduced by the hydraulic pump's start / stop occurs (acoustic degradation alone), one curve declines while the other remains stable; neither exhibits a co-variation trend within the window. Approaching 0 indicates heterogeneous degradation—that is, the two degradation paths are caused by their own independent interference sources. Simultaneously, the instantaneous value of the visual degradation time-series curve at the current time step is extracted as the raw visual quality score, and the instantaneous value of the acoustic degradation time-series curve at the current time step is extracted as the raw acoustic quality score, for use in subsequent correction steps.
[0040] Furthermore, based on the cross-modal degradation synchronicity coefficient, homogeneous degradation and heterogeneous degradation scenarios are distinguished. Causal perception adaptive correction is applied to the original visual quality score and the original acoustic quality score to obtain the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score. Here, it should be understood that for homogeneous degradation scenarios (transient events such as slag splashing simultaneously contaminate two sensors), the sensors themselves do not suffer permanent failure, and the signal will recover on its own after the transient interference subsides. Therefore, it is necessary to moderately increase the quality scores of the two paths to avoid the fusion layer falling into an information vacuum due to the confidence of both paths dropping to the bottom at the same time. For heterogeneous degradation scenarios (only one path is affected by an independent interference source), the signal quality of the degraded path is indeed unreliable, and it is necessary to apply additional penalty attenuation to it, while keeping the quality score of the healthy path unchanged, so as to strengthen the dominant position of the healthy path in subsequent fusion decisions.
[0041] The correction process for the visual channel can be expressed by the following formula: in, The output visual feature quality score is the result of correction for degradation causes. The original visual quality score is the uncorrected value obtained directly from the information entropy. This is the cross-modal degradation synchronization coefficient; The homogeneity degradation compensation intensity coefficient serves as the upper limit of the quality score recovery rate under the preset hyperparameter control in homogeneous scenarios. This is a penalty term for visual heterogeneous degradation. It is negative when visual degradation alone (acoustic health) is detected to apply additional attenuation, and zero when visual health is detected.
[0042] The correction process for the acoustic channel can be expressed by the following formula: in, The signal-to-noise ratio score of the output acoustic signal after correction for degradation causes; The acoustic raw quality score is the uncorrected value obtained directly from the frequency domain signal-to-noise ratio. This is the cross-modal degradation synchronization coefficient; The homogeneity degradation compensation intensity coefficient shares the same hyperparameter as the visual correction formula; This is a penalty term for acoustic heterogeneous degradation. It is negative when acoustic degradation alone is detected (visual health) to apply additional attenuation, and zero when acoustic health is good.
[0043] It should be noted that when When it approaches 1, in the formula The term has a positive amplification effect on the original quality score, achieving compensation and recovery in the same degradation scenario, while When the term approaches zero, the penalty term is automatically deactivated; when When it approaches 0, The term degenerates into an identity mapping close to 1, while the original mass score remains essentially unchanged, and When the term approaches 1, the penalty term is fully activated, applying a strong decay to the degenerate path.
[0044] In the blast furnace start-up scenario, this means that when a large-scale slag splash simultaneously causes the lens to be blocked and the acoustic sensor to be impacted, the correction mechanism identifies the physical cause of homogeneous transient degradation, appropriately restores the confidence of both paths, and ensures that the control system can still maintain the basic perception capability of the combustion state in the swirl zone under extreme transient interference. When the start-up and shutdown of the external hydraulic pump only introduces acoustic noise, the correction mechanism identifies heterogeneous independent degradation, imposes a strong penalty on the acoustic channel while maintaining the dominance of the visual channel, and ensures that the fusion layer can correctly rely on undisturbed visual information to judge the combustion state.
[0045] By introducing a cross-modal degradation synchronicity coefficient as a quantitative diagnostic basis for degradation causes, and accordingly making differential corrections to the original quality evaluation scores of the two modes based on causal perception, the original independent evaluation mechanism overcomes the two types of core failure modes when facing extreme blast furnace start-up conditions.
[0046] Specifically, in homogeneous degradation scenarios, the modified quality assessment score avoids the information vacuum problem in the fusion layer caused by the simultaneous drop in confidence levels of both paths to the bottom. This allows subsequent multimodal fusion steps to still obtain effective combustion state information input during transient extreme events such as slag splashing, thus maintaining the continuity and robustness of the oxygen lance control system. In heterogeneous degradation scenarios, the modified quality assessment score can accurately identify the degradation path and impose a stronger penalty on it, while preserving the information dominance of the healthy path. This allows the subsequent fusion layer to correctly shift the decision focus to the undisturbed modal channel, avoiding incorrect combustion state judgments due to single-path noise contamination. Overall, the improved modal quality assessment mechanism transforms the degradation mode itself into a diagnostic signal carrying physical causal information, achieving a paradigm shift from passively quantifying the degree of degradation to actively inferring the causes of degradation and responding differently. This significantly improves the perception reliability and control continuity of the blast furnace start-up oxygen lance intelligent control system under non-steady-state harsh conditions.
[0047] Specifically, the acoustic-optical cross-modal spatiotemporal feature adaptive fusion module 330 is used to adaptively fuse the visual spatiotemporal feature tensor of the tuyeres and the acoustic spectral feature matrix of the swirl zone based on the confidence weights of the visual and acoustic modes to obtain a swirl zone combustion dynamics state representation vector. It should be understood that during the blast furnace start-up stage, the combustion dynamics state of the swirl zone is a complex system state driven by the coupling of multiple physical processes, and the feature representation of a single mode can only provide a partial characterization from a certain perceptual dimension. The visual spatiotemporal feature tensor of the tuyeres encodes the spatial morphology and temporal evolution characteristics of the combustion flame, but cannot perceive the acoustic oscillation characteristics of airflow pulsation and combustion intensity; the acoustic spectral feature matrix of the swirl zone encodes the time-frequency energy distribution characteristics of the combustion acoustic oscillations, but cannot perceive the spatial structure and morphological changes of the flame. By deeply fusing the feature information of the two modes, the system can obtain a comprehensive and robust representation of the combustion dynamics state of the swirl zone. However, simple feature concatenation or fixed-weighting cannot capture the deep semantic connections and complementary relationships between the two modal features, and cannot dynamically adjust the fusion strategy according to the real-time reliability of each modality under harsh conditions. Therefore, in the technical solution of this application, the two heterogeneous modal features are deeply integrated while considering their respective reliability, and finally output a compact low-dimensional vector to comprehensively represent the combustion dynamics state of the current vortex zone, providing high-quality environmental observation input for subsequent deep reinforcement learning policy networks.
[0048] In practice, firstly, cross-modal correlation interaction and attention feature extraction are performed on the visual spatiotemporal feature tensor of the vent and the acoustic spectral feature matrix of the swirl zone to obtain a cross-modal interaction feature tensor. Specifically, firstly, the two heterogeneous modal features are unified to the same dimensional space for interactive computation. Specifically, the visual spatiotemporal feature tensor of the vent is flattened, merging its spatial and temporal dimensions into a sequence dimension to obtain a visual feature sequence matrix, where each row corresponds to a feature vector of a spatiotemporal location; similarly, the acoustic spectral feature matrix of the swirl zone undergoes a dimensional transformation, using its time frame dimension as the sequence dimension, with each row corresponding to a spectral feature vector of a time frame. If the channel dimensions of the two features are inconsistent, they are mapped to the same hidden dimension space through a linear projection layer.
[0049] After dimensionality unification, a cross-modal attention mechanism is employed to facilitate the correlation interaction between the two modal features. Specifically, for visual modality-guided cross-attention, a query matrix is generated from the visual feature sequence through linear transformation, and a key matrix and value matrix are generated from the acoustic feature sequence through linear transformation. During execution, for each position in the visual feature sequence, the dot product similarity between the query vector at that position and the key vectors at all positions in the acoustic feature sequence is calculated. All similarity values are divided by a scaling factor and then transformed into a normalized weight distribution using the Softmax function. This weight distribution is then used to perform a weighted summation of the value vectors at all positions in the acoustic feature sequence to obtain the feature representation of that visual position enhanced by acoustic information. Performing the above operation on each position in the visual sequence yields the complete visual-to-acoustic cross-attention output matrix. Symmetrically, for acoustic modality-guided cross-attention, a query matrix is generated from the acoustic feature sequence, and a key matrix and value matrix are generated from the visual feature sequence to obtain the acoustic-to-visual cross-attention output. During execution, for each position in the acoustic feature sequence, the dot product similarity between the query vector at that position and the key vectors at all positions in the visual feature sequence is calculated. After scaling and Softmax normalization, the value vectors at all positions in the visual feature sequence are weighted and summed using a weighted distribution to obtain the feature representation of the acoustic position after visual information guidance enhancement.
[0050] After bidirectional cross-attention computation, the cross-attention output of vision on acoustics (i.e., visual features enhanced by acoustic information) and the cross-attention output of acoustics on vision (i.e., acoustic features enhanced by visual information) are obtained. These two are then concatenated along the feature dimension to form a cross-modal interactive feature tensor. This tensor simultaneously encodes the enhanced features of the visual modality under acoustic semantic guidance and the enhanced features of the acoustic modality under visual semantic guidance, fully capturing the deep semantic connections and complementary information between the two modalities.
[0051] Next, based on the visual modality confidence weight and the acoustic modality confidence weight, the cross-modal interaction feature tensor is adaptively reweighted based on multimodal confidence to obtain a confidence-weighted feature combination. Here, it should be understood that although the cross-modal interaction feature tensor encodes the interaction information of two modalities, the contribution ratio of the visual guidance part and the acoustic guidance part is fixed, without considering the actual reliability differences of the signals of each modality at the current moment. Under adverse conditions, if the visual signal is severely degraded, noise information introduced by the degraded visual features will be mixed into the cross-attention output of the visual to the acoustic signal, and the reliability of this part of the feature should be reduced; conversely, if the acoustic signal is degraded, the noise contribution in the cross-attention output of the acoustic to the visual signal should be suppressed. Therefore, in the technical solution of this application, the feature components from different modal guidance in the cross-modal interaction feature tensor are adaptively reweighted differently using the visual modality confidence weight and the acoustic modality confidence weight.
[0052] Specifically, the cross-modal interaction feature tensor is split along the feature dimension into visual guidance feature components and acoustic guidance feature components. Each element of the visual guidance feature component is multiplied by the visual modality confidence weight and scaled to ensure that the component maintains a large amplitude when the visual modality is reliable and its amplitude is compressed when the visual modality degenerates. Similarly, each element of the acoustic guidance feature component is multiplied by the acoustic modality confidence weight and scaled to ensure that the component maintains a large amplitude when the acoustic modality is reliable and its amplitude is compressed when the acoustic modality degenerates. The scaled two feature components are then concatenated along the feature dimension to obtain a confidence-weighted feature combination. Through adaptive modulation of the confidence weights, the visual guidance feature component dominates the combination when the visual modality is highly reliable; the acoustic guidance feature component dominates when the acoustic modality is highly reliable; and when both modalities are reliable, the two feature components contribute approximately equally. This mechanism ensures that the fusion result is always dominated by the more reliable modality information at the current moment, effectively suppressing the negative impact of degenerate modalities on the fusion quality.
[0053] Furthermore, a multilayer perceptron with a residual structure is used to perform deep nonlinear feature compression on the confidence-weighted feature combination to obtain the combustion dynamics state representation vector of the swirl zone. Specifically, the confidence-weighted feature combination is first flattened into a one-dimensional vector as input to the multilayer perceptron. After layer-by-layer nonlinear transformation and dimensional compression through multiple residual blocks, the final output layer produces a low-dimensional vector with a fixed dimension, namely the combustion dynamics state representation vector of the swirl zone. This vector comprehensively encodes the combustion dynamics state information of the swirl zone at the current moment in a compact numerical form, including comprehensive features such as combustion intensity, flame stability, and airflow distribution uniformity, and can be directly used as the environmental observation input for the deep reinforcement learning policy network.
[0054] Specifically, the strategy reasoning module 340 is used to input the combustion dynamics state representation vector of the swirl zone as the current environmental observation into a pre-trained deep reinforcement learning strategy network for strategy reasoning to obtain the continuous control action vector of the oxygen lance. It should be understood that during the blast furnace start-up phase, control parameters such as the oxygen lance flow rate, gas flow distribution, and insertion depth need to be continuously and precisely dynamically adjusted according to the real-time changes in the combustion dynamics state of the swirl zone, rather than simple discrete gear switching or fixed rule responses. Although the combustion dynamics state representation vector of the swirl zone has fully encoded the combustion state information at the current moment in a compact numerical form, there is a highly complex nonlinear decision mapping relationship between state perception and control action output—the same combustion state may correspond to different optimal control strategies under different historical trajectories and future expectations, and the effect of the control action has time delay and cumulative characteristics. The action selection at the current moment not only affects the immediate combustion state but will also affect the furnace condition evolution in multiple future time steps through the propagation of physical processes. Traditional rule-based or supervised learning-based control methods are difficult to effectively handle this sequential decision-making problem with delayed rewards, long-term cumulative effects, and continuous action space characteristics. Deep reinforcement learning policy networks, through trial and error learning on a large amount of historical furnace start-up data or in simulation environments, can automatically discover the complex mapping relationship between combustion state observations and optimal continuous control actions. They output oxygen lance control strategies that balance immediate effects and long-term impacts while maximizing long-term cumulative rewards. Therefore, in the technical solution of this application, the combustion dynamics state representation vector of the swirl zone is used as the input of the current environmental observation to a pre-trained deep reinforcement learning policy network for policy inference, in order to obtain an oxygen lance continuous control action vector that can accurately specify the adjustment amounts of each control parameter of the oxygen lance in the continuous action space.
[0055] Figure 4 This is a block diagram of the strategy reasoning module in the intelligent control system for oxygen lance operation based on the combustion state inside the blast furnace, according to an embodiment of this application. Figure 4 As shown, the strategy reasoning module 340 includes: a deep state feature encoding unit 341, used to perform agent state space observation and hidden layer feature mapping on the combustion dynamics state representation vector of the swirling zone to obtain the hidden layer state features of reinforcement learning; a Gaussian distribution parameter regression unit 342, used to perform Gaussian distribution parameter regression on the hidden layer state features of reinforcement learning to obtain the mean vector and standard deviation vector of the action distribution; and a reparameterization sampling and legal boundary truncation unit 343, used to perform continuous action reparameterization sampling and legal boundary truncation on the mean vector and standard deviation vector of the action distribution based on random noise variables sampled by standard normal distribution to obtain the continuous control action vector of the oxygen lance.
[0056] Specifically, the deep state feature encoding unit 341 is used to perform agent state space observation and hidden layer feature mapping on the combustion dynamics state representation vector of the swirling zone to obtain the hidden layer state features of reinforcement learning. Specifically, the combustion dynamics state representation vector of the swirling zone is taken as input and sequentially passed through multiple fully connected layers for layer-by-layer nonlinear feature transformation. Each fully connected layer performs a linear transformation, followed by nonlinear mapping through a nonlinear activation function. This process can be expressed by the following formula: in, This is the characterization vector for the combustion dynamics state in the swirl zone. , ,..., They are respectively the first to the second floor. The weight matrix of the fully connected layer; , ,..., These are the bias vectors for the corresponding layers. It is a non-linear activation function; , ,..., These are the output feature vectors of each layer, and the final... This refers to the hidden layer state features in reinforcement learning. The weight matrix and bias vector of each layer are parameters learned by the policy network during the pre-training phase through reinforcement learning algorithms (such as the Proximal Policy Optimization algorithm, PPO) and interaction with the environment, and remain fixed during the inference phase without being updated. Through multi-layer nonlinear transformations, the original combustion state observations are gradually mapped to a higher-level semantic feature space that is more discriminative for control decisions, laying the foundation for accurate regression of subsequent action distribution parameters.
[0057] Specifically, the Gaussian distribution parameter regression unit 342 is used to perform Gaussian distribution parameter regression on the reinforcement learning hidden layer state features in the continuous action space to obtain the action distribution mean vector and action distribution standard deviation vector. Specifically, the Gaussian distribution parameter regression unit consists of two parallel linear output heads, which respectively regress the mean parameter and standard deviation parameter of the action distribution from the reinforcement learning hidden layer state features. The mean output head linearly maps the hidden layer state features to the action distribution mean vector through a fully connected layer, the dimension of which is equal to the dimension of the continuous action space (i.e., the number of oxygen gun control parameters). This process can be expressed by the following formula: in, Let the mean vector of the action distribution be _____. The weight matrix of the mean output head. For the corresponding bias vector, To reinforce the learning of hidden layer state features, each component in the action distribution mean vector corresponds to an optimal action estimate in the dimension of the oxygen lance control parameters.
[0058] The standard deviation output head linearly maps the hidden layer state features to a logarithmic standard deviation vector through another fully connected layer, and then transforms it into a positive action distribution standard deviation vector through an exponential function. The reason for using logarithmic space regression followed by exponentiation is that the standard deviation must be positive. Direct regression to positive values can easily lead to numerical instability during optimization, while regression in logarithmic space followed by exponentiation naturally ensures a positive output and a more stable optimization process. This process can be expressed by the formula: in, Let the standard deviation be a vector of logarithms. The standard deviation is the weight matrix of the output head. For the corresponding bias vector, Let be the standard deviation vector of the action distribution. This is an element-wise exponential function operation. Each component in the action distribution standard deviation vector corresponds to the degree of policy uncertainty in a dimension of oxygen lance control parameters. The smaller the standard deviation, the more certain the policy's action selection in that dimension (i.e., it has fully learned the optimal action), while the larger the standard deviation, the more exploration space the policy still has in that dimension.
[0059] Specifically, the reparameterized sampling and legal boundary truncation unit 343 is used to perform continuous action reparameterized sampling and legal boundary truncation on the action distribution mean vector and action distribution standard deviation vector based on random noise variables sampled from a standard normal distribution to obtain the oxygen gun continuous control action vector. Specifically, firstly, a random noise variable vector with the same action dimension is sampled from a standard normal distribution (a Gaussian distribution with a mean of 0 and a standard deviation of 1). Then, this noise variable vector is multiplied element-wise by the action distribution standard deviation vector and added to the action distribution mean vector to obtain the action value sampled from the target Gaussian distribution. This process can be expressed by the following formula: in, Let be a vector of random noise variables sampled from a standard normal distribution. This represents a multidimensional standard normal distribution with a mean of zero and a covariance matrix equal to the identity matrix. The original action vector obtained by reparameterized sampling. Let the mean vector of the action distribution be _____. Let be the standard deviation vector of the action distribution. This represents a positional dot product. Through this reparameterization transformation, the original action vector obtained by sampling is mathematically equivalent to the vector obtained from the mean value. Standard deviation is It is directly sampled from a Gaussian distribution, but the deterministic path in its computation graph is differentiable with respect to the network parameters.
[0060] However, the original motion vector obtained by reparameterized sampling theoretically ranges from negative infinity to positive infinity, while actual oxygen lance control parameters have clear physical legal boundaries (e.g., oxygen flow rate cannot be negative, and insertion depth cannot exceed physical limits). Therefore, the original motion vector is further truncated to restrict its components within a preset legal motion range. Specifically, the legal boundary truncation uses a truncation function (i.e., the Clip operation) to apply upper and lower limits to each component of the original motion vector. This process can be expressed by the following formula: in, The first continuous control motion vector of the oxygen lance Each dimension component (i.e., the truncated legal action value). The first parameter of the original action vector obtained by reparameterized sampling Each dimension component For the first Preset legal lower bound for each action dimension For the first A preset legal upper bound for each action dimension. The function's purpose is to output the lower bound when the input value is less than the lower bound, output the upper bound when the input value is greater than the upper bound, and keep the original value unchanged when the input value is between the upper and lower bounds. After performing legal boundary truncation on each of the action dimensions, the resulting components together constitute the oxygen gun's continuous control action vector. Each component in this vector is within the legal range of its corresponding physical control parameters and can be safely used for subsequent physical control command parsing.
[0061] Specifically, the physical control command parsing module 350 is used to parse the oxygen lance continuous control action vector to obtain the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value. It should be understood that the oxygen lance continuous control action vector is a normalized dimensionless numerical vector, with the values of its components limited to a standardized range of -1 to +1. This is a common numerical representation convention used by deep reinforcement learning policy networks to facilitate training and optimization. However, the control commands received by the actual blast furnace oxygen lance physical actuators (including the oxygen flow regulating valve, multi-ring airflow distributor, and oxygen lance lifting drive device) must be actual values with clear physical dimensions and engineering units. For example, the total oxygen flow rate is expressed in standard cubic meters per hour, the multi-ring airflow distribution ratio in percentage, and the oxygen lance insertion depth in millimeters. Furthermore, the values of each parameter must be strictly limited within the engineering safety range of the corresponding actuator to ensure equipment safety and process compliance. Therefore, in the technical solution of this application, the normalized action vector output by the strategy network is transformed into a control value with actual physical meaning through engineering quantity inverse normalization mapping, establishing a precise correspondence from the abstract decision space to the physical execution space, so that the decision output of the intelligent control system can be directly received by the programmable logic controller and drive the physical execution mechanism to perform the corresponding adjustment action.
[0062] In practical implementation, based on the preset upper and lower limits of the engineering safety range for each physical actuator, the components of each dimension in the oxygen lance continuous control action vector are subjected to engineering quantity inverse normalization mapping to obtain the total oxygen flow control value, the multi-ring gas flow distribution ratio value, and the oxygen lance insertion depth control value. Specifically, the oxygen lance continuous control action vector contains three dimensions, corresponding to the three physical control parameters: total oxygen flow, multi-ring gas flow distribution ratio, and oxygen lance insertion depth. Each physical control parameter has an upper and lower limit of the engineering safety range determined jointly by the equipment hardware characteristics and process safety specifications. These upper and lower limits are configured as preset parameters of the system according to the actual specifications of the specific blast furnace equipment during deployment. Through engineering quantity inverse normalization mapping, the standardized interval from negative 1 to positive 1 in the normalized action space is linearly mapped to the actual engineering range interval corresponding to each physical parameter, so that the normalized value of negative 1 precisely corresponds to the lower limit of the engineering safety range, the normalized value of positive 1 precisely corresponds to the upper limit of the engineering safety range, and the normalized value of 0 corresponds to the midpoint value of the range interval, with the intermediate values corresponding in a linear proportional relationship.
[0063] Specifically, for the first dimension component of the continuous control action vector of the oxygen lance (corresponding to the total oxygen flow control parameter), the normalized action value of this dimension is incremented by 1 and divided by 2, linearly transforming it from the range of negative 1 to positive 1 to the range of 0 to 1. Then, it is multiplied by the difference between the upper and lower limits of the engineering safety range of the total oxygen flow (i.e., the range span), and the lower limit of the engineering safety range is added to obtain the total oxygen flow control value with actual physical dimensions. This process can be expressed by the formula: in, This is the total oxygen flow control value. This is the first dimension component of the continuous control motion vector for the oxygen lance. This is the preset lower limit of the engineering safety range for the total oxygen flow rate. This is the upper limit of the preset engineering safety range for the total oxygen flow rate.
[0064] For the second dimension component of the oxygen lance continuous control action vector (corresponding to the multi-ring airflow distribution ratio control parameter), it is exactly the same as the first dimension. The normalized action value of this dimension is incremented by 1 and divided by 2, linearly transforming it from the range of negative 1 to positive 1 to the range of 0 to 1. Then, it is multiplied by the difference between the upper and lower limits of the engineering safety range for the multi-ring airflow distribution ratio, and the lower limit of the engineering safety range is added to obtain the multi-ring airflow distribution ratio value with actual physical dimensions. This process can be expressed by the formula: in, This represents the multi-ring airflow distribution ratio. This is the second dimension component of the oxygen lance continuous control motion vector. This is the lower limit of the preset engineering safety range for the multi-ring airflow distribution ratio. The upper limit of the preset engineering safety range for the multi-ring airflow distribution ratio.
[0065] For the third dimension component of the continuous control motion vector of the oxygen lance (corresponding to the oxygen lance insertion depth control parameter), the same linear transformation logic is followed. The normalized motion value of this dimension is incremented by 1 and divided by 2, linearly transforming it from the range of negative 1 to positive 1 to the range of 0 to 1. Then, it is multiplied by the difference between the upper and lower limits of the engineering safety range for the oxygen lance insertion depth, and the lower limit of the engineering safety range is added to obtain the oxygen lance insertion depth control value with actual physical dimensions. This process can be expressed by the formula: in, This is the control value for the oxygen lance insertion depth. This is the third dimension component of the oxygen lance continuous control motion vector. This is the preset lower limit of the engineering safety range for the oxygen lance insertion depth. The preset engineering safety range upper limit for the oxygen lance insertion depth.
[0066] The above three dimensions of engineering quantity inverse normalization mapping can be uniformly represented by a general linear mapping formula, for any _th_ element in the oxygen lance continuous control action vector. This applies to all dimensional components. The process can be expressed by the formula: in, For the first Engineering quantity control values for each physical control parameter The first continuous control motion vector of the oxygen lance Each dimension component For the first The preset lower limit of the engineering safety range for each physical control parameter. For the first The preset engineering safety range upper limit of each physical control parameter. Since the components of each dimension in the oxygen lance continuous control action vector have been strictly limited to the range of negative 1 to positive 1 in the aforementioned reparameterized sampling and legal boundary truncation unit, the values of each physical control parameter obtained after the above linear mapping will necessarily fall strictly between the corresponding engineering safety range upper and lower limits, without the need for additional safety limiting processing, thus fundamentally ensuring the engineering safety of the output control command.
[0067] Specifically, the multi-dimensional coordinated adjustment module 360 is used to send the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value to the programmable logic controller to drive the oxygen lance for multi-dimensional coordinated adjustment during blast furnace start-up. It should be understood that in the intelligent control system for blast furnace start-up oxygen lances, the physical control command parsing module has already converted the normalized action vector output by the strategy network into total oxygen flow control values, multi-ring airflow distribution ratio values, and oxygen lance insertion depth control values with actual physical dimensions. However, these three control values are still at the digital signal level of the host computer and have not yet been converted into electrical control signals that can drive the actual actions of the physical actuators. During blast furnace start-up, the three control dimensions of oxygen lance flow regulation, airflow distribution regulation, and insertion depth regulation are not independent and isolated operations, but rather have a close physical coupling relationship—changes in the total oxygen flow will affect the pressure distribution of each ring pipeline, thus changing the actual airflow distribution ratio; changes in the oxygen lance insertion depth will change the outlet back pressure of the oxygen jet, thus affecting the actual flow; and changes in the multi-ring airflow distribution ratio will change the momentum balance of each ring jet, thus affecting the force state of the oxygen lance. If the adjustments of the three control parameters are asynchronous in time or inconsistent in logic, transient instability in the combustion field of the swirl zone may occur, potentially even leading to a safety accident. Therefore, a programmable logic controller (PLC) is used to simultaneously send the setpoints of the three control parameters to each physical actuator in a strictly synchronized manner. This ensures that the oxygen flow regulating valve, multi-ring airflow distributor, and oxygen lance lifting drive device work together to complete their respective adjustment actions within the same control cycle, achieving time synchronization and logical coordination of the oxygen lance's multi-dimensional control parameters.
[0068] In practice, firstly, the host computer (i.e., the industrial server running the intelligent control algorithm) encapsulates the oxygen total flow control value, multi-loop airflow distribution ratio value, and oxygen lance insertion depth control value output by the physical control instruction parsing module into a standard industrial communication protocol data frame. This data frame is then sent to the communication interface module of the programmable logic controller (PLC) via industrial Ethernet or fieldbus (such as Profinet, Modbus TCP, etc.). The data frame contains the setpoints for the three control parameters and their corresponding timestamps. The timestamps ensure that the PLC can recognize that this set of control instructions belongs to the coordinated output of the same decision cycle and should be executed synchronously within the same control cycle. Upon receiving the data frame, the PLC's communication interface module performs protocol parsing and data integrity verification, extracts the three setpoints (oxygen total flow control value, multi-loop airflow distribution ratio value, and oxygen lance insertion depth control value), and writes them into the PLC's internal control parameter register, awaiting read by the execution logic in the next scan cycle.
[0069] Next, in each fixed-cycle program scan, the programmable logic controller (PLC) simultaneously reads three setpoints from the control parameter register: the total oxygen flow control value, the multi-loop airflow distribution ratio value, and the oxygen lance insertion depth control value. Within the same scan cycle, it also synchronously calculates the drive signals for each physical actuator. For the total oxygen flow control channel, the PLC uses the total oxygen flow control value as the target setpoint for the flow regulation loop, compares it with the actual flow measurement value fed back by the oxygen flow sensor, calculates the deviation, and then uses a PID control algorithm to calculate the opening adjustment of the oxygen flow regulating valve. Specifically, first, the deviation between the target setpoint and the actual measurement value is calculated as the control error. Then, this control error is fed into the proportional, integral, and derivative stages for parallel processing. The proportional stage multiplies the current control error by the proportional gain coefficient to obtain the immediate response component. The integral stage integrates the control error from the initial time to the current time and multiplies it by the integral gain coefficient to obtain the component that accumulates and eliminates the steady-state error. The derivative stage multiplies the rate of change of the control error with respect to time by the derivative gain coefficient to obtain the lead correction component. Finally, the output components of the three stages are added together to obtain the opening adjustment of the oxygen flow regulating valve. Similarly, for the multi-loop airflow distribution ratio control channel, the programmable logic controller (PLC) uses the multi-loop airflow distribution ratio value as the target setpoint for the airflow distribution adjustment loop, compares it with the actual distribution ratio measurement value fed back by the flow sensors in each loop, calculates the deviation, and uses a PID control algorithm to calculate the opening adjustment amount of the regulating valve in each loop of the multi-loop airflow distributor. Likewise, for the oxygen lance insertion depth control channel, the PLC uses the oxygen lance insertion depth control value as the target setpoint for the position adjustment loop, compares it with the actual insertion depth measurement value fed back by the oxygen lance displacement sensor, calculates the deviation, and uses a PID control algorithm to calculate the drive signal for the oxygen lance lifting drive device.
[0070] After completing the PID calculations for the three control channels within the same scan cycle, the programmable logic controller (PLC) synchronously outputs the drive signal adjustment values of the three channels to their respective analog-to-digital converters (ADCs). This ensures that the drive signals of the three control channels arrive at each physical actuator in strict synchronization, achieving coordinated adjustment of multi-dimensional control parameters. Specifically, the PLC's ADC module synchronously converts the digital drive signals of the three channels into analog electrical signals, which are then output to the electric actuators of the oxygen flow regulating valve, the electric actuators of the regulating valves in each loop of the multi-loop airflow distributor, and the servo driver of the oxygen lance lifting drive device. The electric actuator of the oxygen flow regulating valve adjusts the valve opening according to the received drive signal, changing the flow cross-sectional area of the oxygen pipeline to adjust the total oxygen flow to the target set value. The electric actuators of the regulating valves in each loop of the multi-loop airflow distributor adjust the valve opening of each loop according to the received drive signal, changing the flow distribution ratio between the loops to the target set value. The servo driver of the oxygen lance lifting drive device controls the rotation direction and speed of the servo motor according to the received drive signal, driving the oxygen lance axially to the target insertion depth position through a mechanical transmission mechanism. The three physical actuators respond synchronously to their respective drive signals within the same control cycle, and work together to adjust the three control parameters of total oxygen flow, multi-ring gas distribution ratio and oxygen lance insertion depth, thereby achieving multi-dimensional coordinated control of the blast furnace oxygen lance during start-up.
[0071] As described above, the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace according to the embodiments of this application can be implemented in various wireless terminals, such as servers with intelligent oxygen lance control algorithms based on the combustion state inside the blast furnace. In one possible implementation, the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace can also be one of many hardware modules of the wireless terminal.
[0072] Alternatively, in another example, the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace and the wireless terminal can also be separate devices, and the intelligent oxygen lance control system 300 based on the combustion state inside the blast furnace can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.
[0073] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A smart oxygen lance control system based on the combustion state inside a blast furnace, characterized in that, include: The multimodal data synchronization and characterization module is used to perform spatiotemporal synchronization and feature extraction on the acquired acoustic and optical multimodal sensing data to obtain the visual spatiotemporal feature tensor of the vent and the acoustic spectrum feature matrix of the swirling zone. The acoustic and optical multimodal sensing data includes high frame rate image sequences of vent combustion and broadband acoustic oscillation signals of the direct-blowing pipe. The modal immunity confidence assessment module is used to perform modal immunity confidence assessment based on harsh working conditions on the visual spatiotemporal feature tensor of the wind vent and the acoustic spectrum feature matrix of the swirling zone to obtain the visual modal confidence weight and the acoustic modal confidence weight. The acoustic-optical cross-modal spatiotemporal feature adaptive fusion module is used to perform acoustic-optical cross-modal spatiotemporal feature adaptive fusion of the visual modal confidence weight and the acoustic modal confidence weight on the air vent visual spatiotemporal feature tensor and the acoustic spectrum feature matrix of the swirl zone to obtain the combustion dynamics state characterization vector of the swirl zone. The strategy reasoning module is used to input the combustion dynamics state representation vector of the swirl zone as the current environmental observation into the pre-trained deep reinforcement learning policy network to perform strategy reasoning in order to obtain the continuous control action vector of the oxygen lance. The physical control command parsing module is used to parse the oxygen lance's continuous control action vector to obtain the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value.
2. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 1, characterized in that, It also includes a multi-dimensional coordinated adjustment module, which sends the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value to the programmable logic controller to drive the oxygen lance to perform multi-dimensional coordinated adjustment during furnace start-up.
3. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 1, characterized in that, The multimodal data synchronization and characterization module is used for: Based on the microsecond-level global clock stamp of the blast furnace distributed control system, the video frame timestamp of the high frame rate image sequence of tuyere combustion and the sampling clock stamp of the broadband acoustic oscillation signal of the direct blowing pipe are resampled, interpolated and aligned to obtain the aligned original acoustic and optical data stream. By using a three-dimensional convolutional neural network, joint feature extraction of spatial morphology and temporal evolution of image components in the aligned acoustic-optical raw data stream is performed to obtain the visual spatiotemporal feature tensor of the wind vent. The acoustic components in the aligned acousto-optic raw data stream are subjected to time-frequency domain power spectral density mapping to obtain the acoustic spectrum feature matrix of the cyclotron region.
4. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 1, characterized in that, The modal immunity confidence assessment module includes: The signal quality sensing unit is used to perform signal quality sensing on the visual spatiotemporal feature tensor of the wind vent and the acoustic spectrum feature matrix of the swirling zone to obtain the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score. The modal confidence nonlinear mapping unit is used to perform modal confidence nonlinear mapping on the visual feature quality evaluation score and the acoustic signal signal-to-noise ratio evaluation score based on a preset modal sensitivity adjustment coefficient to obtain the visual mapping confidence score and the acoustic mapping confidence score. The competitive normalization unit is used to competitively normalize the visual mapping confidence score and the acoustic mapping confidence score to obtain the visual modality confidence weight and the acoustic modality confidence weight.
5. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 1, characterized in that, The acousto-optic cross-modal spatiotemporal feature adaptive fusion module is used for: Cross-modal correlation interaction and attention feature extraction are performed on the visual spatiotemporal feature tensor of the wind vent and the acoustic spectrum feature matrix of the vortex region to obtain the cross-modal interaction feature tensor; Based on the confidence weights of the visual modality and the acoustic modality, the cross-modal interaction feature tensor is subjected to adaptive reweighting based on multimodal confidence to obtain a confidence-weighted feature combination. By using a multilayer perceptron with residual structure, a deep nonlinear feature compression is performed on the confidence-weighted feature combination to obtain the combustion dynamics state characterization vector of the swirl zone.
6. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 1, characterized in that, The strategy reasoning module includes: The deep state feature encoding unit is used to perform agent state space observation and hidden layer feature mapping on the combustion dynamics state representation vector in the swirling zone to obtain the hidden layer state features of reinforcement learning. The Gaussian distribution parameter regression unit is used to perform Gaussian distribution parameter regression on the state features of the hidden layer in reinforcement learning to obtain the mean vector and standard deviation vector of the action distribution. The reparameterization sampling and legal boundary truncation unit is used to perform continuous motion reparameterization sampling and legal boundary truncation on the motion distribution mean vector and motion distribution standard deviation vector based on random noise variables sampled from a standard normal distribution to obtain the oxygen gun continuous control motion vector.
7. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 1, characterized in that, The physical control command parsing module is used for: Based on the preset upper and lower limits of the engineering safety range of each physical actuator, the components of each dimension in the continuous control action vector of the oxygen lance are mapped by engineering quantity inverse normalization to obtain the total oxygen flow control value, the multi-ring airflow distribution ratio value, and the oxygen lance insertion depth control value.
8. The intelligent oxygen lance control system based on the combustion state inside the blast furnace according to claim 4, characterized in that, Signal quality sensing unit, used for: By calculating the local spatial information entropy and frequency domain signal-to-noise ratio frame by frame along the time axis, the degradation trend sequence extraction of the visual spatiotemporal feature tensor of the wind vent and the acoustic spectral feature matrix of the swirling zone is performed to obtain the visual degradation time series curve and the acoustic degradation time series curve. Within a sliding window, temporal synchronicity quantization is performed on the visual degradation time-series curve and the acoustic degradation time-series curve, and the instantaneous value of the current frame is extracted to obtain the cross-modal degradation synchronicity coefficient, the visual raw quality score, and the acoustic raw quality score. Based on the cross-modal degradation synchronicity coefficient, homogeneous degradation and heterogeneous degradation scenarios are distinguished. Causal perception adaptive correction is performed on the original visual quality score and the original acoustic quality score to obtain the visual feature quality evaluation score and the acoustic signal-to-noise ratio evaluation score.