Rotary kiln coal injection quantity prediction method and device based on multi-modal data dynamic gating and readable storage medium thereof

By combining the dynamic gated cross-attention mechanism and the latent clustering deconstruction mechanism, the problems of multimodal data quality fluctuation and time series data non-stationarity are solved, and the accurate prediction and real-time control of the coal injection amount of the rotary kiln are achieved, thereby improving combustion efficiency and production stability.

CN120687915AActive Publication Date: 2025-09-23CHINA JILIANG UNIV

Patent Information

Application Number
CN202511174726.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-23
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing technologies are unable to dynamically adapt to the quality fluctuations of multimodal data and have difficulty handling the non-stationarity of time series data, resulting in difficulty in accurately capturing the sudden change points in operating conditions when controlling the coal injection rate of the rotary kiln, thus affecting combustion efficiency and production stability.

Method used

A dynamic gated cross-attention mechanism is used to realize quality-aware adaptive fusion of multimodal data. The latent clustering deconstruction mechanism is combined to perform intelligent adaptive segmentation and hierarchical feature modeling of time series data. The dynamic gated cross-attention mechanism is used to evaluate modal confidence and perform adaptive weighted fusion. The latent clustering deconstruction mechanism is combined to monitor the displacement changes of the hidden state clustering center of the time series model to identify the working condition change points.

Benefits of technology

It significantly improves the accuracy and robustness of the prediction of coal injection amount in the rotary kiln, can respond to changes in operating conditions in real time, improves combustion efficiency and production stability, and promotes the intelligent and refined management of rotary kiln operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687915A_ABST
    Figure CN120687915A_ABST
Patent Text Reader

Abstract

The invention provides a rotary kiln coal injection quantity prediction method and device based on multi-modal data dynamic gating and a readable storage medium thereof, and aims to solve the problems of unstable quality of multi-modal data and non-stability of time sequence data in an industrial field. Multi-modal feature adaptive fusion is realized by evaluating modal confidence; monitoring hidden state clustering center displacement in combination with a hidden state clustering deconstruction mechanism, and recognizing working condition change points in combination with a dynamic threshold value to realize adaptive segmentation of time series data; and constructing hierarchical features through local and global attention, and finally outputting a predicted value of the coal injection quantity. According to the method, the accuracy, robustness and real-time performance of prediction are improved, and support is provided for industrial intelligent control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent prediction and control of industrial processes, and in particular to a method and device for predicting the coal injection amount of a rotary kiln based on multimodal data dynamic gating, and a readable storage medium thereof. Background Art

[0002] The rotary kiln is the core heat treatment equipment in the cement, metallurgy, chemical and other industrial fields. It mainly realizes the physical and chemical transformation of materials through high-temperature calcination. The coal injection amount is a key process parameter that directly affects the combustion efficiency and production stability.

[0003] Currently, the control of coal injection rate in rotary kilns mainly relies on the operator's experience. This has the following problems: difficulty in grasping the complex correlation between multiple parameters, prone to misjudgment during long-term monitoring, and delayed response to sudden changes in operating conditions. Existing intelligent prediction methods face two major bottlenecks: First, the quality of multimodal data (images, numerical values, etc.) in industrial sites is unstable due to sensor failures, environmental interference, and other factors. Traditional fixed-weight fusion methods cannot dynamically adapt to data reliability. Second, time series data is non-stationary, and traditional fixed sliding windows are difficult to accurately capture sudden changes in operating conditions, resulting in insufficient identification of key state changes.

[0004] Therefore, developing an intelligent prediction system that can dynamically integrate multimodal data and adaptively process time series characteristics is of great significance to improving the operating efficiency of rotary kilns and saving energy and reducing emissions. Summary of the Invention

[0005] The embodiments of the present invention provide a method, device and readable storage medium for predicting the coal injection amount of a rotary kiln based on dynamic gating of multimodal data. They address the problems existing in current technology, such as the inability to dynamically adapt to the quality fluctuations of multimodal data and the difficulty in effectively processing the non-stationarity of time series data to accurately capture the sudden changes in operating conditions.

[0006] The core technology of this invention is to propose a dynamic gated cross-attention mechanism to realize the quality-aware adaptive fusion of multimodal data, combine it with the latent clustering deconstruction mechanism to realize the intelligent adaptive segmentation and hierarchical feature modeling of time series data, and construct an end-to-end prediction device for the coal injection amount of a rotary kiln.

[0007] In a first aspect, the present invention provides a method for predicting the amount of coal injected into a rotary kiln based on dynamic gating of multimodal data, the method comprising the following steps: Collecting multimodal data during the rotary kiln process, the multimodal data including at least combustion image data and process numerical data; the process numerical data including at least combustion state parameters and historical coal injection amount data, the combustion state parameters including parameters corresponding to under-combustion, normal combustion and over-combustion states; Extract spatial features from combustion image data and perform high-dimensional feature encoding on process numerical data; A dynamic gated cross-attention mechanism is used to fuse spatial features and encoded high-dimensional features. The dynamic gated cross-attention mechanism generates modal confidence by evaluating the quality of different modal data and adaptively weights the different modal features based on the modal confidence. The hidden clustering deconstruction mechanism is used to perform time series processing on the fused feature sequence. By monitoring the displacement changes of the hidden state clustering center of the time series model and combining it with the dynamic threshold algorithm to identify the working condition change points, the adaptive segmentation of the time series data is achieved. A local attention mechanism is used to extract fine-grained features within adaptive segments, while a global attention mechanism is used to integrate information across segments to model long-term dependencies. The global time series features processed by hierarchical attention are nonlinearly transformed to output the predicted value of the rotary kiln coal injection amount.

[0008] Furthermore, extracting spatial features from the combustion image data includes: Extract the brightness features, clarity features, smoke concentration features and inter-frame consistency features of the combustion image; High-dimensional feature encoding of process numerical data includes: Extract process parameter change rate characteristics, short-term volatility characteristics and combustion state coding characteristics of process values.

[0009] Furthermore, in the dynamic gated cross-attention mechanism, the modal confidence is generated by nonlinearly mapping the extracted features through a multi-layer perceptron, and the confidence weights of different modalities are balanced through regularization constraints.

[0010] Furthermore, in the latent clustering deconstruction mechanism, the latent state of the time series model is extracted through the long short-term memory network, and the identification of the operating condition change point is based on whether the displacement between the current latent state and the historical cluster center exceeds the dynamic threshold. The dynamic threshold is the sum of the mean of the historical displacement and the preset sensitivity coefficient multiplied by the standard deviation of the historical displacement.

[0011] In a second aspect, the present invention provides a rotary kiln coal injection amount prediction device based on multimodal data dynamic gating, comprising: A data acquisition module is used to collect multimodal data during the rotary kiln process, wherein the multimodal data includes at least combustion image data and process numerical data; the process numerical data includes at least combustion state parameters and historical coal injection amount data, and the combustion state parameters include parameters corresponding to under-combustion, normal combustion and over-combustion states; Feature extraction module, used to extract spatial features from combustion image data and perform high-dimensional feature encoding on process numerical data; The multimodal fusion module uses a dynamic gated cross-attention mechanism to fuse spatial features and high-dimensional features. The dynamic gated cross-attention mechanism generates modal confidence by evaluating the quality of different modal data and performs adaptive weighted fusion of different modal features based on the modal confidence. The time series processing module uses a hidden clustering deconstruction mechanism to perform time series processing on the fused feature sequence. The hidden clustering deconstruction mechanism monitors the displacement changes of the hidden state clustering center of the time series model and combines it with a dynamic threshold algorithm to identify the working condition change points, thereby achieving adaptive segmentation of the time series data. Hierarchical attention module, including local attention mechanism and global attention mechanism. The local attention mechanism is used to capture local mutation features and small state transitions within the adaptive segment, while the global attention mechanism is used to model long-term dependencies and global change trends through cross-segment information interaction; The prediction module performs nonlinear transformation on the global temporal features after hierarchical attention processing and outputs the predicted value of the rotary kiln coal injection amount.

[0012] Furthermore, the feature extraction module includes: An image feature extraction unit is used to extract the brightness features, clarity features, smoke concentration features, and inter-frame consistency features of combustion image data through a deep learning network; The numerical feature extraction unit is used to extract the process parameter change rate characteristics, short-term volatility characteristics and combustion state coding characteristics of process numerical data through a deep learning network.

[0013] Furthermore, the multimodal fusion module includes: a confidence evaluation unit, configured to generate an image modal confidence metric based on the features output by the image feature extraction unit, and to generate a numerical modal confidence metric based on the features output by the numerical feature extraction unit; The dynamic weighting unit is used to construct a confidence weighting matrix based on the image modal confidence and the numerical modal confidence. The confidence weighting matrix is ​​used to perform weighted adjustment on the key features and value features in the feature fusion process to achieve adaptive enhancement of high-quality modal features.

[0014] Furthermore, the confidence assessment unit performs nonlinear mapping on image features and numerical features through a multi-layer perceptron to generate image modal confidence and numerical modal confidence; the multimodal fusion module also includes a regularization unit, which is used to balance the image modal confidence and numerical modal confidence through regularization constraints to avoid excessive dependence of the model on a single modality.

[0015] Furthermore, the timing processing module includes: Hidden state extraction unit, used to process the fused feature sequence through the long short-term memory network to generate a temporal hidden state; Cluster center calculation unit, used to calculate the cluster center of the time series hidden state and the displacement between the current hidden state and the historical cluster center in real time; A change point identification unit is used to compare the displacement with the dynamic threshold and identify the working condition change point to complete the adaptive segmentation of the time series data. The dynamic threshold is determined based on the statistical characteristics of the historical displacement; The dynamic threshold is the sum of the mean of the historical displacements and the preset sensitivity coefficient multiplied by the standard deviation of the historical displacements.

[0016] In a third aspect, the present invention provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process, and the process includes the above-mentioned method for predicting the coal injection amount of a rotary kiln based on dynamic gating of multimodal data.

[0017] The main contributions and innovations of the present invention are as follows: 1. Effectively address the issue of unstable multimodal data quality: A dynamic gated cross-attention mechanism is used to evaluate modal data quality and generate confidence in real time. The weights of different modal features are dynamically adjusted to automatically enhance the influence of high-quality modalities and suppress interference from low-quality data. This prevents the model from over-reliance on a single modality, balances the complementary value of image morphological features and process numerical values, and significantly improves the robustness and adaptability of multimodal feature fusion.

[0018] 2. Breaking through the limitations of traditional time series modeling: Through the latent cluster deconstruction mechanism, the displacement of the latent cluster center is monitored and combined with dynamic thresholds to accurately identify the points of change in working conditions and achieve adaptive segmentation of time series data. This overcomes the defect of traditional fixed sliding windows in capturing mutation points. The hierarchical architecture that combines local and global attention can not only capture local mutation characteristics but also model long-term dependencies, effectively solving the prediction lag problem of non-stationary time series data and improving the ability to identify and predict changes in key working conditions.

[0019] 3. Overall improvement of prediction performance: Through the synergistic effect of the above mechanisms, the accuracy, robustness and real-time performance of the prediction of the coal injection amount of the rotary kiln are significantly improved, providing reliable support for parameter optimization and intelligent decision-making in the industrial production process, and effectively promoting the intelligent and refined management level of the rotary kiln operation.

[0020] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below so that other features, objects, and advantages of the invention are more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flow chart of a method for predicting coal injection amount of a rotary kiln based on multimodal data dynamic gating according to an embodiment of the present invention; Figure 2 is an overall structural diagram according to an embodiment of the present invention; Figure 3 is a structural diagram of a dynamic gated cross-attention mechanism according to an embodiment of the present invention; Figure 4 is a scatter plot of coal injection amount prediction according to an embodiment of the present invention; Figure 5 This is a graph showing the effect of time series prediction of coal injection amount according to an embodiment of the present invention; Figure 6 FIG. 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.

[0023] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0024] As core equipment in the fields of cement, metallurgy, etc., the prediction of coal injection volume of rotary kiln is affected by multimodal data of combustion images (flame state) and process values ​​(temperature, flow, etc.), and needs to cope with dynamic working conditions in complex industrial environments.

[0025] Based on this, the present invention realizes quality-aware adaptive fusion of multimodal data based on a dynamic gated cross-attention mechanism to solve the problems existing in the prior art.

[0026] Example 1 The present invention aims to propose a method for predicting the amount of coal injection in a rotary kiln based on dynamic gating of multimodal data. Figure 1 , the method comprises the following steps: Step 1: Collect multimodal data during the rotary kiln process, the multimodal data including at least combustion image data and process numerical data; the process numerical data including at least combustion state parameters and historical coal injection amount data, the combustion state parameters including parameters corresponding to under-combustion, normal combustion and over-combustion states; In this embodiment, during the dataset construction phase, combustion images, combustion state data, and coal injection amount data during the rotary kiln process are acquired to construct a multimodal dataset.

[0027] Step 2: Extract spatial features from combustion image data and perform high-dimensional feature encoding on process numerical data; In this embodiment, if Figure 2 As shown in the figure, in the feature extraction and representation learning stage, a deep convolutional neural network (CNN) is used to extract spatial features of the combustion image to capture the flame shape, brightness distribution and boundary features; at the same time, a multi-layer perceptron (MLP) network is used to perform nonlinear mapping and representation learning on the combustion state parameters and coal injection amount values ​​to achieve high-dimensional semantic encoding of numerical features. For example: (1) For the image data, the flame brightness, clarity, dust concentration, and inter-frame consistency attributes are extracted as follows: 1. Flame brightness:

[0028] Among them, H and W represent the height and width of the burning image, that is, the image contains pixels; is the three-channel (red, green, and blue channel) value of the pixel point (i, j), and the result is linearly normalized to the [0, 1] interval.

[0029] Used to evaluate the quality of image modalities and provide basic features for confidence calculation in the dynamic gated cross-attention mechanism.

[0030] 2. Clarity: The Laplace operator is used to calculate the mean value of the image gradient and quantify the sharpness of the flame edge. The formula is:

[0031] Among them, H and W represent the height and width of the burning image, that is, the image contains pixels; f(i,j) represents the grayscale value (or color channel value) of the image at pixel (i,j); It is the calculation result of the Laplace operator on the pixel. Its essence is the second-order derivative of the image, which reflects the curvature change of the pixel value in the local area. The absolute value of the second-order derivative is larger in the edge area due to the sudden change of pixel value. The result is normalized to [0,1]. The larger the value, the clearer the image.

[0032] The calculated result C is normalized to a value range of [0, 1]. A larger value indicates a sharper flame edge and a clearer outline in the image. This feature is used for subsequent quality assessment of image modalities and provides a key basis for the calculation of image confidence in the dynamic gated cross-attention mechanism.

[0033] 3. Dust concentration: The formula for counting the proportion of gray pixels in the HSV color space is:

[0034] Among them, H and W represent the height and width of the burning image, that is, the image contains pixels; Is a discriminant function used to determine whether a pixel (i, j) is a gray pixel: when the parameters of the pixel in the HSV color space meet the "hue H is 0~180, saturation S<30, lightness V<100", The value is 1 (gray pixel), otherwise it is 0 (non-gray pixel); H, S, and V are the hue, saturation, and brightness values ​​in the HSV space, respectively; the value range of S is [0,1]. The larger the value, the higher the proportion of gray pixels in the image, which corresponds to a higher smoke concentration in the combustion image (because smoke often appears as a gray area in the image). This indicator is used to evaluate the quality of image data - too high a smoke concentration will cause the flame features to be blurred, thereby reducing the reliability of the image modality, providing a key feature basis for the subsequent dynamic gated cross-attention mechanism to calculate the image confidence.

[0035] This is equivalent to counting the number of all gray pixels in the image (obtained by summing them up) and dividing it by the total number of pixels. , and get the gray pixel proportion SC.

[0036] 4. Inter-frame consistency: The structural similarity index (SSIM) is used to calculate the similarity between adjacent frames. The formula is:

[0037] Among them, x and y represent two adjacent frames of images (such as the current frame and the previous frame); are the mean of the pixel values ​​of the two frames, reflecting the overall brightness level; are the variances of the pixel values ​​of the two frames, reflecting the contrast characteristics; is the covariance of the pixel values ​​of the two frames, reflecting the structural correlation between pixels; This is a constant (usually set based on the image dynamic range) used to avoid zero denominators and ensure computational stability. SSIM values ​​range from [0, 1]. Values ​​closer to 1 indicate greater similarity between adjacent frames and better inter-frame consistency (e.g., stable flames with no severe jitter or noise). Lower values ​​indicate greater inter-frame discrepancies (possibly due to image instability caused by high-temperature dust or sensor jitter).

[0038] This indicator, as a key feature of the image modality, is used to evaluate the stability and quality of image data, and provide a basis for the subsequent dynamic gated cross-attention mechanism to calculate the image modality confidence. The higher the inter-frame consistency, the stronger the image reliability and the higher the confidence; otherwise, it decreases, prompting the model to dynamically enhance its attention to high-quality numerical modalities and achieve multimodal adaptive fusion.

[0039] The multi-dimensional image features extracted above are input into the deep convolutional neural network (CNN), and the spatial features are integrated and mapped to high dimensions through operations such as convolution and pooling, and the fused image feature vector is finally output. , which contains the comprehensive semantic information of the image modality.

[0040] Preferably, the method further includes nonlinearly mapping the image features into image modality confidence through a two-layer multi-layer perceptron (MLP):

[0041] in, The modal confidence factor of the output image is used to quantify the "healthiness" or reliability of the combustion image data. Its value range is usually [0,1] (the closer the value is to 1, the higher the data quality); It is an activation function (such as the sigmoid function), which limits the network output to the interval [0,1], which conforms to the quantitative characteristics of confidence. is the input image feature set, including the key features extracted above, such as flame brightness, clarity, smoke concentration, and inter-frame consistency; are the weight matrices of the two-layer MLP, is the corresponding bias term, which is used to capture the association between features through linear transformation.

[0042] (2) For the numerical data, the process parameter change rate, short-term volatility, and combustion state coding values ​​are extracted as follows: 1. Process parameter change rate: Reflects the adjustment range of process parameters between units:

[0043] in, Indicates the process parameter values ​​at the current moment (such as time t) (such as coal injection amount, temperature, pressure and other key process indicators); Indicates the same process parameter value at the previous moment (such as moment t-1); calculation result Characterizes the relative adjustment range of process parameters per unit time. The larger the value, the more drastic the parameter change (such as a sudden and significant increase / decrease in coal injection amount).

[0044] 2. Process parameter volatility: The stability of the coal injection system is quantified by the standard deviation of short-term data, and the formula is as follows:

[0045] in, It represents the i-th process parameter value (such as coal injection rate, secondary air temperature and other key numerical parameters) within a short-term time window (such as from t-n+1 to time t); n is the window size (usually the number of data points within 10 minutes), representing the short-term data range used to evaluate stability; It is the mean of all process parameter values ​​within the window, reflecting the overall level in the short term. Intuitively reflects the severity of fluctuations in short-term process parameters: The larger the value is, the more significant the fluctuation of the parameter value is in the short term, and the worse the stability of the coal injection system is. The smaller it is, the smoother the parameter changes and the smoother the system runs.

[0046] 3. Combustion status code value: The combustion state is converted into a One-Hot vector, the specific form is as follows: When the combustion state is "under-combustion", the corresponding vector is [1,0,0]; When the combustion state is "normal combustion", the corresponding vector is [0,1,0]; When the combustion state is "overburning", the corresponding vector is [0,0,1].

[0047] Among them, the characteristics of the one-hot vector are: only one position in each vector is "1", and the rest are "0". The "1" position uniquely corresponds to a combustion state. This encoding method avoids the misleading "value size implies priority" that may be introduced when directly representing categories with numbers (such as 1, 2, 3) (for example, it will not cause the model to mistakenly believe that "overburning" (3) is more "important" than "underburning" (1)). The categories are distinguished only by position differences.

[0048] The converted one-hot vector serves as a key feature of the numerical mode and contributes to the confidence assessment of the numerical mode. For example, if the combustion state frequently switches between "underburn" and "overburn" (corresponding to rapid switching of the vector), this may indicate sensor anomalies or unstable operating conditions, reducing the reliability of the numerical mode. The dynamic gated cross-attention mechanism uses this feature to adjust the weight of the numerical mode, ensuring that the model responds appropriately to changes in the combustion state and improving the accuracy of the coal injection rate prediction.

[0049] The multi-dimensional numerical features extracted above are input into the multi-layer perceptron (MLP), and high-dimensional semantic mapping and fusion are performed through nonlinear transformations (such as activation functions and fully connected layers), and finally the numerical feature vector is output. , which contains the comprehensive dynamic information of the numerical mode.

[0050] Preferably, the numerical modal features are nonlinearly mapped to the numerical modal confidence by a two-layer multi-layer perceptron (MLP):

[0051] in, The output numerical modal confidence level is used to quantify the reliability of process numerical data (such as coal injection rate, temperature and other parameters). The value range is usually [0,1] (the closer the value is to 1, the higher the data quality); It is an activation function (such as the sigmoid function), which limits the network output to the interval [0,1], which conforms to the quantitative characteristics of confidence. is the input numerical feature set, including key features extracted above, such as process parameter change rate, short-term volatility, combustion state coding, etc. are the weight matrices of the two-layer MLP, is the corresponding bias term, which is used to capture the association between features through linear transformation.

[0052] Finally, for the image feature vector ( is the image feature dimension) and the numerical feature vector ( is a numerical feature dimension), calculate its Query, Key, and Value:

[0053] in, : The weight matrix corresponding to the image modality is used to map image features to the "query", "key" and "value" spaces respectively; : The weight matrix corresponding to the numerical mode is used to map the numerical features to the "query", "key" and "value" spaces respectively; Among them, the Query, Key, and Value of the image mode are generated: Image feature vector It is a high-dimensional representation of the image modality after feature extraction (such as the fusion of flame brightness and clarity features extracted by CNN). In order to enable it to participate in cross-modal attention calculation, it needs to be linearly transformed through three learnable weight matrices: :By image features through query weight matrix The mapping is obtained and used for “active query” to determine the degree of correlation with the numerical modal features.

[0054] : Image features through key weight matrix The mapping is obtained and used to match the query with the numerical mode and calculate the association weight; :Image features are passed through the value weight matrix The mapping is the core feature information in the image modality used for final fusion; Query, Key, and Value generation for numerical mode: Numerical eigenvector It is a high-dimensional representation of the numerical mode after feature encoding (such as the fusion of process parameter change rate and volatility features processed by MLP). Similarly, linear transformation is performed through three dedicated learnable weight matrices: Query ): The numerical features are passed through the query weight matrix The mapping is obtained and used to determine the degree of correlation between “active query” and image modality features; Key ): by numerical features through key weight matrix The mapping is obtained and used to match the query of the image modality and calculate the association weight; Value ): The numerical features are passed through the value weight matrix The mapped information is the core feature information used for final fusion in the numerical mode.

[0055] This set of formulas is used to map the raw feature vectors of the image and numerical modalities into the "query," "key," and "value" vectors required by the dynamic gated crisscross attention mechanism. In short, these mappings are the "preprocessing" step of the attention mechanism. By converting features from different modalities into vectors suitable for calculating associated weights, they lay the foundation for the dynamic gated crisscross attention mechanism to achieve "quality-aware adaptive fusion."

[0056] Step 3: A dynamic gated cross-attention mechanism is used to fuse the spatial features and the encoded high-dimensional features. The dynamic gated cross-attention mechanism generates modal confidence by evaluating the quality of different modal data and adaptively weights the different modal features based on the modal confidence. In this embodiment, during the multimodal dynamic gated fusion stage, the quality and reliability of different modal data are evaluated in real time through a dynamic gated cross-attention mechanism, modal confidence scores are calculated, and adaptive weighted fusion of different modal features is performed accordingly, effectively suppressing the interference of low-quality data and significantly improving the robustness and adaptability of feature fusion. For example: like Figure 3 As shown in the figure, in the dynamic gated cross attention mechanism, the weight calculation combines feature relevance and modality reliability, essentially coupling the influencing factors of the two through a mathematical structure. The weight of the traditional attention mechanism is determined only by the semantic association between features, and the formula is:

[0057] The dynamic gated crisscross attention mechanism upgrades the crisscross attention mechanism by introducing the modal confidence g:

[0058] in, is the diagonal matrix corresponding to the confidence level. This improved method optimizes the attention weight distribution from two aspects: feature relevance and modality reliability: 1. Feature Correlation: (query vector after image feature mapping) and The dot product of (the key vector after numerical feature mapping) measures the semantic association between image features and numerical features.

[0059] (query vector after numerical feature mapping) and The dot product of (the key vector after image feature mapping) captures the logical association between numerical features and image features.

[0060] 2. Modal reliability: When calculating When the value of the numerical mode is quilt Weighted, if the numerical modal confidence Low (such as sensor failure causing data jump), then The weight of is attenuated, weakening the influence of unreliable numerical features; Similarly, Medium Image Mode quilt Weighted, when the image is blurred by dust When is low, its feature weight decreases accordingly.

[0061] By regularizing the constraints on the modal confidence Avoid the model from being overly dependent on a certain modality and losing the advantages of multimodal fusion.

[0062] Finally, the two attention outputs are concatenated in the channel dimension and then feature mapped through a nonlinear activation function and a fully connected layer:

[0063] in, is the output mapping matrix; Concat represents the feature concatenation operation in the channel dimension; ReLU is the nonlinear activation function.

[0064] The final output It is the core achievement of the dynamic gated cross-attention mechanism, which not only contains the dominant information of high-quality modalities, but also balances the complementary value of the two modalities, laying a reliable multimodal feature foundation for subsequent processing.

[0065] Step 4: Use the hidden clustering deconstruction mechanism to perform time series processing on the fused feature sequence. By monitoring the displacement changes of the hidden state clustering center of the time series model and combining it with the dynamic threshold algorithm to identify the working condition change points, the time series data can be adaptively segmented. In this embodiment, an optimized sliding window strategy is applied to the fused feature sequence in the latent clustering and time series decomposition stage, the latent state representation is extracted through LSTM, the displacement change of the latent cluster center is monitored and combined with the dynamic threshold algorithm to achieve accurate identification of working condition change points and adaptive segmentation of time series data.

[0066] Specifically, to address the non-stationary nature of industrial time series data, this paper proposes a Latent State Clustering Decomposition Mechanism (LSCDM) for time series data. This mechanism achieves adaptive segmentation by monitoring the displacement of LSTM latent state cluster centers and combines dynamic thresholds to identify operating condition change points, thus avoiding the failure of traditional fixed windows to capture mutation points. This mechanism employs a two-layer attention architecture: local attention enhances mutation feature extraction within dynamic segments, while global attention models long-term dependencies across segments, effectively addressing the prediction lag caused by the non-stationarity of industrial data. For example: 1. Take the feature sequence output by the dynamic gated crisscross attention mechanism as input:

[0067] in, : Represents the fusion feature output by the dynamic gated cross attention mechanism at the tth time step, which is the comprehensive feature after the concatenation and nonlinear transformation of the bidirectional attention results of "image→value" and "value→image" (including the complementary information of the image modality and the numerical modality, and the low-quality data interference has been filtered out by the dynamic gating mechanism).

[0068] : The fusion features of the t-th time step Defined as the input features of the time series processing module .

[0069] : t is the time step index, T is the total length of the sequence, indicating that the entire input is a feature sequence arranged in chronological order , that is, the A temporal feature sequence composed by time dimension.

[0070] 2. Use LSTM to process the fused feature sequence and generate the hidden state of each time step:

[0071] in, : The LSTM hidden state at the current time step (time t) is a compressed representation of the "current fusion features + historical time series information", including the key time series dynamic features up to time t; LSTM(·): The computational unit of a long short-term memory network. It controls the inflow, retention, and output of information through a gating mechanism (input gate, forget gate, and output gate), solving the long-range dependency forgetting problem of traditional recurrent neural networks (RNNs). : The LSTM hidden state of the previous time step (t-1 moment), which contains the historical time series information up to t-1 moment.

[0072] 3. For a fixed segment at time t, calculate the cluster center of the segment before time t:

[0073] in, : represents the cluster center before time t in a fixed segment; k: the starting time step of the current fixed segment (i.e., the segment starts at time k); t: The current time step to be evaluated (it is necessary to determine whether time t still belongs to the current segment); : The hidden state output by LSTM at time j (including the temporal dynamic information up to time j).

[0074] 4. Calculate the hidden state and Euclidean distance (displacement):

[0075] in, It represents the cluster center of the hidden state in the current fixed segment as of time t-1, which is the "typical feature benchmark" of the historical hidden state in the segment.

[0076] It serves as a "quantitative indicator" for judging operating condition changes within the latent state clustering deconstruction mechanism: by measuring the deviation of the current latent state from the historical benchmark, it accurately captures mutation points in time series data, providing a basis for adaptive segmentation of time series data. This dynamic monitoring approach overcomes the limitations of traditional fixed windows, which are unable to flexibly respond to sudden changes in operating conditions. It ensures that subsequent hierarchical attention models can target the operating condition characteristics of different segments, improving the predictive performance of non-stationary time series data.

[0077] 5. Compare the cluster center offset of the current stage with the dynamic threshold to determine whether the current node is a segmentation point:

[0078] in, is the dynamic threshold at time t, used to determine the current offset Whether it is an "abnormal fluctuation" (i.e., a sudden change in operating conditions) is essentially based on the statistical characteristics of historical offsets:

[0079] in : represents the set of all historical offsets from time 2 to time t-1 (i.e. ), these offsets reflect the historical fluctuations of the hidden state from the cluster center under normal working conditions; : The mean of the historical offset set, reflecting the typical level of offset under normal operating conditions (central trend); : The standard deviation of the historical offset set, reflecting the fluctuation range (discreteness) of the offset under normal working conditions; : A hyperparameter (usually 2 to 3) used to adjust the contribution of the standard deviation to the threshold (the larger the weight, the more tolerant the threshold is to historical fluctuations).

[0080] This dynamic threshold based on historical statistical characteristics overcomes the defect that fixed thresholds (such as presetting a constant) cannot adapt to dynamic changes in working conditions (for example, fixed thresholds may be too loose under stable working conditions, resulting in missed judgments, and may be too strict under fluctuating working conditions, resulting in misjudgments). It makes time series segmentation more accurate, provides a reliable basis for subsequent hierarchical attention models to model different segments, and ultimately improves the prediction robustness of non-stationary time series data.

[0081] Step 5: Use local attention mechanism to extract fine-grained features within the adaptive segment, and use global attention mechanism to integrate information across segments to model long-term dependencies; In this embodiment, during the hierarchical attention feature modeling stage, a local attention mechanism is deployed within the identified dynamic segments to extract and enhance fine-grained features, accurately capturing local mutation patterns; at the same time, a global attention mechanism is used to enable cross-segment information interaction and integration, effectively modeling long-term dependencies and constructing a complete temporal representation. For example: 1. For each dynamic segment S i , firstly apply the local attention (LocalAtt) mechanism to extract fine-grained features within the segment:

[0082] Among them, S i It represents the i-th local time series segment output by the hidden state clustering deconstruction mechanism, which is the subsequence obtained after the time series data is adaptively segmented. It comes from the hidden state offset in step 4. With dynamic threshold When the comparison Time-triggered segmentation divides the time series data into multiple continuous subsequences , each S i Corresponding to a relatively stable operating condition (such as "normal combustion stability period", "coal injection amount adjustment transition period", etc.).

[0083] 2. Global modeling of representations of all segments: In the fixed window L, the local features of all segments Perform global modeling:

[0084] Global attention calculation:

[0085] Local modeling first extracts the core features of each operating condition segment, and then global modeling integrates inter-segment correlations. This allows the model to accurately capture local details while grasping overall time series trends, significantly improving its modeling capabilities for non-stationary industrial time series data. Global attention calculation integrates information from multiple local segments within a window, capturing long-range dependencies across segments (such as causal relationships between different operating conditions), providing global perspective feature support for subsequent coal injection rate prediction.

[0086] Among them, the fixed window L refers to a time range containing multiple continuous local segments (such as a window consisting of the last three stable operating conditions), which is used to limit the time range of global modeling, avoid information redundancy caused by too long time series, and ensure focus on recent operating condition associations.

[0087] : represents n local segment features within the window L, each It is achieved by paying local attention (LocalAtt) to the i-th stable working condition segment ( ) after processing (i.e. ), which contains the key time step information within the segment (such as the core characteristics of the "normal combustion segment", the key parameter changes of the "adjustment segment", etc.).

[0088] : Concatenate the n local features within the window in the sequence dimension to form a global feature sequence, integrating the information of all local segments within the window.

[0089] : A learnable global weight matrix is ​​used to map the concatenated global feature sequence to the "query", "key" and "value" spaces, respectively, to achieve the unification of feature dimensions and the extraction of global semantics.

[0090] : The generated global query, key, and value vectors are used for subsequent global attention calculations to measure the degree of association between different local segments (such as the dependency between the "normal combustion segment" and the "overburning segment").

[0091] 3. Compress the output of global attention into a fixed-dimensional feature representation through pooling operation:

[0092] Among them, GlobalAtt: The output result of global attention is all local segment features within the fixed window L ( ) after modeling cross-segment associations. Its dimensionality is typically related to the number of segments n in the window (e.g., n segments correspond to n feature vectors, with a dimension of [B,n,d], where B is the batch size and d is the feature dimension). Therefore, it changes with the number of segments in the window (i.e., "variable dimensionality").

[0093] Pool(·): Pooling operations (commonly used are average pooling, max pooling, or self-attention pooling) compress the GlobalAtt sequence dimensions, eliminating dimensionality differences caused by the number of segments within a window. For example, average pooling calculates the average of all feature vectors in the GlobalAtt sequence to obtain a composite feature; max pooling selects the most significant feature vector in the sequence (e.g., the dimension with the largest value), preserving key information.

[0094] : The fixed-dimensional feature representation obtained after pooling has a dimension that is independent of the number of segments in the window (e.g., fixed to [B, d]). It is a "condensed version" of the global features in the window and contains the core information associated with cross-segments within the window.

[0095] Step 6: Perform nonlinear transformation on the global time series features after hierarchical attention processing and output the predicted value of the rotary kiln coal injection amount.

[0096] In this embodiment, during the end-to-end prediction and decision support phase, the global time series features processed by hierarchical attention are input into the prediction layer, and a high-precision prediction value of the rotary kiln coal injection rate is generated through nonlinear transformation, providing real-time and accurate process parameter optimization suggestions for the production process, thereby realizing intelligent control and optimization of the rotary kiln combustion process. For example: The end-to-end prediction and decision support phase can be broken down into three core steps: "prediction layer modeling," "high-precision prediction," and "decision support generation," ultimately enabling intelligent control of the rotary kiln combustion process. 1. Input global timing characteristics The input features are fixed-dimensional features obtained by hierarchical attention processing and compression by pooling operations. (Right now ). This feature integrates the following key information: The local core characteristics of a single stable operating segment (through Extraction, such as key values ​​of temperature and pressure in the “normal combustion section”); Cross-segment correlation of multiple consecutive operating conditions (captured through global attention GlobalAtt, such as the impact of the "coal injection amount adjustment segment" on the subsequent "combustion stabilization segment"); Global trends within the window (retained through pooling, such as the overall combustion efficiency change over the last three operating periods).

[0097] 2. Nonlinear transformation of the prediction layer The prediction layer usually uses a multi-layer perceptron (MLP) or a fully connected network with an activation function (such as ReLU, Swish). Perform a nonlinear transformation:

[0098] in is the predicted value of coal injection amount. The role of nonlinear transformation is: Exploring the complex mapping relationship between global features and pulverized coal injection rate (e.g., "combustion temperature decrease + pressure increase" may correspond to "need to increase pulverized coal injection rate"); Fitting nonlinear and strongly coupled process characteristics in industrial scenarios (such as the nonlinear relationship between coal injection amount and parameters such as rotary kiln speed and material humidity).

[0099] 3. High-precision prediction and decision support High-precision prediction: Through model training (such as adjusting the weight of MLP with historical data), The prediction is as close as possible to the actual coal injection amount y (through loss function such as MSE optimization), and finally a high-precision prediction is achieved within the industrial allowable error range (such as error <2%).

[0100] Decision support generation: based on predicted values and production goals (such as "ensuring stable combustion temperature" and "reducing energy consumption") to generate specific process optimization suggestions: If the predicted coal injection rate is too low (which may result in insufficient temperature), it is recommended to "increase the coal injection rate by X kg / h"; if the predicted coal injection rate is too high (which may result in energy waste or overburning), it is recommended to "reduce the coal injection rate by Y kg / h".

[0101] Real-time performance: Since the previous feature extraction (hierarchical attention) and prediction layer calculations are both end-to-end neural network forward propagation and can be completed in milliseconds, they can meet the real-time requirements of industrial production (such as updating predictions and recommendations every 5 seconds).

[0102] 4. Intelligent control and optimization Predicted value The decision suggestions will be fed back to the rotary kiln control system (such as PLC, DCS) to achieve two optimization modes: Passive optimization: The operator manually adjusts the coal injection valve opening according to the suggestions; Active optimization: The system automatically and dynamically adjusts the coal injection amount according to the predicted value (for example, when "combustion efficiency decreases" is predicted, the coal injection amount is fine-tuned in advance to maintain stability), ultimately achieving the goal of "energy saving and consumption reduction, and reducing failures" (such as reducing coal consumption per ton of clinker and reducing the risk of kiln shutdown due to unstable combustion).

[0103] like Figure 4As shown in the figure, it can be seen from the fitting curve that the model shows excellent results in the time series data prediction task. The scatter plot of the true value and the predicted value is closely aligned with the ideal prediction line. The RMSE (0.0729) and MAE (0.0262) values ​​are low, indicating that the prediction deviation is small. The R² reaches 0.9600, indicating that the model can explain 96% of the data fluctuations and can effectively capture the changing trend of time series data. It has high reliability in modeling and prediction of non-stationary time series data of rotary kiln system. Although there are a few sparse points or room for optimization, the overall adaptation is suitable for actual application needs, providing strong support for the prediction of complex time series scenarios.

[0104] like Figure 5 As shown in the figure, the time series prediction effect diagram of coal injection amount shows the good performance of the model of the present invention. In the time series prediction curve above, the blue true value and the red predicted value are highly consistent with each other in overall trend, and the pink prediction interval effectively covers the fluctuation. The RMSE (0.0729) and MAE (0.0262) are low, and the R² reaches 0.96, indicating that the model accurately captures data changes. In the error analysis diagram below, the errors mostly fluctuate slightly around the 0 line, with only a few peaks at a few moments, indicating that the model prediction is stable. Although there may be errors at extreme moments, the overall prediction adaptability of the rotary kiln coal injection amount is strong, which can provide reliable support for process monitoring and decision-making.

[0105] In summary, this paper proposes a dynamic gated cross-attention mechanism, which uses a lightweight gating module to assess the quality of data from different modalities. For image modalities, this mechanism extracts key features such as flame brightness, clarity, smoke concentration, and inter-frame consistency; for numerical modalities, it extracts process-related indicators such as rate of change, volatility, and combustion state encoding. These features are then fed into their respective confidence assessment modules, which quantify and output a confidence score for the modal data, effectively reflecting the "health" of the data.

[0106] The innovation of this solution lies in deeply embedding the modal confidence score into the cross-attention mechanism, so that the weight calculation simultaneously integrates feature relevance and modal reliability. Specifically, when the system detects through the confidence assessment model that the confidence of the combustion image has decreased due to high-temperature dust, or when the numerical sensor has data anomalies due to aging, it will dynamically adjust the attention allocation based on the confidence score - weighting the Key and Value through the confidence diagonal matrix, automatically enhancing the attention to the other high-quality modality, and realizing the adaptive fusion of data quality perception. The confidence assessment process is optimized through an end-to-end loss function, and regularization constraints are introduced to avoid the model's over-reliance on a single modality, ensuring a balance between the complementary value of image morphological features and process values ​​in the prediction of coal injection amount.

[0107] Example 2 Based on the same concept, the present invention also proposes a rotary kiln coal injection amount prediction device based on multimodal data dynamic gating, comprising: A data acquisition module is used to collect multimodal data during the rotary kiln process, wherein the multimodal data includes at least combustion image data and process numerical data; the process numerical data includes at least combustion state parameters and historical coal injection amount data, and the combustion state parameters include parameters corresponding to under-combustion, normal combustion and over-combustion states; Feature extraction module, used to extract spatial features from combustion image data and perform high-dimensional feature encoding on process numerical data; The multimodal fusion module uses a dynamic gated cross-attention mechanism to fuse spatial features and high-dimensional features. The dynamic gated cross-attention mechanism generates modal confidence by evaluating the quality of different modal data and performs adaptive weighted fusion of different modal features based on the modal confidence. The time series processing module uses a hidden clustering deconstruction mechanism to perform time series processing on the fused feature sequence. The hidden clustering deconstruction mechanism monitors the displacement changes of the hidden state clustering center of the time series model and combines it with a dynamic threshold algorithm to identify the working condition change points, thereby achieving adaptive segmentation of the time series data. Hierarchical attention module, including local attention mechanism and global attention mechanism. The local attention mechanism is used to capture local mutation features and small state transitions within the adaptive segment, while the global attention mechanism is used to model long-term dependencies and global change trends through cross-segment information interaction; The prediction module performs nonlinear transformation on the global temporal features after hierarchical attention processing and outputs the predicted value of the rotary kiln coal injection amount.

[0108] In this embodiment, the feature extraction module includes: An image feature extraction unit is used to extract the brightness features, clarity features, smoke concentration features, and inter-frame consistency features of combustion image data through a deep learning network; The numerical feature extraction unit is used to extract the process parameter change rate characteristics, short-term volatility characteristics and combustion state coding characteristics of process numerical data through a deep learning network.

[0109] In this embodiment, the multimodal fusion module includes: a confidence evaluation unit, configured to generate an image modal confidence level based on the features output by the image feature extraction unit, and to generate a numerical modal confidence level based on the features output by the numerical feature extraction unit; The dynamic weighting unit is used to construct a confidence weighting matrix based on the image modal confidence and the numerical modal confidence. The confidence weighting matrix is ​​used to perform weighted adjustment on the key features and value features in the feature fusion process to achieve adaptive enhancement of high-quality modal features.

[0110] In this embodiment, the confidence assessment unit performs nonlinear mapping on image features and numerical features through a multi-layer perceptron to generate image modal confidence and numerical modal confidence; the multimodal fusion module also includes a regularization unit for balancing the image modal confidence and numerical modal confidence through regularization constraints to avoid excessive reliance on a single modality.

[0111] In this embodiment, the timing processing module includes: Hidden state extraction unit, used to process the fused feature sequence through the long short-term memory network to generate a temporal hidden state; Cluster center calculation unit, used to calculate the cluster center of the time series hidden state and the displacement between the current hidden state and the historical cluster center in real time; A change point identification unit is used to compare the displacement with the dynamic threshold and identify the working condition change point to complete the adaptive segmentation of the time series data. The dynamic threshold is determined based on the statistical characteristics of the historical displacement; The dynamic threshold is the sum of the mean of the historical displacements and the preset sensitivity coefficient multiplied by the standard deviation of the historical displacements.

[0112] Example 3 This embodiment also provides an electronic device, referring to Figure 6 , includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.

[0113] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.

[0114] Memory 404 may include a large-capacity memory 404 for data or instructions. By way of example, and not limitation, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to the data processing device. In certain embodiments, memory 404 is non-volatile memory. In certain embodiments, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. In appropriate circumstances, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0115] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .

[0116] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the rotary kiln coal injection amount prediction methods based on multimodal data dynamic gating in the above embodiments.

[0117] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .

[0118] Transmission device 406 can be used to receive or transmit data via a network. Specific examples of such networks may include wired or wireless networks provided by the electronic device's communications provider. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0119] The input and output devices 408 are used to input or output information.

[0120] Example 4 This embodiment also provides a readable storage medium, which stores a computer program. The computer program includes a program code for controlling a process to execute a process. The process includes a rotary kiln coal injection amount prediction method based on multimodal data dynamic gating according to embodiment one.

[0121] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.

[0122] In general, various embodiments may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0123] The embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components that are configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, it should be noted at this point that, for example, Figure 1 Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.

[0124] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0125] The above embodiments merely illustrate several embodiments of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A method for predicting the amount of coal injected into a rotary kiln based on dynamic gating of multimodal data, characterized in that: The following steps are involved: Collecting multimodal data during the rotary kiln process, the multimodal data including at least combustion image data and process numerical data; the process numerical data including at least combustion state parameters and historical coal injection amount data, the combustion state parameters including parameters corresponding to under-combustion, normal combustion, and over-combustion states; Performing spatial feature extraction on the combustion image data and performing high-dimensional feature encoding on the process numerical data; A dynamic gated cross-attention mechanism is used to fuse the spatial features and the encoded high-dimensional features. The dynamic gated cross-attention mechanism generates a modality confidence by evaluating the quality of different modal data and adaptively weights the different modal features based on the modality confidence. The hidden clustering deconstruction mechanism is used to perform time series processing on the fused feature sequence. By monitoring the displacement changes of the hidden state clustering center of the time series model and combining it with the dynamic threshold algorithm to identify the working condition change points, the adaptive segmentation of the time series data is achieved. A local attention mechanism is used to extract fine-grained features within the adaptive segment, while a global attention mechanism is used to integrate information across segments to model long-term dependencies. The global time series features processed by hierarchical attention are nonlinearly transformed to output the predicted value of the rotary kiln coal injection amount.

2. A rotary kiln coal injection amount prediction method based on multimodal data dynamic gating according to claim 1, characterized in that: Spatial feature extraction of combustion image data includes: Extract the brightness features, clarity features, smoke concentration features and inter-frame consistency features of the combustion image; High-dimensional feature encoding of process numerical data includes: The process parameter change rate characteristics, short-term volatility characteristics and combustion state coding characteristics of the process values ​​are extracted.

3. The method for predicting the amount of coal injected into a rotary kiln based on multimodal data dynamic gating according to claim 1, wherein: In the dynamic gated cross-attention mechanism, the modal confidence is generated by nonlinear mapping of the extracted features through a multi-layer perceptron, and the confidence weights of different modalities are balanced through regularization constraints.

4. A rotary kiln coal injection rate prediction method based on multimodal data dynamic gating according to any one of claims 1 to 3, characterized in that: In the latent clustering deconstruction mechanism, the latent state of the time series model is extracted through a long short-term memory network. The identification of the operating condition change point is based on whether the displacement between the current latent state and the historical cluster center exceeds a dynamic threshold. The dynamic threshold is the sum of the mean of the historical displacement and the preset sensitivity coefficient multiplied by the standard deviation of the historical displacement.

5. An end-to-end prediction device for coal injection amount of a rotary kiln, characterized in that: include: A data acquisition module, configured to collect multimodal data during the rotary kiln process, wherein the multimodal data includes at least combustion image data and process numerical data; The process numerical data at least includes combustion state parameters and historical coal injection amount data, wherein the combustion state parameters include parameters corresponding to under-combustion, normal combustion and over-combustion states; A feature extraction module, configured to extract spatial features from the combustion image data and perform high-dimensional feature encoding on the process numerical data; a multimodal fusion module that fuses the spatial features and high-dimensional features using a dynamic gated cross-attention mechanism that generates a modality confidence by evaluating the quality of data from different modalities and adaptively weights the features of different modalities based on the modality confidence; The time series processing module uses a hidden clustering deconstruction mechanism to perform time series processing on the fused feature sequence. The hidden clustering deconstruction mechanism monitors the displacement changes of the hidden state clustering center of the time series model and combines it with a dynamic threshold algorithm to identify the working condition change points, thereby realizing adaptive segmentation of the time series data. A hierarchical attention module, including a local attention mechanism and a global attention mechanism. The local attention mechanism is used to capture local mutation features and small state transitions within the adaptive segment, and the global attention mechanism is used to model long-term dependencies and global change trends through cross-segment information interaction; The prediction module performs nonlinear transformation on the global temporal features after hierarchical attention processing and outputs the predicted value of the rotary kiln coal injection amount.

6. The end-to-end prediction device for coal injection amount of a rotary kiln according to claim 5, characterized in that: The feature extraction module includes: An image feature extraction unit, configured to extract brightness features, clarity features, smoke concentration features, and inter-frame consistency features of the combustion image data through a deep learning network; The numerical feature extraction unit is used to extract the process parameter change rate characteristics, short-term volatility characteristics and combustion state coding characteristics of the process numerical data through a deep learning network.

7. The end-to-end prediction device for coal injection amount of a rotary kiln according to claim 6, characterized in that: The multimodal fusion module includes: a confidence evaluation unit, configured to generate an image modal confidence metric based on the features output by the image feature extraction unit, and to generate a numerical modal confidence metric based on the features output by the numerical feature extraction unit; A dynamic weighting unit is used to construct a confidence weighting matrix based on the image modal confidence and the numerical modal confidence, and to perform weighted adjustment on key features and value features in the feature fusion process through the confidence weighting matrix to achieve adaptive enhancement of high-quality modal features.

8. The end-to-end prediction device for coal injection amount of a rotary kiln according to claim 7, characterized in that: The confidence assessment unit performs nonlinear mapping on the image features and numerical features through a multi-layer perceptron to generate the image modal confidence and numerical modal confidence; the multimodal fusion module also includes a regularization unit for balancing the image modal confidence and numerical modal confidence through regularization constraints to avoid excessive reliance on a single modality.

9. The end-to-end prediction device for coal injection amount of a rotary kiln according to claim 5, characterized in that: The timing processing module includes: Hidden state extraction unit, used to process the fused feature sequence through the long short-term memory network to generate a temporal hidden state; A cluster center calculation unit, used to calculate the cluster center of the temporal hidden state and the displacement between the current hidden state and the historical cluster center in real time; a change point identification unit, configured to compare the displacement with a dynamic threshold value to identify a change point of the operating condition so as to complete adaptive segmentation of the time series data, wherein the dynamic threshold value is determined based on the statistical characteristics of the historical displacement; The dynamic threshold is the sum of the mean of the historical displacements and the preset sensitivity coefficient multiplied by the standard deviation of the historical displacements.

10. A readable storage medium, characterized in that: A computer program is stored in the readable storage medium, and the computer program includes a program code for controlling a process to execute a process, and the process includes a rotary kiln coal injection amount prediction method based on multimodal data dynamic gating according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Coal injection quantity prediction method based on thermal parameters of rotary kiln

    CN116929052A

  • Zinc rotary kiln digital twin model parameter updating method, device, equipment and medium

    CN118551802A

  • Construction and application method of rotary kiln combustion state recognition model based on bimodal

    CN119068200A

  • Pyrolysis control apparatus and method using image information of raw materials and products

    US20250188352A1

  • Method and system for predicting the outside temperature of a rotary KILN shell

    WO2024022793A1

Cited By

  • Multi-modal emotion recognition method based on electroencephalogram signal and facial expression fusion

    CN121861735A

  • Industrial flue gas multi-modal prediction method based on gating multi-dimensional information fusion converter

    CN122224313A